Faisal Rahutomo scite author profile

Faisal Rahutomo

4Publications

51Citation Statements Received

9Citation Statements Given

How they've been cited

127

How they cite others

Affiliations

Sebelas Maret University, State University of Malang, Kumamoto University

Publications

Order By: Most citations

Test collection recycling for semantic text similarity

Rahutomo

Kitasuka

Aritsugi

2012

View full text Add to dashboard Cite

Semantic text similarity (STS) uses specific test collections as its performance evaluation measurement. The test collections consist of text pairs with the same meaning even though in different text form. The existence is scarce compared with information retrieval (IR) test collections. This paper investigates the possibility to reuse IR test collections for STS tasks. Text pairs are derived from the relevant pair of IR test collections. Latent semantic analysis (LSA) and explicit semantic analysis (ESA) evaluate Glasgow's test collections, which are provided by ACM SIGIR community. Jaccard index measures the lexical similarity. Recall metric measures retrievability of recycling test collection with two existing test collections, Microsoft research paraphrase corpus and Microsoft research video description corpus, as evaluation baselines. Evaluation yields a promising outcome; the evaluated test collections have low Jaccard index and their recall values between the two baselines.

show abstract

Study of hoax news detection using naïve bayes classifier in Indonesian language

Pratiwi

Asmara

Rahutomo

2017

View full text Add to dashboard Cite

Evaluasi Daftar Stopword Bahasa Indonesia

Rahutomo

Ririd

2019

JTIIK

View full text Add to dashboard Cite

Pada sistem temu kembali informasi berbentuk teks maupun text mining, terdapat proses pengindeksan. Teks diproses dengan tujuan mengintisarikan informasi berbentuk teks tersebut. Salah satu proses yang dilakukan adalah stopword filtering, beberapa kata yang tidak layak diindeks diabaikan berdasar sebuah daftar. Di dalam sistem berbahasa Indonesia, terdapat beberapa versi daftar stopword yang tersedia bebas. Penelitian ini bertujuan mengevaluasi daftar yang telah tersedia tersebut. Tujuan akhir dari penelitian ini adalah telaah daftar yang tersedia berdasarkan tata bahasa Indonesia, cara penyusunan, dan kebiasaan perambah internet. Dari hasil telaah diperoleh fakta bahwa daftar yang tersedia dibangun dengan analisis frekuensi kemunculan kata pada sebuah korpus (corpus) teks, tanpa memperhatikan jenis kata ataupun kebiasaan pengguna internet. Hasil lain penelitian ini adalah beberapa rekomendasi lebih lanjut bagi para peneliti di bidang ini ketika membutuhkan daftar stopword bahasa Indonesia, yaitu daftar yang memperhatikan jenis kata dan kebiasaan pengguna internet melalui mesin perambah yang tersedia.AbstractMost of text-based information retrieval system uses indexing process. The system processes the texts in order to obtain the information essence. One of the process is stopword filtering, several words are being ignored based on a stopword list. Several Indonesian stopword list are available openly. Therefore, this paper evaluates the available lists based on Indonesian formal grammar, its preparation technique, and internet surfer habit. The results show all of the list are developed by term frequency analysis based on a text corpus. This paper also provides several recommendations for researcher both in text mining and text-based information retrieval field, developing stoplist by the word type and internet surfer habit.

show abstract

Eksperimen Naïve Bayes Pada Deteksi Berita Hoax Berbahasa Indonesia

Rahutomo¹,

Pratiwi²,

Ramadhani³

2019

Jurnal PKOP

View full text Add to dashboard Cite

Website and blog are popular as a media to spread news. The validity of an article of news’s can either be valid or fake. A fake article of news is usually called a hoax news article. The purpose of making hoax news is to persuade, manipulate, affect to people to do something that contradicts or prevents the right action. A hoax news usually used threats or misleading information to make them believe things that are not real. This research proposes an experiment using naïve Bayes to detect hoax news in Bahasa Indonesia. In this research, we use our own dataset consisting of a total of 600 valid and hoax articles. We asked three reviewers to conduct manual classification for our dataset. Final tagging was obtained by adopting the maximum score from the three reviewers. In our experiment, we show that naïve Bayes can classify Indonesian online news articles with term frequency feature using the PHP-ML library component’s. We obtained an accuracy is 82.6% with static testing and 68.33% with dynamic testing. We give free access to the dataset so the future research can replicate, comparing the result and make a baseline testing.Keywords : Hoax News Detection, Naïve Bayes Classifier.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

customersupport@researchsolutions.com

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Faisal Rahutomo

Test collection recycling for semantic text similarity

Study of hoax news detection using naïve bayes classifier in Indonesian language

Evaluasi Daftar Stopword Bahasa Indonesia

Eksperimen Naïve Bayes Pada Deteksi Berita Hoax Berbahasa Indonesia

Contact Info

Product

Resources

About