Evaluating M-Pajak Service through Human-Validated IndoRoBERTa Analysis of Google Play Reviews

Authors

  • Hana Tamara Putri Universitas Batanghari, Indonesia
  • Mufidah Universitas Batanghari, Indonesia

DOI:

https://doi.org/10.61255/jeemba.v4i4.1356

Keywords:

M-Pajak, Mobile Tax Application, IndoRoBERTa, Sentiment Analysis, Digital Tax Administration

Abstract

Purpose – This study evaluates user-perceived experience in Indonesia's M-Pajak mobile tax application using Google Play reviews and identifies temporal, sentiment, and lexical patterns in users' expressed evaluations.

Design/methodology/approach – The study analyzed a relevance-ranked, retrievable corpus scraped on 24 May 2026. Of 7,009 retrieved reviews, 6,980 valid reviews from 2021–2026 were retained. IndoRoBERTa sentiment labels were evaluated against 1,500 reviews selected through proportionate year-stratified sampling and manually adjudicated by two annotators. The analysis combined temporal ratings, temporal sentiment, rating–sentiment alignment, and transparent bigram-to-dimension coding.

Findings/Results – Human annotators achieved 97.20% agreement (Cohen's kappa = 0.9132), while IndoRoBERTa achieved 94.20% agreement with the human gold standard (kappa = 0.8243). Within the retrieved corpus, 78.81% of reviews were classified as negative, and later-year review shares were increasingly concentrated in low ratings and negative sentiment. These patterns describe expressed review-based evaluation and do not establish that overall application quality or population-level satisfaction necessarily declined. Authentication, OTP, login, NPWP, EFIN, payment, and system-error barriers dominated negative expressions.

Originality/Value – The study offers a human-validated, multi-layer review-analysis framework for evaluating mobile public services, makes the lexical grouping procedure auditable, and translates recurrent complaints into operational service-improvement priorities.

Abstract views: 5 , PDF downloads: 4

Downloads

Download data is not yet available.

References

Al-Natour, S., & Turetken, O. (2020). A comparative assessment of sentiment analysis and star ratings for consumer reviews. International Journal of Information Management, 54, 102132. https://doi.org/10.1016/j.ijinfomgt.2020.102132

Artstein, R., & Poesio, M. (2008). Inter-Coder Agreement for Computational Linguistics. Computational Linguistics, 34(4), 555–596. https://doi.org/10.1162/coli.07-034-R2

Budianto, T. A. C., Dermawan, B. A., & Jajuli, M. (2025). Analisis Sentimen Ulasan Aplikasi M-Pajak pada Google Play Store Menggunakan XGBoost. Jurnal Sistem Informasi Dan Aplikasi, 3(1), 11–25. https://doi.org/10.52958/jsia.v3i1.12440

Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20(1), 37–46. https://doi.org/10.1177/001316446002000104

Dąbrowski, J., Letier, E., Perini, A., & Susi, A. (2022). Analysing app reviews for software engineering: a systematic literature review. Empirical Software Engineering, 27(2), 43. https://doi.org/10.1007/s10664-021-10065-7

Dechamps, S., Simonofski, A., & Burnay, C. (2025). Citizen-centricity in digital government: A theoretical and empirical typology. Government Information Quarterly, 42(1), 102005. https://doi.org/10.1016/j.giq.2024.102005

Duwairi, R., & El-Orfali, M. (2014). A study of the effects of preprocessing strategies on sentiment analysis for Arabic text. Journal of Information Science, 40(4). https://doi.org/10.1177/0165551514534143

Genc-Nayebi, N., & Abran, A. (2017). A systematic literature review: Opinion mining studies from mobile app store user reviews. Journal of Systems and Software, 125, 207–219. https://doi.org/10.1016/j.jss.2016.11.027

Gliniecka, Martyna. (2023). The Ethics of Publicly Available Data Research: A Situated Ethics Framework for Reddit. Social Media + Society, 9(3), 20563051231192020. https://doi.org/10.1177/20563051231192021

Google Play, G. P. (2026). M-Pajak [Mobile application listing]. https://play.google.com/store/apps/details?id=id.go.pajak.djp&hl=en-US

Hicks, S. A., Strümke, I., Thambawita, V., Hammou, M., Riegler, M. A., Halvorsen, P., & Parasa, S. (2022). On evaluation metrics for medical applications of artificial intelligence. Scientific Reports, 12(1), 5979. https://doi.org/10.1038/s41598-022-09954-8

Hunter, R. F., Gough, A., O’Kane, N., McKeown, G., Fitzpatrick, A., Walker, T., McKinley, M., Lee, M., & Kee, F. (2018). Ethical Issues in Social Media Research for Public Health. American Journal of Public Health, 108(3), 343–348. https://doi.org/10.2105/AJPH.2017.304249

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. In D. Scott, N. Bel, & C. Zong (Eds.), Proceedings of the 28th International Conference on Computational Linguistics (pp. 757–770). International Committee on Computational Linguistics. https://doi.org/10.18653/v1/2020.coling-main.66

Leem, B.-H., & Eum, S.-W. (2021). Using text mining to measure mobile banking service quality. Industrial Management & Data Systems, 121(5), 993–1007. https://doi.org/10.1108/IMDS-09-2020-0545

Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. https://doi.org/10.48550/arXiv.1907.11692

Maalej, W., Kurtanović, Z., Nabil, H., & Stanik, C. (2016). On the automatic classification of app reviews. Requirements Engineering, 21(3), 311–331. https://doi.org/10.1007/s00766-016-0251-9

Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M., & Gao, J. (2021). Deep Learning--based Text Classification: A Comprehensive Review. ACM Comput. Surv., 54(3). https://doi.org/10.1145/3439726

Nakamura, W. T., de Oliveira, E. C., de Oliveira, E. H. T., Redmiles, D., & Conte, T. (2022). What factors affect the UX in mobile apps? A systematic mapping study on the analysis of app store reviews. Journal of Systems and Software, 193, 111462. https://doi.org/10.1016/j.jss.2022.111462

Nandwani, P., & Verma, R. (2021). A review on sentiment analysis and emotion detection from text. Social Network Analysis and Mining, 11(1), 81. https://doi.org/10.1007/s13278-021-00776-6

Putra, A., & Fathurrahman, R. (2022). Improving the Quality of the Mobile Tax Service Apps in Indonesia: A Delphi Study. Budapest International Research and Critics Institute-Journal (BIRCI-Journal), 5(2), 8319–8330. https://doi.org/10.33258/birci.v5i2.4614

Ramzy, M., & Ibrahim, B. (2024). User satisfaction with Arabic COVID-19 apps: Sentiment analysis of users’ reviews using machine learning techniques. Information Processing & Management, 61(3), 103644. https://doi.org/10.1016/j.ipm.2024.103644

Samuel, Gabrielle, & Buchanan, Elizabeth. (2020). Guest Editorial: Ethical Issues in Social Media Research. Journal of Empirical Research on Human Research Ethics, 15(1–2), 3–11. https://doi.org/10.1177/1556264619901215

Saptono, P. B., Hodžić, S., Khozen, I., Mahmud, G., Pratiwi, I., Purwanto, D., Aditama, M. A., Haq, N., & Khodijah, S. (2023). Quality of E-Tax System and Tax Compliance Intention: The Mediating Role of User Satisfaction. Informatics, 10(1). https://doi.org/10.3390/informatics10010022

Setiyani, L., Indahsari, A. N., Monica, S., Poerwati, A. E., Ezrafel, A., & Rustam, R. (2022). Analysis of M-Tax Mobile Application Adoption on Tax Compliance in Indonesia Using Diffusion of Innovation Theory (DIT). Central European Management Journal, 30(4), 1202–1212. https://doi.org/10.57030/23364890.cemj.30.4.121

Sharma, S. K., Al-Badi, A., Rana, N. P., & Al-Azizi, L. (2018). Mobile applications in government services (mG-App) from user’s perspectives: A predictive modelling approach. Government Information Quarterly, 35(4), 557–568. https://doi.org/10.1016/j.giq.2018.07.002

Sidiq, D., Raharjo, T., & Trisnawaty, N. W. (2024). Towards Tax Administration 3.0: Bracing the Challenges in Mobile Application Development. Jurnal Informatika Ekonomi Bisnis, 6(2), 410–417. https://doi.org/10.37034/infeb.v6i2.915

Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427–437. https://doi.org/10.1016/j.ipm.2009.03.002

Symeonidis, S., Effrosynidis, D., & Arampatzis, A. (2018). A comparative evaluation of pre-processing techniques and their interactions for twitter sentiment analysis. Expert Systems with Applications, 110, 298–310. https://doi.org/10.1016/j.eswa.2018.06.022

Twizeyimana, J. D., & Andersson, A. (2019). The public value of E-Government – A literature review. Government Information Quarterly, 36(2), 167–178. https://doi.org/10.1016/j.giq.2019.01.001

Verkijika, S. F., & Neneh, B. N. (2021). Standing up for or against: A text-mining study on the recommendation of mobile payment apps. Journal of Retailing and Consumer Services, 63, 102743. https://doi.org/10.1016/j.jretconser.2021.102743

Wang, C., & Teo, T. S. H. (2020). Online service quality and perceived value in mobile government success: An empirical study of mobile police in China. International Journal of Information Management, 52, 102076. https://doi.org/10.1016/j.ijinfomgt.2020.102076

Wicaksono, P. T., Tjen, C., & Indriani, V. (2021). Improving the tax e-filing system in Indonesia: An exploration of individual taxpayers’ opinions. Jurnal Akuntansi Dan Auditing Indonesia, 25(2). https://doi.org/10.20885/jaai.vol25.iss2.art4

Wongso, W. (2023). indonesian-roberta-base-sentiment-classifier (Revision e402e46). Hugging Face. https://doi.org/10.57967/hf/0644

Yi, J., & Oh, Y. K. (2022). The informational value of multi-attribute online consumer reviews: A text mining approach. Journal of Retailing and Consumer Services, 65, 102519. https://doi.org/10.1016/j.jretconser.2021.102519

Zhang, Y., Fu, J., Lai, J., Deng, S., Guo, Z., Zhong, C., Tang, J., Cao, W., & Wu, Y. (2024). Reporting of Ethical Considerations in Qualitative Research Utilizing Social Media Data on Public Health Care: Scoping Review. J Med Internet Res, 26, e51496. https://doi.org/10.2196/51496

Downloads

Published

2026-07-18

How to Cite

Putri, H. T., & Mufidah, M. (2026). Evaluating M-Pajak Service through Human-Validated IndoRoBERTa Analysis of Google Play Reviews. Journal of Economics, Entrepreneurship, Management Business and Accounting, 4(4), 1145–1182. https://doi.org/10.61255/jeemba.v4i4.1356