Classification Analysis of Market Service Levy Payment Compliance in South Tangerang Using Random Forest Algorithm

Authors

  • Tiara Diba Department of Information Systems, Faculty of Science and Technology, Syarif Hidayatullah Jakarta University
  • Eri Rustamaji Department of Information Systems, Faculty of Science and Technology, Syarif Hidayatullah Jakarta University
  • Eva khudzaeva UIN Syarif Hidayatullah Jakarta
  • Nida’ul Hasanati Department of Information Systems, Faculty of Science and Technology, Syarif Hidayatullah Jakarta University
  • Nia Kumaladewi Department of Information Systems, Faculty of Science and Technology, Syarif Hidayatullah Jakarta University

DOI:

https://doi.org/10.25077/TEKNOSI.v12i2.2026.190-202

Keywords:

CRISP-DM, Classification,, Machine Learning, Random Forest

Abstract

Market service levies are one of the primary sources of Local Own-Source Revenue (PAD) used to develop and revitalize markets, as well as improve comfort, security, and local economic growth. Although the Electronic Transaction System for Local Government (ETPD) has been implemented in South Tangerang City to facilitate payments through various digital channels, the payment compliance rate remains low. In 2023, South Tangerang's market service levy revenue only reached 29.94% of the set target. Of this amount, only 16.86% of levy obligors paid regularly and on time, while 83.14% failed to make timely payments. This condition may hinder optimal regional revenue generation. This study aims to develop a classification model that can predict compliance in market service levy payments in South Tangerang City. The research utilizes the Random Forest algorithm and follows the Cross-Industry Standard Process for Data Mining (CRISP-DM) framework to classify market service levy data from November 2021 to November 2024 obtained through the ETPD system. The developed model classifies payment compliance into two categories: timely and delinquent. Analysis results using K-Fold Cross Validation show that the Random Forest model achieves strong performance metrics with 94.94% accuracy, 99.49% precision, 90.34% recall, 94.69% F1-score, and 95% AUC, indicating the model's excellent classification capability. These findings can serve as recommendations for the South Tangerang City government to formulate strategies for improving levy payment compliance. KEYWORDS

References

V. P. Tappi, “Analisis Pendapatan Asli Daerah (PAD) Kabupaten Jayapura,” Jurnal Ekonomu & Bisnis, vol. 12, no. 1, pp. 16–24, 2021, Accessed: Mar. 07, 2024. [Online]. Available: https://doi.org/10.55049/jeb.v12i1.66

Sumiati, Sulkarnain, Sitti Jamilah Amin, and Damirah, “Analisis Potensi Pasar Tradisional dalam Meningkatkan Perekonomian Daerah,” BALANCA : Jurnal Ekonomi dan Bisnis Islam, vol. 4, no. 2, pp. 8–15, Nov. 2023, doi: 10.35905/balanca.v4i2.4823.

S. Ruddin and A. Y. Nasution, “Analisis Revitalisasi Pasar Tradisional untuk Meningkatkan Pendapatan Daerah Kota Tangerang Selatan, Provinsi Banten,” Jurnal Mandiri : Ilmu Pengetahuan, Seni, dan Teknologi, vol. 3, no. 2, pp. 294–306, Dec. 2019, doi: 10.33753/mandiri.v3i2.91.

O. B. Saputri, “Analisis SWOT Transformasi Digital Transaksi Keuangan Pemerintah Daerah dalam Mendukung Inklusi Keuangan,” Jurnal Ekonomi Keuangan dan Manajemen, vol. 17, no. 3, pp. 482–494, 2021, Accessed: Mar. 18, 2024. [Online]. Available: https://doi.org/10.30872/jinv.v17i3.9339

L. Breiman, “Random Forests,” Mach Learn, vol. 45, no. 1, pp. 5–32, 2001, Accessed: Sep. 17, 2024. [Online]. Available: https://doi.org/10.1023/A:1010933404324

G. Biau and E. Scornet, “A Random Forest Guided Tour,” Test, vol. 25, no. 2, pp. 197–227, Jun. 2016, doi: 10.1007/s11749-016-0481-7.

Y. Ao, H. Li, L. Zhu, S. Ali, and Z. Yang, “The Linear Random Forest Algorithm and Its Advantages in Machine Learning Assisted Logging Regression Modeling,” J Pet Sci Eng, vol. 174, pp. 776–789, Mar. 2019, doi: 10.1016/j.petrol.2018.11.067.

C. Janiesch, P. Zschech, and K. Heinrich, “Machine Learning and Deep Learning,” Electronic Markets, vol. 31, no. 3, pp. 685–695, Sep. 2021, doi: 10.1007/s12525-021-00475-2.

Narayanan and Jayashree, “Implementation of Efficient Machine Learning Techniques for Prediction of Cardiac Disease using SMOTE,” in Procedia Computer Science, Elsevier B.V., 2024, pp. 558–569. doi: 10.1016/j.procs.2024.03.245.

Y. Kanoun, A. M. Aghbash, T. Belem, B. Zouari, and H. Mrad, “Failure Prediction in The Refinery Piping System Using Machine Learning Algorithms: Classification and Comparison,” in Procedia Computer Science, Elsevier B.V., 2024, pp. 1663–1672. doi: 10.1016/j.procs.2024.01.164.

C. Schröer, F. Kruse, and J. M. Gómez, “A Systematic Literature Review on Applying CRISP-DM Process Model,” in Procedia Computer Science, Elsevier B.V., 2021, pp. 526–534. doi: 10.1016/j.procs.2021.01.199.

P. Chapman et al., CRISP-DM 1.0 Step-By-Step Data Mining Guide. Amsterdam: SPSS Inc, 2000.

I. N. Hidayati, T. F. Kusumasari, and F. Hamami, “Comparison of Apache SparkSQL and Oracle Performance: Case Study of Data Cleansing Process,” International Journal on Informatics Visualization, vol. 6, no. 2, pp. 208–213, 2022, Accessed: Jul. 24, 2024. [Online]. Available: http://dx.doi.org/10.30630/joiv.6.1-2.928

F. Ridzuan, W. M. Nazmee, and W. Zainon, “A Review on Data Cleansing Methods for Big Data,” in Procedia Computer Science, Elsevier B.V., 2019, pp. 731–738. doi: 10.1016/j.procs.2019.11.177.

U. Suriani, “Penerapan Data Mining untuk Memprediksi Tingkat Kelulusan Mahasiswa Menggunakan Algoritma Decision Tree C4.5,” Journal of Computer and Information Systems Ampera, vol. 3, no. 2, pp. 55–66, 2023, doi: 10.51519/journalcisa.v4i2.393.

K. Potdar, T. S. Pardawala, and C. D. Pai, “A Comparative Study of Categorical Variable Encoding Techniques for Neural Network Classifiers,” Int J Comput Appl, vol. 175, no. 4, pp. 7–9, 2017, Accessed: Jul. 24, 2024. [Online]. Available: https://doi.org/10.5120/ijca2017915495

R. Larose and B. Coyle, “Robust data encodings for quantum classifiers,” Phys Rev A (Coll Park), vol. 102, no. 3, pp. 1–24, Sep. 2020, doi: 10.1103/PhysRevA.102.032420.

M. S. Irwanto, F. A. Bachtiar, and N. Yudistira, “Klasifikasi Aktivitas Manusia Menggunakan Algoritme Computed Input Weight Extreme Learning Machine dengan Reduksi Dimensi Principal Component Analysis,” Jurnal Teknologi Informasi dan Ilmu Komputer, vol. 9, no. 6, pp. 1195–1202, 2022, doi: 10.25126/jtiik.202295504.

F. Rachmawati, J. Jaenudin, N. B. Ginting, and P. Laksono, “Machine Learning for the Model Prediction of Final Semester Assessment (FSA) using the Multiple Linear Regression Method,” Jurnal Teknik Informatika, vol. 17, no. 1, pp. 1–9, May 2024, doi: 10.15408/jti.v17i1.28652.

F. Aldi, F. Hadi, N. A. Rahmi, and S. Defit, “Standardscaler’s Potential in Enhancing Breast Cancer Accuracy Using Machine Learning,” Journal of Applied Engineering and Technological Science, vol. 5, no. 1, pp. 401–413, 2023, Accessed: Jul. 25, 2024. [Online]. Available: https://doi.org/10.37385/jaets.v5i1.3080

Ridwan, E. H. Hermaliani, and M. Ernawati, “Penerapan Metode SMOTE Untuk Mengatasi Imbalanced Data pada Klasifikasi Ujaran Kebencian,” Computer Science , vol. 4, no. 1, pp. 80–88, 2024, Accessed: Jul. 25, 2024. [Online]. Available: https://doi.org/10.31294/coscience.v4i1.2990

E. Sutoyo and M. A. Fadlurrahman, “Penerapan SMOTE untuk Mengatasi Imbalance Class dalam Klasifikasi Television Advertisement Performance Rating Menggunakan Artificial Neural Network,” Jurnal Edukasi dan Penelitian Informatika, vol. 6, no. 3, pp. 379–385, 2020, Accessed: Sep. 13, 2024. [Online]. Available: http://dx.doi.org/10.26418/jp.v6i3.42896

V. R. Joseph, “Optimal Ratio for Data Splitting,” Stat Anal Data Min, vol. 15, no. 4, pp. 531–538, Aug. 2022, doi: 10.1002/sam.11583.

N. M. Kebonye, “Exploring The Novel Support Points-Based Split Method On A Soil Dataset,” Measurement (Lond), vol. 186, no. 58, pp. 1–4, Dec. 2021, doi: 10.1016/j.measurement.2021.110131.

C. Zucco, “Multiple learners combination: Bagging,” Encyclopedia of Bioinformatics and Computational Biology: ABC of Bioinformatics, vol. 1, no. 3, pp. 525–530, Jan. 2018, doi: 10.1016/B978-0-12-809633-8.20345-2.

D. Nishfi, I. Huda, C. Prianto, and R. M. Awangga, “Analisis Sentimen Perbandingan Layanan Jasa Pengiriman Kurir pada Ulasan Play Store Menggunakan Metode Random Forest dan Descision Tree,” Jurnal Ilmiah Informatika (JIF), vol. 11, no. 2, pp. 150–158, 2023, Accessed: Aug. 07, 2024. [Online]. Available: https://doi.org/10.33884/jif.v11i02.7952

F. Y. Pamuji and V. P. Ramadhan, “Jurnal Teknologi dan Manajemen Informatika Komparasi Algoritma Random Forest Dan Decision Tree untuk Memprediksi Keberhasilan Immunotheraphy,” Jurnal Teknologi dan Manajemen Informatika, vol. 7, no. 1, pp. 46–50, 2021, Accessed: Aug. 07, 2024. [Online]. Available: https://doi.org/10.26905/jtmi.v7i1.5982

L. Langsetmo et al., “Advantages and Disadvantages of Random Forest Models for Prediction of Hip Fracture Risk Versus Mortality Risk in the Oldest Old,” Journal of Bone and Mineral Research Plus, vol. 7, no. 8, pp. 1–11, Aug. 2023, doi: 10.1002/jbm4.10757.

T. M. Khoshgoftaar, M. Golawala, and J. Van Hulse, “An empirical study of learning from imbalanced data using random forest,” in Proceedings - International Conference on Tools with Artificial Intelligence, ICTAI, 2007, pp. 310–317. doi: 10.1109/ICTAI.2007.46.

E. Renata and M. Ayub, “Penerapan Metode Random Forest untuk Analisis Risiko pada Dataset Peer to Peer Lending,” Jurnal Teknik Informatika dan Sistem Informasi, vol. 6, no. 3, pp. 462–474, Dec. 2020, doi: 10.28932/jutisi.v6i3.2890.

V. Nguyen, “Bayesian optimization for accelerating hyper-parameter tuning,” in Proceedings - IEEE 2nd International Conference on Artificial Intelligence and Knowledge Engineering, AIKE 2019, Institute of Electrical and Electronics Engineers Inc., Jun. 2019, pp. 302–305. doi: 10.1109/AIKE.2019.00060.

Q. Yang et al., “Research on Surrogate Models and Optimization Algorithms of Compressor Characteristic Based on Digital Twins,” Journal of Engineering Research (Kuwait), vol. 5, no. 4, pp. 1–13, 2024, doi: 10.1016/j.jer.2024.01.025.

E. Elgeldawi, A. Sayed, A. R. Galal, and A. M. Zaki, “Hyperparameter tuning for machine learning algorithms used for arabic sentiment analysis,” Informatics, vol. 8, no. 4, pp. 1–21, Dec. 2021, doi: 10.3390/informatics8040079.

D. M. Belete and M. D. Huchaiah, “Grid Search in Hyperparameter Optimization of Machine Learning Models for Prediction of HIV/AIDS Test Results,” International Journal of Computers and Applications, vol. 44, no. 9, pp. 875–886, 2022, doi: 10.1080/1206212X.2021.1974663.

S. Andradóttir, “A Review of Random Search Methods,” International Series in Operations Research Management & Science, vol. 7, no. 4, pp. 277–292, 2014, doi: 10.1007/978-1-4939-1384-8__10.

S. Greenhill, S. Rana, S. Gupta, P. Vellanki, and S. Venkatesh, “Bayesian Optimization for Adaptive Experimental Design: A Review,” IEEE Access, vol. 8, no. 5, pp. 13937–13948, 2020, doi: 10.1109/ACCESS.2020.2966228.

D. Normawati and S. A. Prayogi, “Implementasi Naïve Bayes Classifier Dan Confusion Matrix Pada Analisis Sentimen Berbasis Teks Pada Twitter,” Jurnal Sains Komputer & Informatika (J-SAKTI, vol. 5, no. 2, pp. 697–711, 2021, Accessed: Jul. 31, 2024. [Online]. Available: http://dx.doi.org/10.30645/j-sakti.v5i2.369

Z. A. Dwiyanti and C. Prianto, “Prediksi Cuaca Kota Jakarta Menggunakan Metode Random Forest,” Jurnal Tekno Insentif, vol. 17, no. 2, pp. 127–137, Oct. 2023, doi: 10.36787/jti.v17i2.1136.

A. H. Bik, F. T. Anggraeny, and E. Y. Puspaningrum, “Klasifikasi Penyakit Ginjal Menggunakan Algoritma Hibrida CNN-ELM,” Jurnal Mahasiswa Teknik Informatika, vol. 8, no. 3, pp. 3836–3844, 2024, Accessed: Aug. 02, 2024. [Online]. Available: https://doi.org/10.36040/jati.v8i3.9807

D. T. Wilujeng, M. Fatekurohman, and I. M. Tirta, “Analisis Risiko Kredit Perbankan Menggunakan Algoritma K-Nearest Neighbor dan Nearest Weighted K-Nearest Neighbor,” Indonesian Journal of Applied Statistics, vol. 5, no. 2, pp. 142–148, Oct. 2023, doi: 10.13057/ijas.v5i2.58426.

R. N. Irawan, K. M. Hindrayani, and M. Idhom, “Penerapan Cross Validation sebagai Analisis Sentimen Pelayanan Publik Kereta Api Lokal Daop 8 Menggunakan Metode Multinomial Naïve Bayes,” G-Tech: Jurnal Teknologi Terapan, vol. 8, no. 2, pp. 954–963, Apr. 2024, doi: 10.33379/gtech.v8i2.4117.

H. Azis, Purnawansyah, F. Fattah, and I. P. Putri, “Performa Klasifikasi K-NN dan Cross Validation Pada Data Pasien Pengidap Penyakit Jantung,” ILKOM Jurnal Ilmiah, vol. 12, no. 2, pp. 81–86, Aug. 2020, doi: 10.33096/ilkom.v12i2.507.81-86.

H. Hafid, “Penerapan K-Fold Cross Validation untuk Menganalisis Kinerja Algoritma K-Nearest Neighbor pada Data Kasus Covid-19 di Indonesia,” Journal of Mathematics, Computations, and Statistics, vol. 6, no. 2, pp. 161–168, 2023, Accessed: Jul. 29, 2024. [Online]. Available: https://doi.org/10.35580/jmathcos.v6i2.53043

Downloads

Submitted

2025-06-03

Accepted

2026-04-27

Published

2026-08-23

How to Cite

[1]
T. Diba, E. Rustamaji, E. khudzaeva, N. Hasanati, and N. Kumaladewi, “Classification Analysis of Market Service Levy Payment Compliance in South Tangerang Using Random Forest Algorithm”, TEKNOSI, vol. 12, no. 2, pp. 190–202, Aug. 2026.

Similar Articles

1 2 3 4 5 6 7 8 > >> 

You may also start an advanced similarity search for this article.