Comparative Analysis of XGBoost and CatBoost for Detecting AI-Generated Synthetic Faces Using Texture and Color Features
DOI:
https://doi.org/10.61255/decoding.v4i3.1677Keywords:
CatBoost, GLCM, Image classification, XGBoost, YCbCrAbstract
Purpose – The rapid advancement of Generative AI produces near-perfect synthetic face images, posing fraud risks. This study compares XGBoost and CatBoost for detecting real versus synthetic faces using texture and color features.
Methods – Texture features were extracted using GLCM with four orientations (0°, 45°, 90°, 135°), generating 16 features, while YCbCr mean and standard deviation provided 6 color features (22 total). A balanced 750-image dataset (real and AI from GPT Image 2, Nano Banana, Leonardo AI, Canva, Dreamina AI) was split 80:20 and normalized with StandardScaler.
Findings – On an asymmetric, source-shifted test split, where real test images were captured independently while synthetic test images came from the same generators as training, CatBoost outperformed XGBoost, achieving 84.67% accuracy (vs. 80.67%), 83.33% precision (vs. 78.05%), 86.67% recall (vs. 85.33%), and 84.97% F1-score (vs. 81.53%), reducing False Positives by 27.78% and False Negatives by 9.09%. However, repeated 5-fold stratified cross-validation (5×10 repeats) revealed only marginal, non-significant differences (0.05-0.87 percentage points, p > 0.09).
Research implications – Although the moderate dataset excludes recent architectures such as StyleGAN, findings affirm the feasibility of efficient, interpretable detection through feature engineering. Since within-domain cross-validation shows no significant algorithmic difference, the observed cross-domain gap should be regarded as preliminary, pending validation through repeated cross-domain evaluations across diverse AI platforms.
Originality – The novelty lies in comparing XGBoost and CatBoost on multi-orientation GLCM and YCbCr feature fusion for contemporary AI-generated images, identifying cr_std and contrast_0 as stable, interpretable discriminative markers.
Abstract views: 9
,
PDF downloads: 14
Downloads
References
E. J. Miller, B. A. Steward, Z. Witkower, C. A. M. Sutherland, E. G. Krumhuber, and A. Dawel, “AI Hyperrealism : Why AI Faces Are Perceived as More Real Than Human Ones,” Psychol. Sci., vol. 34, no. 12, pp. 1390–1403, 2023, doi: 10.1177/09567976231207095.
A. Hardiyanto, “Classification of Synthetic Images Generated by Diffusion Models Using Gray Level Co-Occurrence Matrix (Glcm) and Convolutional Neural Network (CNN),” JIKO (Jurnal Inform. dan Komputer), vol. 9, no. 3, p. 572, 2025, doi: 10.26798/jiko.v9i3.2033.
H. Prayoga and H. Tuasikal, “Penyebaran Konten Deepfake Sebagai Tindak Pidana: Analisis Kritis Terhadap Penegakan Hukum Dan Perlindungan Publik Di Indonesia [Deepfake Content Distribution as a Criminal Act: Critical Analysis of Law Enforcement and Public Protection in Indonesia],” Abdurrauf Law Sharia, vol. 2, no. 2, pp. 22–38, 2024, doi: 10.70742/arlash.v2i1.194.
D. Velásquez-salamanca, M. Á. Martín-pascual, and C. Andreu-sánchez, “Interpretation of AI-Generated vs . Human-Made Images,” J. Imaging, vol. 11, pp. 1–16, 2025, doi: 10.3390/jimaging11070227.
V. Osińska, W. Kortas, A. Szalach, and M. Welter, “AI Images vs . Real Photographs : Investigating Visual Recognition and Perception,” J. Eye Mov. Res., vol. 18, p. 61, 2025, doi: 10.3390/jemr18060061.
J. Mu, M. Adrezo, and A. N. Haikal, “Identifikasi Wajah Asli dan Buatan Deepfake Menggunakan Metode Convolutional Neural Network [Identification of Real and Deepfake Faces Using Convolutional Neural Network Method],” TEKNIKA, vol. 13, no. 1, pp. 45–50, 2024, doi: 10.34148/teknika.v13i1.705.
M. S. A. Aria, C. Slamet, and M. D. Firdaus, “Klasifikasi Fake dan Real Menggunakan Vision Transformer dan EfficientNet-B0 pada Gambar Asli dan Generatif AI [Fake and Real Classification Using Vision Transformer and EfficientNet-B0 on Real and Generative AI Images],” SMATIKA J., vol. 15, no. 1, pp. 179–192, 2025, doi: 10.32664/smatika.v15i01.1531.
A. Gulati, A. Felahatpisheh, and C. E. Valderrama, “Feature engineering through two-level genetic algorithm,” Mach. Learn. with Appl., vol. 21, no. July, p. 100696, 2025, doi: 10.1016/j.mlwa.2025.100696.
T. Qamar and N. Z. Bawany, “Understanding the black-box : towards interpretable and reliable deep learning models,” PeerJ Comput. Sci., vol. 9, p. e1629, 2023, doi: 10.7717/peerj-cs.1629.
J. Huang and T. Handhayani, “Klasifikasi Telapak Tangan Menggunakan Deep Learning Dan Machine Learning [Palm Classification Using Deep Learning and Machine Learning],” J. Mnemon., vol. 8, no. 2, pp. 293–300, 2025, doi: 10.36040/mnemonic.v8i2.15716.
R. F. Naibaho and I. P. Sari, “Implementasi Metode Gray Level Co-Occurrence Matrix Menganalisis Tekstur Kulit Wajah [Implementation of Gray Level Co-Occurrence Matrix Method for Analyzing Facial Skin Texture],” sudo J. Tek. Inform., vol. 3, no. 4, pp. 172–182, 2025, doi: 10.56211/sudo.v3i4.668.
C. I. Purba, A. Alrizal, and Y. Fendriani, “Klasifikasi Kanker Kulit dari Citra Dermoskopi Menggunakan Fitur Gray Level Co-occurrence Matrix (GLCM) dengan Algoritma Machine Learning [Skin Cancer Classification from Dermoscopy Images Using Gray Level Co-occurrence Matrix (GLCM) Features with Machine,” JFT J. Fis. dan Ter., vol. 12, pp. 30–44, 2025, doi: 10.24252/jft.v12i1.56651.
A. Prasetio and Y. Utami, “Analisis Warna Kulit Menggunakan Citra Digital Berbasis Ruang Warna YCbCr [Skin Color Analysis Using Digital Images Based on YCbCr Color Space],” J. Nas. Komputasi dan Teknol. Inf., vol. 8, no. 3, pp. 1728–1732, 2025, doi: 10.32672/jnkti.v8i3.9228.
S. Alharthi and A. Gutub, “Adjusting image stego practicality via YCbCr color space formation,” J. Eng. Res., vol. 14, no. 1, pp. 756–764, 2026, doi: 10.1016/j.jer.2025.07.008.
K. Letou and S. Thiyagarajan, “Analysis of Enhanced Adaptive Gradient Boosting Regression in Business Intelligence,” Int. J. Adv. SIGNAL IMAGE Sci., vol. 12, no. 2, pp. 1075–1088, 2026, doi: 10.29284/1f9fcb03.
T. Chen and C. Guestrin, “XGBoost : A Scalable Tree Boosting System,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, pp. 785–794. doi: 10.1145/2939672.2939785.
F. I. Adani and H. Amalia, “Penerapan Metode Algoritma XGBoost untuk Prediksi Risiko Penyakit Jantung [Application of the XGBoost Algorithm Method for Heart Disease Risk Prediction],” J. Insa. J. Inf. Syst. Manag. Innov., vol. 5, no. 2, pp. 117–125, 2025, doi: 10.31294/j-insan.v5i2.10000.
L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin, “CatBoost : unbiased boosting with categorical features,” in Advances in Neural Information Processing Systems (NeurIPS 2018), 2018, pp. 6638–6648.
G. E. Saputra, M. Hanindia, P. Swari, and A. L. Nurlaili, “Implementasi Algoritma XGBoost, CatBoost, dan LGBM untuk Klasifikasi Pencemaran Udara [Implementation of XGBoost, CatBoost, and LGBM Algorithms for Air Pollution Classification],” JIIP - J. Ilm. Ilmu Pendidik., vol. 8, no. 12, pp. 14135–14139, 2025, doi: 10.54371/jiip.v8i12.10102.
I. Setiawan, R. Dahlan, A. I. Basuki, H. Susanto, and D. Rosiyadi, “Trade-off between Image Quality and Computational Complexity : Image Resizing Perspective,” J. Tek. Elektro, vol. 14, no. 1, pp. 24–28, 2022, doi: 10.15294/jte.v14i1.37629.
A. Setiawan and A. A. Budiman, “Optimization of Gray Level Co-occurrence Matrix ( GLCM ) Texture Feature Parameters in Determining Rice Seed Quality,” Emit. Int. J. Eng. Technol., vol. 13, no. 1, pp. 110–123, 2025, doi: 10.24003/emitter.v13i1.928.
S. K. Abdulateef and A. N. Hasoon, “Comparison of the Components of Different Colour Spaces to Enhanced Image Representation,” J. Image Process. Intell. Remote Sens., vol. 03, no. 01, pp. 11–17, 2023, doi: 10.55529/jipirs.31.11.17.
R. R. Laska and A. M. Yolanda, “A Comparative Study of Z-Score and Min-Max Normalization for Rainfall Classification in Pekanbaru,” J. Data Sci., vol. 2024, no. 1, pp. 1–8, 2024, doi: 10.61453/jods.v2024no04.
K. M. Sujon, R. Hassan, K. Choi, and M. A. Samad, “Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models,” J. Big Data, vol. 12, no. 1, p. 268, 2025, doi: 10.1186/s40537-025-01313-4.
A. A. Huang and S. Y. Huang, “Use of machine learning to identify risk factors for insomnia,” PLoS One, vol. 18, no. 4, pp. 1–16, 2023, doi: 10.1371/journal.pone.0282622.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 indra firmansyah, Defry Hamdhana , Wahyu Fuadi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
