High Confidence, Low Accuracy: Probing Ownership Bias in LLMs on Banjarese and Dayak Ngaju

Penulis

DOI:

https://doi.org/10.25077/TEKNOSI.v12i2.2026.386-392

Kata Kunci:

Confidence calibration, Ownership bias, Banjarese Dayak Ngaju language, Large language models

Abstrak

AI assistants often give wrong answers with high confidence, especially when answering in regional languages. One cause is ownership bias: a model trusts its own answer more when that answer appears in the assistant role. This effect is well documented in English, but it is unclear whether it also appears in low-resource languages such as Banjarese and Dayak Ngaju. This study measures the extent of ownership bias and the calibration quality of four open-weight 8-bit LLMs when tested on school exam questions in Banjarese and Dayak Ngaju, drawn from the IndoMMLU benchmark. Four models (Qwen3-8B, Llama-3.1-8B, SahabatAI-Llama-8B, and SahabatAI-Gemma2-9B) are evaluated under two conditions: Own, in which the model’s answer is placed in the assistant role, and Reframed, in which the same answer is attributed to the user. We report accuracy, ownership bias, Expected Calibration Error (ECE), and Brier Score. Llama-3.1-8B shows the highest ownership bias (49.93 pp on Banjarese and 53.83 pp on Dayak Ngaju), while SahabatAI-Gemma2-9B shows almost no bias (0.91–1.33 pp) but assigns 100% confidence to every answer. The same pattern appears in both languages, indicating that ownership bias is driven by the model's architecture and training, not by the language being tested. AI assistants that are overconfident in English remain overconfident in Banjarese and Dayak Ngaju as well. Practitioners should treat verbalized confidence as an unreliable signal and use model-level calibration such as temperature scaling or logprob-based methods, rather than relying on prompt-level reframing alone.

Referensi

M. Sanz-Guerrero, M. Mager, and K. von der Wense, “Large Language Models Are Overconfident in Their Own Responses,” in Findings of the Association for Computational Linguistics: ACL 2026, 2026, pp. 31406–31418, doi: 10.18653/v1/2026.findings-acl.1570.

F. Koto, N. Aisyah, H. Li, and T. Baldwin, “Large Language Models Only Pass Primary School Exams in Indonesia: A Comprehensive Test on IndoMMLU,” in Proc. 2023 Conf. Empirical Methods Natural Language Processing (EMNLP), 2023, pp. 12359–12374, doi: 10.18653/v1/2023.emnlp-main.760.

C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in Proc. 34th Int. Conf. Machine Learning (ICML), vol. 70, 2017, pp. 1321–1330. [Online]. Available: https://proceedings.mlr.press/v70/guo17a.html

S. Liu et al., “Can LLMs Learn Uncertainty on Their Own? Expressing Uncertainty Effectively in a Self-Training Manner,” in Proc. 2024 Conf. Empirical Methods Natural Language Processing (EMNLP), 2024, pp. 21635–21645, doi: 10.18653/v1/2024.emnlp-main.1205.

J. Ni, H. Zhao, Y. Yang, and D. Guo, “Deep Neural Network Calibration by Reducing Classifier Shift with Stochastic Masking,” Pattern Recognit., vol. 177, Art. no. 113217, 2026, doi: 10.1016/j.patcog.2026.113217.

S. Zhang and L. Xie, “Advancing Neural Network Calibration: The Role of Gradient Decay in Large-Margin Softmax Optimization,” Neural Netw., vol. 178, Art. no. 106457, 2024, doi: 10.1016/j.neunet.2024.106457.

K. R. M. Fernando and C. P. Tsokos, “Dynamically Weighted Balanced Loss: Class Imbalanced Learning and Confidence Calibration of Deep Neural Networks,” IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 7, pp. 2940–2951, 2022, doi: 10.1109/TNNLS.2020.3047335.

S. A. Balanya, J. Maroñas, and D. Ramos, “Adaptive Temperature Scaling for Robust Calibration of Deep Neural Networks,” Neural Comput. Appl., vol. 36, pp. 8073–8095, 2024, doi: 10.1007/s00521-024-09505-4.

Z. Lin, S. Trivedi, and J. Sun, “Generating with Confidence: Uncertainty Quantification for Black-Box Large Language Models,” Trans. Mach. Learn. Res., 2024. [Online]. Available: https://openreview.net/forum?id=DWkJCSxKU5

Y. Huang et al., “Look Before You Leap: An Exploratory Study of Uncertainty Analysis for Large Language Models,” IEEE Trans. Softw. Eng., vol. 51, no. 2, pp. 413–429, 2025, doi: 10.1109/TSE.2024.3519464.

O. Shorinwa, Z. Mei, J. Lidard, A. Z. Ren, and A. Majumdar, “A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions,” ACM Comput. Surv., vol. 58, no. 3, Art. no. 63, pp. 1–38, 2026, doi: 10.1145/3744238.

R. Vashurin et al., “Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph,” Trans. Assoc. Comput. Linguist., vol. 13, pp. 220–248, 2025, doi: 10.1162/tacl_a_00737.

R. Bentegeac, B. Le Guellec, G. Kuchcinski, P. Amouyel, and A. Hamroun, “Token Probabilities to Mitigate Large Language Models Overconfidence in Answering Medical Questions: Quantitative Study,” J. Med. Internet Res., vol. 27, Art. no. e64348, 2025, doi: 10.2196/64348.

I. A. Qazi et al., “Large Language Models Show Dunning–Kruger-Like Effects in Multilingual Fact-Checking,” Sci. Rep., vol. 16, Art. no. 7594, 2026, doi: 10.1038/s41598-026-39046-w.

J. C. Bauer, S. Trattnig, F. Vieltorf, and R. Daub, “Handling Data Drift in Deep Learning-Based Quality Monitoring: Evaluating Calibration Methods Using the Example of Friction Stir Welding,” J. Intell. Manuf., vol. 37, pp. 759–774, 2026, doi: 10.1007/s10845-025-02569-6.

K. Mohanarangan and P. Palanisamy, “Confidence-Aware Cascaded Learning for Real-Time Weakly Supervised Anomaly Detection in Surveillance Video,” IEEE Access, vol. 14, pp. 56388–56411, 2026, doi: 10.1109/ACCESS.2026.3679742.

M. Thelwall, “Evaluating Research Quality with Large Language Models: An Analysis of ChatGPT’s Effectiveness with Different Settings and Inputs,” J. Data Inf. Sci., vol. 10, no. 1, pp. 7–25, 2025, doi: 10.2478/jdis-2025-0011.

J. Q. J. Liu et al., “The Great Detectives: Humans Versus AI Detectors in Catching Large Language Model-Generated Medical Writing,” Int. J. Educ. Integr., vol. 20, Art. no. 8, 2024, doi: 10.1007/s40979-024-00155-6.

S. Ramamoorthy et al., “MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking,” Trans. Assoc. Comput. Linguist., vol. 14, pp. 399–417, 2026, doi: 10.1162/tacl.a.633.

H. Zhang, T. Feng, P. Han, and J. You, “AcademicEval: Live Long-Context LLM Benchmark,” Trans. Mach. Learn. Res., 2025. [Online]. Available: https://openreview.net/forum?id=LjQ4voE5bs

I. de Zarzà, M. Liz, J. de Curtò, and C. T. Calafate, “Energy-Aware Multilingual Evaluation of Large Language Models,” Electronics, vol. 15, no. 7, Art. no. 1395, 2026, doi: 10.3390/electronics15071395.

L. Qin et al., “A Survey of Multilingual Large Language Models,” Patterns, vol. 6, no. 1, Art. no. 101118, 2025, doi: 10.1016/j.patter.2024.101118.

Telah diserahkan

09-07-2026

Diterima

20-08-2026

Diterbitkan

16-09-2026

Cara Mengutip

[1]
M. Ihsan dan D. Andriawan, “High Confidence, Low Accuracy: Probing Ownership Bias in LLMs on Banjarese and Dayak Ngaju”, TEKNOSI, vol. 12, no. 2, hlm. 386–392, Sep 2026.

Terbitan

Bagian

Articles

Artikel Serupa

1 2 3 4 > >> 

Anda juga bisa Mulai pencarian similarity tingkat lanjut untuk artikel ini.