High Confidence, Low Accuracy: Probing Ownership Bias in LLMs on Banjarese and Dayak Ngaju

Penulis

DOI:

https://doi.org/10.25077/TEKNOSI.v12i2.2026.386-392

Kata Kunci:

Confidence calibration, Ownership bias, Banjarese Dayak Ngaju language, Large language models

Abstrak

AI assistants often give wrong answers with high confidence, especially when answering in regional languages. One cause is ownership bias: a model trusts its own answer more when that answer appears in the assistant role. This effect is well documented in English, but it is unclear whether it also appears in low-resource languages such as Banjarese and Dayak Ngaju. This study measures the extent of ownership bias and the calibration quality of four open-weight 8-bit LLMs when tested on school exam questions in Banjarese and Dayak Ngaju, drawn from the IndoMMLU benchmark. Four models (Qwen3-8B, Llama-3.1-8B, SahabatAI-Llama-8B, and SahabatAI-Gemma2-9B) are evaluated under two conditions: Own, in which the model’s answer is placed in the assistant role, and Reframed, in which the same answer is attributed to the user. We report accuracy, ownership bias, Expected Calibration Error (ECE), and Brier Score. Llama-3.1-8B shows the highest ownership bias (49.93 pp on Banjarese and 53.83 pp on Dayak Ngaju), while SahabatAI-Gemma2-9B shows almost no bias (0.91–1.33 pp) but assigns 100% confidence to every answer. The same pattern appears in both languages, indicating that ownership bias is driven by the model's architecture and training, not by the language being tested. AI assistants that are overconfident in English remain overconfident in Banjarese and Dayak Ngaju as well. Practitioners should treat verbalized confidence as an unreliable signal and use model-level calibration such as temperature scaling or logprob-based methods, rather than relying on prompt-level reframing alone.

Referensi

J. Sanz-Guerrero and others, “Ownership Bias in Instruction-Tuned Large Language Models,” 2026.

F. Koto and others, “IndoMMLU: A Massive Multilingual Language Understanding Benchmark for Indonesian,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023.

C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in International Conference on Machine Learning, 2017, pp. 1321–1330.

Z. Liu and others, “UAIT: Uncertainty-Aware Instruction Tuning for Large Language Models,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024.

J. Ni, H. Zhao, Y. Yang, and D. Guo, “Deep neural network calibration by reducing classifier shift with stochastic masking,” Pattern Recognit., vol. 177, p. 113217, 2026.

S. Zhang and L. Xie, “Advancing neural network calibration: The role of gradient decay in large-margin Softmax optimization,” Neural Networks, vol. 178, p. 106457, 2024.

K. R. M. Fernando and C. P. Tsokos, “Dynamically Weighted Balanced Loss: Class Imbalanced Learning and Confidence Calibration of Deep Neural Networks,” IEEE Trans. Neural Netw. Learn. Syst., 2022.

S. A. Balanya, J. Maroñas, and D. Ramos, “Adaptive temperature scaling for robust calibration of deep neural networks,” Neural Comput. Appl., vol. 36, pp. 8073–8095, 2024.

Z. Lin, S. Trivedi, and J. Sun, “Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models,” Transactions on Machine Learning Research, 2024.

Y. Huang et al., “Look Before You Leap: An Exploratory Study of Uncertainty Analysis for Large Language Models,” IEEE Transactions on Software Engineering, 2024.

O. Shorinwa, Z. Mei, J. Lidard, A. Z. Ren, and A. Majumdar, “A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions,” ACM Comput. Surv., vol. 58, no. 3, p. 63, 2025.

R. Vashurin, E. Fadeeva, A. Vazhentsev, and others, “Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph,” in Transactions of the Association for Computational Linguistics, 2025.

R. Bentegeac, B. Le Guellec, G. Kuchcinski, and others, “Token Probabilities to Mitigate Large Language Models Overconfidence in Answering Medical Questions: Quantitative Study,” J. Med. Internet Res., 2025.

I. A. Qazi, Z. Khan, A. Ghani, and others, “Large language models show Dunning-Kruger-like effects in multilingual fact-checking,” Sci. Rep., 2026.

J. C. Bauer, S. Trattnig, F. Vieltorf, and R. Daub, “Handling data drift in deep learning-based quality monitoring: evaluating calibration methods using the example of friction stir welding,” J. Intell. Manuf., vol. 37, pp. 759–774, 2026.

K. Mohanarangan and P. Palanisamy, “Confidence-Aware Cascaded Learning for Real-Time Weakly Supervised Anomaly Detection in Surveillance Video,” IEEE Access, 2026.

M. Thelwall, “Evaluating research quality with Large Language Models: An analysis of ChatGPT’s effectiveness with different settings and inputs,” Journal of Data and Information Science, vol. 10, no. 1, pp. 7–25, 2025.

J. Q. J. Liu, K. T. K. Hui, and others, “The great detectives: humans versus AI detectors in catching large language model-generated medical writing,” International Journal for Educational Integrity, vol. 20, no. 8, 2024.

S. Ramamoorthy, V. Shah, S. Khanuja, Z. Sheikh, and others, “MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking,” in Transactions of the Association for Computational Linguistics, 2026.

H. Zhang, T. Feng, P. Han, and J. You, “AcademicEval: Live Long-Context LLM Benchmark,” Transactions on Machine Learning Research, 2025.

I. de Zarzà, M. Liz, J. de Curtò, and C. T. Calafate, “Energy-Aware Multilingual Evaluation of Large Language Models,” Energies (Basel)., 2026.

L. Qin, Q. Chen, Y. Zhou, Z. Chen, and others, “A survey of multilingual large language models,” Patterns, 2025.

Telah diserahkan

09-07-2026

Diterima

20-08-2026

Diterbitkan

16-09-2026

Cara Mengutip

[1]
M. Ihsan dan D. Andriawan, “High Confidence, Low Accuracy: Probing Ownership Bias in LLMs on Banjarese and Dayak Ngaju”, TEKNOSI, vol. 12, no. 2, hlm. 386–392, Sep 2026.

Terbitan

Bagian

Articles

Artikel Serupa

1 2 3 4 > >> 

Anda juga bisa Mulai pencarian similarity tingkat lanjut untuk artikel ini.