Tinjauan Literatur Sistematis: Teknik Optical Character Recognition Dan In-formation Retrieval Pada Sistem Tanya Jawab

Authors

  • Sekar Arum Chrysantia Program Studi D4 Teknik Informatika, Universitas Logistik dan Bisnis Internasional
  • Roni Habibi Universitas Logistik dan Bisnis Internasional

DOI:

https://doi.org/10.25077/TEKNOSI.v12i2.2026.284-299

Keywords:

Sistem Tanya Jawab Berbasis Gambar, Optical Character Recognition, Information Retrieval

Abstract

Berikut terjemahan yang telah disesuaikan dengan bahasa akademik yang lazim digunakan dalam jurnal internasional (Scopus/Q1): The development of Image-based Question Answering (IQA) systems has advanced rapidly alongside the evolution of Optical Character Recognition (OCR), Information Retrieval (IR), and multimodal artificial intelligence technologies. However, the integration of OCR, retrieval mechanisms, and multimodal reasoning continues to face various conceptual and methodological challenges. Furthermore, comprehensive studies synthesizing the interrelationships among these components remain limited. This study aims to map the development, research trends, challenges, and future directions of OCR- and Information Retrieval-based techniques in Image-based Question Answering systems. A Systematic Literature Review (SLR) was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A total of 61 indexed articles were analyzed using the Theory, Context, Characteristics, and Methodology (TCCM) framework. The findings indicate that Multimodal Deep Learning, Transformer architectures, and Attention Mechanisms constitute the primary foundations of modern Question Answering systems. In addition, the integration of Knowledge Graphs, Large Language Models (LLMs), and Retrieval-Augmented Generation (RAG) has become increasingly prominent due to their ability to enhance reasoning capabilities and contextual understanding. The review also reveals a paradigm shift from traditional feature engineering–based approaches toward knowledge-driven multimodal reasoning systems. These findings provide a foundation for the development of more adaptive, reliable, and context-aware image-based question answering systems. Despite the rapid progress of IQA research, several key challenges remain, including the semantic gap, the complexity of multimodal fusion, hallucination issues, limited availability of multilingual datasets, and the lack of system interpretability.  

References

ased framework for chest X–ray follow–up medical visual question answering,” Biomed. Signal Process. Control, vol. 120, p. 109908, 2026, doi: 10.1016/j.bspc.2026.109908.

A. Salaberria, G. Azkune, O. Lopez De Lacalle, A. Soroa, and E. Agirre, “Image captioning for effective use of language models in knowledge-based visual question answering,” Expert Syst. Appl., vol. 212, p. 118669, 2023, doi: 10.1016/j.eswa.2022.118669.

L. Jiang and Z. Meng, “Knowledge-Based Visual Question Answering Using Multi-Modal Semantic Graph,” Electronics (Basel)., vol. 12, no. 6, p. 1390, 2023, doi: 10.3390/electronics12061390.

I. Kim, “Visual Experience-Based Question Answering with Complex Multimodal Environments,” Math. Probl. Eng., vol. 2020, pp. 1–18, 2020, doi: 10.1155/2020/8567271.

C. Chen, D. Han, and J. Wang, “Multimodal Encoder-Decoder Attention Networks for Visual Question Answering,” IEEE Access, vol. 8, pp. 35662–35671, 2020, doi: 10.1109/ACCESS.2020.2975093.

S. Zhang, Y. Zhang, Z. Chen, and Z. Li, “VSAM-Based Visual Keyword Generation for Image Caption,” IEEE Access, vol. 9, pp. 27638–27649, 2021, doi: 10.1109/ACCESS.2021.3058425.

P. Gao, H. Sun, G. Chen, R. Wang, and M. Li, “Visual Question Answering for Intelligent Interaction,” Mobile Information Systems, vol. 2022, pp. 1–6, 2022, doi: 10.1155/2022/4232968.

Q. Li, X. Tang, and Y. Jian, “Learning to Reason on Tree Structures for Knowledge-Based Visual Question Answering,” Sensors, vol. 22, no. 4, p. 1575, 2022, doi: 10.3390/s22041575.

R. Wang, S. Wu, and X. Wang, “The Core of Smart Cities: Knowledge Representation and Descriptive Framework Construction in Knowledge-Based Visual Question Answering,” Sustainability, vol. 14, no. 20, p. 13236, 2022, doi: 10.3390/su142013236.

R. Wang et al., “FTN–VQA: MULTIMODAL REASONING BY LEVERAGING A FULLY TRANSFORMER-BASED NETWORK FOR VISUAL QUESTION ANSWERING,” Fractals, vol. 31, no. 06, p. 2340133, 2023, doi: 10.1142/S0218348X23401333.

H. Wang, J. Li, and J. Wang, “Retrieving Chinese Questions and Answers Based on Deep-Learning Algorithm,” Mathematics, vol. 11, no. 18, pp. 1–18, 2023, doi: 10.3390/math11183843.

H. Zhu, R. Togo, T. Ogawa, and M. Haseyama, “Multimodal Natural Language Explanation Generation for Visual Question Answering Based on Multiple Reference Data,” Electronics (Basel)., vol. 12, no. 10, p. 2183, 2023, doi: 10.3390/electronics12102183.

J. He et al., “PERS: Parameter-Efficient Multimodal Transfer Learning for Remote Sensing Visual Question Answering,” IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., vol. 17, pp. 14823–14835, 2024, doi: 10.1109/JSTARS.2024.3447086.

T. Yamane, P. Chun, J. Dang, and T. Okatani, “Deep learning-based bridge damage cause estimation from multiple images using visual question answering,” Structure and Infrastructure Engineering, vol. 22, no. 4, pp. 734–747, 2026, doi: 10.1080/15732479.2024.2355929.

C. Song, “Enhancing Multimodal Understanding With LIUS: A Novel Framework for Visual Question Answering in Digital Marketing,” Journal of Organizational and End User Computing, vol. 36, no. 1, pp. 1–17, 2024, doi: 10.4018/JOEUC.336276.

Y. Cong and H. Mo, “A visual question answering method based on task decomposition,” PLoS One, vol. 20, no. 11, p. e0336623, 2025, doi: 10.1371/journal.pone.0336623.

Z. Liu, K. Yang, Y. Weng, Z. He, X. Liu, and H. Gao, “SCAG: Semantic Co-occurring Attention Guided Alignment for Knowledge-based Visual Question Answering,” ACM Transactions on Multimedia Computing, Communications, and Applications, vol. 21, no. 7, pp. 1–20, 2025, doi: 10.1145/3734220.

C. Huang and Z. Hu, “A multimodal transformer-based visual question answering method integrating local and global information,” PLoS One, vol. 20, no. 7, p. e0324757, 2025, doi: 10.1371/journal.pone.0324757.

M. S. Islam et al., “MRAN-VQA: Multimodal Recursive Attention Network for Visual Question Answering,” Engineering Science and Technology, an International Journal, vol. 72, p. 102232, 2025, doi: 10.1016/j.jestch.2025.102232.

C. Gupta, N. S. Gill, P. Gulia, and G. Pau, “CODNet: Context-based object detection network for multimodal image captioning and virtual question answering,” Image Vis. Comput., vol. 163, p. 105768, 2025, doi: 10.1016/j.imavis.2025.105768.

J. Xiao and Z. Zhang, “EduVQA: A multimodal Visual Question Answering framework for smart education,” Alexandria Engineering Journal, vol. 122, pp. 615–624, 2025, doi: 10.1016/j.aej.2025.03.005.

N. Hu, X. Zhang, Q. Zhang, W. Huo, and S. You, “ZPVQA: Visual Question Answering of Images Based on Zero-Shot Prompt Learning,” IEEE Access, vol. 13, pp. 50849–50859, 2025, doi: 10.1109/ACCESS.2025.3550942.

J. Ma et al., “DeepSORT-OCR: Design and Application Research of a Maritime Ship Target Tracking Algorithm Incorporating Hull Number Features,” Mathematics, vol. 14, no. 6, p. 1062, 2026, doi: 10.3390/math14061062.

H. Mohammadi and S. H. Khasteh, “A fast text similarity measure for large document collections using multireferencecosine and genetic algorithm,” TURKISH JOURNAL OF ELECTRICAL ENGINEERING & COMPUTER SCIENCES, vol. 28, no. 2, pp. 999–1013, 2020, doi: 10.3906/elk-1906-30.

W. Budiharto, V. Andreas, and A. A. S. Gunawan, “Deep learning-based question answering system for intelligent humanoid robot,” J. Big Data, vol. 7, no. 1, 2020, doi: 10.1186/s40537-020-00341-6.

L.-Q. Cai, M. Wei, S.-T. Zhou, and X. Yan, “Intelligent Question Answering in Restricted Domains Using Deep Learning and Question Pair Matching,” IEEE Access, vol. 8, pp. 32922–32934, 2020, doi: 10.1109/ACCESS.2020.2973728.

Y. Wen, X. Zhu, and L. Zhang, “CQACD: A Concept Question-Answering System for Intelligent Tutoring Using a Domain Ontology With Rich Semantics,” IEEE Access, vol. 10, pp. 67247–67261, 2022, doi: 10.1109/ACCESS.2022.3185400.

X. Yang, J. Yang, R. Li, H. Li, H. Zhang, and Y. Zhang, “Complex Knowledge Base Question Answering for Intelligent Bridge Management Based on Multi-Task Learning and Cross-Task Constraints,” Entropy, vol. 24, no. 12, p. 1805, 2022, doi: 10.3390/e24121805.

Y. Fang, J. Deng, F. Zhang, and H. Wang, “An Intelligent Question-Answering Model over Educational Knowledge Graph for Sustainable Urban Living,” Sustainability, vol. 15, no. 2, p. 1139, 2023, doi: 10.3390/su15021139.

M. Gao, M. Li, T. Ji, N. Wang, G. Lin, and Q. Wu, “Key Technologies of Intelligent Question-Answering System for Power System Rules and Regulations Based on Improved BERTserini Algorithm,” Processes, vol. 12, no. 1, p. 58, 2023, doi: 10.3390/pr12010058.

S. Zhang and Q. Jaamour, “English Intelligent Question Answering System Based on elliptic fitting equation,” Applied Mathematics and Nonlinear Sciences, vol. 8, no. 1, pp. 1743–1752, 2023, doi: 10.2478/amns.2022.2.0162.

R. Li, G. Ren, J. Yan, B. Zou, and Q. Liu, “Intelligent question answering system for traditional Chinese medicine based on BSG deep learning model: taking prescription and Chinese materia medica as examples,” Digital Chinese Medicine, vol. 7, no. 1, pp. 47–55, Mar. 2024, doi: 10.1016/j.dcmed.2024.04.006.

P. Lyu, J. Fu, C. Liu, W. Yu, and L. Xia, “GPB and BAC: two novel models towards building an intelligent motor fault maintenance question answering system,” Journal of Engineering Design, vol. 36, no. 11, pp. 2007–2027, 2025, doi: 10.1080/09544828.2024.2335135.

A. Mohamed, K. Abdelqader, and K. Shaalan, “Machine learning and deep learning techniques in Arabic question answering systems: innovations and challenges,” PeerJ Comput. Sci., vol. 11, pp. 1–35, 2025, doi: 10.7717/peerj-cs.3331.

E. Bernasconi, D. Redavid, and S. Ferilli, “Enhancing Personalised Learning with a Context-Aware Intelligent Question-Answering System and Automated Frequently Asked Question Generation,” Electronics (Switzerland), vol. 14, no. 7, pp. 1–26, 2025, doi: 10.3390/electronics14071481.

B. Zhang, X. Zhang, Q. Wang, G. Gui, and L. Shan, “Intelligent Question Answering System Design with Domain-Specific Knowledge Graphs,” IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences, vol. E108.A, no. 3, pp. 546–554, 2025, doi: 10.1587/transfun.2024EAP1090.

D. Sharma, S. Purushotham, and C. K. Reddy, “MedFuseNet: An attention-based multimodal deep learning model for visual question answering in the medical domain,” Sci. Rep., vol. 11, no. 1, p. 19826, 2021, doi: 10.1038/s41598-021-98390-1.

J. D. Silva, B. Martins, and J. Magalhães, “Contrastive training of a multimodal encoder for medical visual question answering,” Intelligent Systems with Applications, vol. 18, p. 200221, 2023, doi: 10.1016/j.iswa.2023.200221.

J. Holland, A. McGarvey, M. Flood, P. Joyce, and T. Pawlikowska, “A Qualitative Exploration of Student Cognition When Answering Text-Only or Image-Based Histology Multiple-Choice Questions,” Med. Sci. Educ., vol. 34, no. 6, pp. 1317–1329, 2024, doi: 10.1007/s40670-024-02104-x.

L. Chen, “Medical education and artificial intelligence: Question answering for medical questions based on intelligent interaction,” Concurr. Comput., vol. 36, no. 14, p. e8079, 2024, doi: 10.1002/cpe.8079.

B. Lei and P. Yin, “Technological Evolution and Research Trends of Intelligent Question-Answering Systems in Healthcare,” Healthcare, vol. 13, no. 18, p. 2269, 2025, doi: 10.3390/healthcare13182269.

F. Lu, S. Liu, W. Lu, P. Chen, and B. Ding, “Medical Knowledge-Based Differential Image Visual Question Answering,” IEEE Access, vol. 13, pp. 93818–93829, 2025, doi: 10.1109/ACCESS.2025.3565695.

T. Seki, Y. Kawazoe, H. Ito, Y. Akagi, T. Takiguchi, and K. Ohe, “Assessing the performance of zero-shot visual question answering in multimodal large language models for 12-lead ECG image interpretation,” Front. Cardiovasc. Med., vol. 12, p. 1458289, 2025, doi: 10.3389/fcvm.2025.1458289.[61] Y. Lu et al., “Application of Multimodal Transformer Model in Intelligent Agricultural Disease Detection and Question-Answering Systems,” Plants, vol. 13, no. 7, p. 972, 2024, doi: 10.3390/plants13070972.

W. Bi, Q. Xiong, X. Chen, Q. Du, J. Wu, and Z. Zhuang, “Intelligent visual question answering in TCM education: An innovative application of IoT and multimodal fusion,” Alexandria Engineering Journal, vol. 118, no. December 2024, pp. 325–336, 2025, doi: 10.1016/j.aej.2024.12.052.

S. Devmane, O. Rana, and C. Perera, “OntoSage: Intelligent Human-Building Smartbot for Semantic Smart Building Question Answering,” World Wide Web, vol. 29, no. 2, p. 17, 2026, doi: 10.1007/s11280-026-01403-0.

Y. Huang, H. Liu, S. Li, W. Wang, and Z. Zhou, “Effective Prediction and Important Counseling Experience for Perceived Helpfulness of Social Question and Answering-Based Online Counseling: An Explainable Machine Learning Model,” Front. Public Health, vol. 10, no. December, pp. 1–19, 2022, doi: 10.3389/fpubh.2022.817570.

Y. Lan, Y. Guo, Q. Chen, S. Lin, Y. Chen, and X. Deng, “Visual question answering model for fruit tree disease decision-making based on multimodal deep learning,” Front. Plant Sci., vol. 13, p. 1064399, 2023, doi: 10.3389/fpls.2022.1064399.

S. H. Kim, J. S. Shin, H. Lee, S. Y. Shin, K. M. Kang, and S. Y. Song, “A Comparative Analysis of GPT-3.5, GPT-4, GPT–4 Omni, Gemini Advanced, and Gemini 1.5 in Answering Frequently Asked Questions Regarding High Tibial Osteotomy,” Orthop. J. Sports Med., vol. 13, no. 11, pp. 1–13, 2025, doi: 10.1177/23259671251385127.

R. Zhu, B. Liu, R. Zhang, S. Zhang, and J. Cao, “OEQA: Knowledge- and Intention-Driven Intelligent Ocean Engineering Question-Answering Framework,” Applied Sciences, vol. 13, no. 23, p. 12915, 2023, doi: 10.3390/app132312915.

Submitted

2026-06-10

Accepted

2026-08-20

Published

2026-08-28

How to Cite

[1]
S. Arum Chrysantia and R. Habibi, “Tinjauan Literatur Sistematis: Teknik Optical Character Recognition Dan In-formation Retrieval Pada Sistem Tanya Jawab”, TEKNOSI, vol. 12, no. 2, pp. 284–299, Aug. 2026.

Similar Articles

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 > >> 

You may also start an advanced similarity search for this article.