Tinjauan Literatur Sistematis: Teknik Optical Character Recognition Dan In-formation Retrieval Pada Sistem Tanya Jawab

Penulis

  • Sekar Arum Chrysantia Program Studi D4 Teknik Informatika, Universitas Logistik dan Bisnis Internasional
  • Roni Habibi bProgram Studi D4 Teknik Informatika, Universitas Logistik dan Bisnis Internasional

DOI:

https://doi.org/10.25077/TEKNOSI.v12i2.2026.284-299

Kata Kunci:

Sistem Tanya Jawab Berbasis Gambar, Optical Character Recognition, Information Retrieval

Abstrak

Perkembangan Sistem Tanya Jawab Berbasis Gambar (Image-based Question Answering) mengalami peningkatan pesat seiring kemajuan teknologi Optical Character Recognition (OCR), information retrieval, dan kecerdasan buatan multimodal. Namun, integrasi OCR, mekanisme retrieval, dan multimodal reasoning masih menghadapi berbagai tantangan konseptual dan metodologis, sementara sintesis penelitian yang mengkaji keterkaitan ketiganya masih terbatas. Penelitian ini bertujuan memetakan perkembangan, tren, tantangan, dan arah penelitian masa depan terkait teknik berbasis OCR dan information retrieval pada sistem Image-based Question Answering. Metode yang digunakan adalah Systematic Literature Review (SLR) dengan mengacu pada pedoman PRISMA. Sebanyak 61 artikel terindeks dianalisis menggunakan kerangka TCCM (Theory, Context, Characteristics, and Methodology). Hasil penelitian menunjukkan bahwa Multimodal Deep Learning, arsitektur Transformer, dan Attention Mechanism menjadi fondasi utama pengembangan sistem QA modern. Selain itu, integrasi Knowledge Graph, Large Language Model (LLM), dan Retrieval-Augmented Generation (RAG) semakin dominan karena mampu meningkatkan kemampuan reasoning dan pemahaman kontekstual. Terdapat temuan yang mengindikasikan pergeseran paradigma dari pendekatan berbasis feature engineering menuju sistem knowledge-driven multimodal reasoning. Temuan ini memberikan dasar bagi pengembangan sistem QA berbasis gambar yang lebih adaptif, andal, dan kontekstual. Meskipun telah berkembang dengan pesat dan cepat, namun pengembangan IQA juga memiliki tantangan utama meliputi semantic gap, kompleksitas multimodal fusion, hallucination, keterbatasan dataset multibahasa, dan rendahnya interpretabilitas sistem.

Referensi

ased framework for chest X–ray follow–up medical visual question answering,” Biomed. Signal Process. Control, vol. 120, p. 109908, 2026, doi: 10.1016/j.bspc.2026.109908.

A. Salaberria, G. Azkune, O. Lopez De Lacalle, A. Soroa, and E. Agirre, “Image captioning for effective use of language models in knowledge-based visual question answering,” Expert Syst. Appl., vol. 212, p. 118669, 2023, doi: 10.1016/j.eswa.2022.118669.

L. Jiang and Z. Meng, “Knowledge-Based Visual Question Answering Using Multi-Modal Semantic Graph,” Electronics (Basel)., vol. 12, no. 6, p. 1390, 2023, doi: 10.3390/electronics12061390.

I. Kim, “Visual Experience-Based Question Answering with Complex Multimodal Environments,” Math. Probl. Eng., vol. 2020, pp. 1–18, 2020, doi: 10.1155/2020/8567271.

C. Chen, D. Han, and J. Wang, “Multimodal Encoder-Decoder Attention Networks for Visual Question Answering,” IEEE Access, vol. 8, pp. 35662–35671, 2020, doi: 10.1109/ACCESS.2020.2975093.

S. Zhang, Y. Zhang, Z. Chen, and Z. Li, “VSAM-Based Visual Keyword Generation for Image Caption,” IEEE Access, vol. 9, pp. 27638–27649, 2021, doi: 10.1109/ACCESS.2021.3058425.

P. Gao, H. Sun, G. Chen, R. Wang, and M. Li, “Visual Question Answering for Intelligent Interaction,” Mobile Information Systems, vol. 2022, pp. 1–6, 2022, doi: 10.1155/2022/4232968.

Q. Li, X. Tang, and Y. Jian, “Learning to Reason on Tree Structures for Knowledge-Based Visual Question Answering,” Sensors, vol. 22, no. 4, p. 1575, 2022, doi: 10.3390/s22041575.

R. Wang, S. Wu, and X. Wang, “The Core of Smart Cities: Knowledge Representation and Descriptive Framework Construction in Knowledge-Based Visual Question Answering,” Sustainability, vol. 14, no. 20, p. 13236, 2022, doi: 10.3390/su142013236.

R. Wang et al., “FTN–VQA: MULTIMODAL REASONING BY LEVERAGING A FULLY TRANSFORMER-BASED NETWORK FOR VISUAL QUESTION ANSWERING,” Fractals, vol. 31, no. 06, p. 2340133, 2023, doi: 10.1142/S0218348X23401333.

H. Wang, J. Li, and J. Wang, “Retrieving Chinese Questions and Answers Based on Deep-Learning Algorithm,” Mathematics, vol. 11, no. 18, pp. 1–18, 2023, doi: 10.3390/math11183843.

H. Zhu, R. Togo, T. Ogawa, and M. Haseyama, “Multimodal Natural Language Explanation Generation for Visual Question Answering Based on Multiple Reference Data,” Electronics (Basel)., vol. 12, no. 10, p. 2183, 2023, doi: 10.3390/electronics12102183.

J. He et al., “PERS: Parameter-Efficient Multimodal Transfer Learning for Remote Sensing Visual Question Answering,” IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., vol. 17, pp. 14823–14835, 2024, doi: 10.1109/JSTARS.2024.3447086.

T. Yamane, P. Chun, J. Dang, and T. Okatani, “Deep learning-based bridge damage cause estimation from multiple images using visual question answering,” Structure and Infrastructure Engineering, vol. 22, no. 4, pp. 734–747, 2026, doi: 10.1080/15732479.2024.2355929.

C. Song, “Enhancing Multimodal Understanding With LIUS: A Novel Framework for Visual Question Answering in Digital Marketing,” Journal of Organizational and End User Computing, vol. 36, no. 1, pp. 1–17, 2024, doi: 10.4018/JOEUC.336276.

Y. Cong and H. Mo, “A visual question answering method based on task decomposition,” PLoS One, vol. 20, no. 11, p. e0336623, 2025, doi: 10.1371/journal.pone.0336623.

Z. Liu, K. Yang, Y. Weng, Z. He, X. Liu, and H. Gao, “SCAG: Semantic Co-occurring Attention Guided Alignment for Knowledge-based Visual Question Answering,” ACM Transactions on Multimedia Computing, Communications, and Applications, vol. 21, no. 7, pp. 1–20, 2025, doi: 10.1145/3734220.

C. Huang and Z. Hu, “A multimodal transformer-based visual question answering method integrating local and global information,” PLoS One, vol. 20, no. 7, p. e0324757, 2025, doi: 10.1371/journal.pone.0324757.

M. S. Islam et al., “MRAN-VQA: Multimodal Recursive Attention Network for Visual Question Answering,” Engineering Science and Technology, an International Journal, vol. 72, p. 102232, 2025, doi: 10.1016/j.jestch.2025.102232.

C. Gupta, N. S. Gill, P. Gulia, and G. Pau, “CODNet: Context-based object detection network for multimodal image captioning and virtual question answering,” Image Vis. Comput., vol. 163, p. 105768, 2025, doi: 10.1016/j.imavis.2025.105768.

J. Xiao and Z. Zhang, “EduVQA: A multimodal Visual Question Answering framework for smart education,” Alexandria Engineering Journal, vol. 122, pp. 615–624, 2025, doi: 10.1016/j.aej.2025.03.005.

N. Hu, X. Zhang, Q. Zhang, W. Huo, and S. You, “ZPVQA: Visual Question Answering of Images Based on Zero-Shot Prompt Learning,” IEEE Access, vol. 13, pp. 50849–50859, 2025, doi: 10.1109/ACCESS.2025.3550942.

J. Ma et al., “DeepSORT-OCR: Design and Application Research of a Maritime Ship Target Tracking Algorithm Incorporating Hull Number Features,” Mathematics, vol. 14, no. 6, p. 1062, 2026, doi: 10.3390/math14061062.

H. Mohammadi and S. H. Khasteh, “A fast text similarity measure for large document collections using multireferencecosine and genetic algorithm,” TURKISH JOURNAL OF ELECTRICAL ENGINEERING & COMPUTER SCIENCES, vol. 28, no. 2, pp. 999–1013, 2020, doi: 10.3906/elk-1906-30.

W. Budiharto, V. Andreas, and A. A. S. Gunawan, “Deep learning-based question answering system for intelligent humanoid robot,” J. Big Data, vol. 7, no. 1, 2020, doi: 10.1186/s40537-020-00341-6.

L.-Q. Cai, M. Wei, S.-T. Zhou, and X. Yan, “Intelligent Question Answering in Restricted Domains Using Deep Learning and Question Pair Matching,” IEEE Access, vol. 8, pp. 32922–32934, 2020, doi: 10.1109/ACCESS.2020.2973728.

Y. Wen, X. Zhu, and L. Zhang, “CQACD: A Concept Question-Answering System for Intelligent Tutoring Using a Domain Ontology With Rich Semantics,” IEEE Access, vol. 10, pp. 67247–67261, 2022, doi: 10.1109/ACCESS.2022.3185400.

X. Yang, J. Yang, R. Li, H. Li, H. Zhang, and Y. Zhang, “Complex Knowledge Base Question Answering for Intelligent Bridge Management Based on Multi-Task Learning and Cross-Task Constraints,” Entropy, vol. 24, no. 12, p. 1805, 2022, doi: 10.3390/e24121805.

Y. Fang, J. Deng, F. Zhang, and H. Wang, “An Intelligent Question-Answering Model over Educational Knowledge Graph for Sustainable Urban Living,” Sustainability, vol. 15, no. 2, p. 1139, 2023, doi: 10.3390/su15021139.

M. Gao, M. Li, T. Ji, N. Wang, G. Lin, and Q. Wu, “Key Technologies of Intelligent Question-Answering System for Power System Rules and Regulations Based on Improved BERTserini Algorithm,” Processes, vol. 12, no. 1, p. 58, 2023, doi: 10.3390/pr12010058.

S. Zhang and Q. Jaamour, “English Intelligent Question Answering System Based on elliptic fitting equation,” Applied Mathematics and Nonlinear Sciences, vol. 8, no. 1, pp. 1743–1752, 2023, doi: 10.2478/amns.2022.2.0162.

R. Li, G. Ren, J. Yan, B. Zou, and Q. Liu, “Intelligent question answering system for traditional Chinese medicine based on BSG deep learning model: taking prescription and Chinese materia medica as examples,” Digital Chinese Medicine, vol. 7, no. 1, pp. 47–55, Mar. 2024, doi: 10.1016/j.dcmed.2024.04.006.

P. Lyu, J. Fu, C. Liu, W. Yu, and L. Xia, “GPB and BAC: two novel models towards building an intelligent motor fault maintenance question answering system,” Journal of Engineering Design, vol. 36, no. 11, pp. 2007–2027, 2025, doi: 10.1080/09544828.2024.2335135.

A. Mohamed, K. Abdelqader, and K. Shaalan, “Machine learning and deep learning techniques in Arabic question answering systems: innovations and challenges,” PeerJ Comput. Sci., vol. 11, pp. 1–35, 2025, doi: 10.7717/peerj-cs.3331.

E. Bernasconi, D. Redavid, and S. Ferilli, “Enhancing Personalised Learning with a Context-Aware Intelligent Question-Answering System and Automated Frequently Asked Question Generation,” Electronics (Switzerland), vol. 14, no. 7, pp. 1–26, 2025, doi: 10.3390/electronics14071481.

B. Zhang, X. Zhang, Q. Wang, G. Gui, and L. Shan, “Intelligent Question Answering System Design with Domain-Specific Knowledge Graphs,” IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences, vol. E108.A, no. 3, pp. 546–554, 2025, doi: 10.1587/transfun.2024EAP1090.

D. Sharma, S. Purushotham, and C. K. Reddy, “MedFuseNet: An attention-based multimodal deep learning model for visual question answering in the medical domain,” Sci. Rep., vol. 11, no. 1, p. 19826, 2021, doi: 10.1038/s41598-021-98390-1.

J. D. Silva, B. Martins, and J. Magalhães, “Contrastive training of a multimodal encoder for medical visual question answering,” Intelligent Systems with Applications, vol. 18, p. 200221, 2023, doi: 10.1016/j.iswa.2023.200221.

J. Holland, A. McGarvey, M. Flood, P. Joyce, and T. Pawlikowska, “A Qualitative Exploration of Student Cognition When Answering Text-Only or Image-Based Histology Multiple-Choice Questions,” Med. Sci. Educ., vol. 34, no. 6, pp. 1317–1329, 2024, doi: 10.1007/s40670-024-02104-x.

L. Chen, “Medical education and artificial intelligence: Question answering for medical questions based on intelligent interaction,” Concurr. Comput., vol. 36, no. 14, p. e8079, 2024, doi: 10.1002/cpe.8079.

B. Lei and P. Yin, “Technological Evolution and Research Trends of Intelligent Question-Answering Systems in Healthcare,” Healthcare, vol. 13, no. 18, p. 2269, 2025, doi: 10.3390/healthcare13182269.

F. Lu, S. Liu, W. Lu, P. Chen, and B. Ding, “Medical Knowledge-Based Differential Image Visual Question Answering,” IEEE Access, vol. 13, pp. 93818–93829, 2025, doi: 10.1109/ACCESS.2025.3565695.

T. Seki, Y. Kawazoe, H. Ito, Y. Akagi, T. Takiguchi, and K. Ohe, “Assessing the performance of zero-shot visual question answering in multimodal large language models for 12-lead ECG image interpretation,” Front. Cardiovasc. Med., vol. 12, p. 1458289, 2025, doi: 10.3389/fcvm.2025.1458289.[61] Y. Lu et al., “Application of Multimodal Transformer Model in Intelligent Agricultural Disease Detection and Question-Answering Systems,” Plants, vol. 13, no. 7, p. 972, 2024, doi: 10.3390/plants13070972.

W. Bi, Q. Xiong, X. Chen, Q. Du, J. Wu, and Z. Zhuang, “Intelligent visual question answering in TCM education: An innovative application of IoT and multimodal fusion,” Alexandria Engineering Journal, vol. 118, no. December 2024, pp. 325–336, 2025, doi: 10.1016/j.aej.2024.12.052.

S. Devmane, O. Rana, and C. Perera, “OntoSage: Intelligent Human-Building Smartbot for Semantic Smart Building Question Answering,” World Wide Web, vol. 29, no. 2, p. 17, 2026, doi: 10.1007/s11280-026-01403-0.

Y. Huang, H. Liu, S. Li, W. Wang, and Z. Zhou, “Effective Prediction and Important Counseling Experience for Perceived Helpfulness of Social Question and Answering-Based Online Counseling: An Explainable Machine Learning Model,” Front. Public Health, vol. 10, no. December, pp. 1–19, 2022, doi: 10.3389/fpubh.2022.817570.

Y. Lan, Y. Guo, Q. Chen, S. Lin, Y. Chen, and X. Deng, “Visual question answering model for fruit tree disease decision-making based on multimodal deep learning,” Front. Plant Sci., vol. 13, p. 1064399, 2023, doi: 10.3389/fpls.2022.1064399.

S. H. Kim, J. S. Shin, H. Lee, S. Y. Shin, K. M. Kang, and S. Y. Song, “A Comparative Analysis of GPT-3.5, GPT-4, GPT–4 Omni, Gemini Advanced, and Gemini 1.5 in Answering Frequently Asked Questions Regarding High Tibial Osteotomy,” Orthop. J. Sports Med., vol. 13, no. 11, pp. 1–13, 2025, doi: 10.1177/23259671251385127.

R. Zhu, B. Liu, R. Zhang, S. Zhang, and J. Cao, “OEQA: Knowledge- and Intention-Driven Intelligent Ocean Engineering Question-Answering Framework,” Applied Sciences, vol. 13, no. 23, p. 12915, 2023, doi: 10.3390/app132312915.

Unduhan

Telah diserahkan

10-06-2026

Diterima

20-08-2026

Diterbitkan

28-08-2026

Cara Mengutip

[1]
S. Arum Chrysantia dan R. Habibi, “Tinjauan Literatur Sistematis: Teknik Optical Character Recognition Dan In-formation Retrieval Pada Sistem Tanya Jawab”, TEKNOSI, vol. 12, no. 2, hlm. 284–299, Agu 2026.

Terbitan

Bagian

Articles

Artikel Serupa

<< < 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 > >> 

Anda juga bisa Mulai pencarian similarity tingkat lanjut untuk artikel ini.