Information System Development through Image Data Extraction using Optical Character Recognition (OCR) Technology

Authors

  • Muckkawa Jeemsuwan College of Engineering and Technology, Dhurakij Pundit University
  • Supakpong Jinarat College of Engineering and Technology, Dhurakij Pundit University
  • Nutchanun Chinpanthana College of Engineering and Technology, Dhurakij Pundit University

Keywords:

data extraction, Optical Character Recognition (OCR), Large Language Model (LLM), digital data storage

Abstract

This research aims to develop an information system for managing and storing fuel receipt data within organizations by integrating Optical Character Recognition: OCR technology with Large Language Models : LLMs to improve data accuracy and reduce manual processing efforts. OCR is used to detect and extract textual information from receipt documents, while LLMs are applied to correct extraction errors such as misspellings and formatting inconsistencies and to generate a structured, ready-to-use database. The system development process consisted of three main stages: (1) analysis of background, significance, and problem context; (2) system design and development; and (3) system testing under real-world conditions. The experimental evaluation assessed data extraction accuracy across receipts with different formats. The samples were divided into two groups for comparative analysis. The results indicate that, in both groups, the system successfully extracted key information—including transaction date, fuel volume, license plate number, and total amount—with an average accuracy rate of 95.98%. Furthermore, the system reduced the time required to compile fuel receipt records from three days to one day and lowered monthly operational costs by 50.35% compared to the traditional manual process. These findings demonstrate that the proposed information system significantly enhances the accuracy, reliability, and efficiency of fuel receipt data management while providing cost-effective advantages for practical organizational and commercial implementation. Additionally, the system facilitates the transition from conventional document-based storage to efficient and sustainable digital data management.

References

Aayush, N., Aayush, L., Ankit, P., & Amam S. (2025). Structured information extraction from Nepali scanned documents using layout transformer and LLMs. In K. Sarveswaran, A. Vaidya, B. Bal, S. Shams & S. Thapa (Eds.), Proceedings of the First Workshop on Challenges in Processing South Asian Languages (CHIPSAL 2025) (pp. 100–110). International Committee on Computational Linguistics.

Abdalla, M., Kasem, M. S., Mahmoud, M., Yagoub, B., Senussi, M. F., Abdallah, A., Kang, S. H., & Kang, H. S. (2025). ReceiptQA: A question-answering dataset for receipt understanding. Mathematics, 13(11), 1760. https://doi.org/10.3390/math13111760.

Anakpluek, N., Pasanta, W., Chantharasukha, L., Chokratansombat, P., Kanjanakaew, P., & Siriborvornratanakul, T. (2025). Improved tesseract optical character recognition performance on Thai document datasets. Big Data Research, 39, 100508. https://doi.org/10.1016/j.bdr.2025.100508

Bharadwaj, A., El Sawy, O. A., Pavlou, P. A., & Venkatraman, N. (2013). Digital business strategy: Toward a next generation of insights. MIS Quarterly, 37(2), 471–482.https://ssrn.com/abstract=2742300

Chompunut, A., & Rajalida, L. (2024). Menu item extraction from Thai receipt images using deep learning and template-based information extraction. In H. Shen, S. C. Tan, X. Jiang, X. Li & R. Latip (Eds.), Proceedings of the 6th International Conference on Information Technology and Computer Communications (ITCC 2024) (pp. 107-113). ACM.

Do, T., Tran, D. P., Vo, A., & Kim, D. (2025). Reference-Based post-OCR processing with LLM for precise diacritic text in historical document recognition. Proceedings of the AAAI Conference on Artificial Intelligence, 39(27), 27951-27959. https://doi.org/10.1609/aaai.v39i27.35012

Google Cloud. (2025). Vision AI: Image and visual AI tools. https://cloud.google.com/vision

Kumar, S. (2024). Autonomous document processing in the business sector using artificial intelligence. International Journal of Technoinformatics Engineering, 1(2), 33-40. https://aimbell.com/wp-content/uploads/2025/08/6-IJTE.pdf

Lefferts, S., & Kozenieski, D. (2025). Preprocessing images to improve OCR & DarkShield results. IRI. https://www.iri.com/blog/data-protection/preprocessing-images-for-ocr-darkshield/

Mankiw, N. G. (2016). Principles of economics (8th ed.). Cengage Learning.

Marangon, J. D. (2025). Google Cloud Vision API for image handling and OCR. Medium. https://medium.com/@johnidouglasmarangon/google-cloud-vision-api-for-image-handling-and-ocr-a6763969a2e6

Martinez, J. (2025). OCR preprocessing: How to improve your OCR extraction outcome. https://www.docuclipper.com/blog/ocr-preprocessing/

Patil, S., & Yadav, S. (2025). Automated expense tracking with OCR. International Advanced Research Journal in Science, Engineering and Technology (IARJSET), 12(1), 209–212. https://iarjset.com/wp-content/uploads/2025/02/IARJSET.2025.12142.pdf

Ramsey, S. (2025). Improving document content extraction with multi-modal LLM. Storytell. https://web.storytell.ai/blog/improving-document-content-extraction-with-multi-modal-llm

Smith, R. (2007). An overview of the Tesseract OCR engine. In F. Bortolozzi and R. Sabourin (Eds.), Ninth International Conference on Document Analysis and Recognition (ICDAR) (pp. 629–633). The Institute of Electrical and Electronics Engineers.

Thammarak, K., Kongkla, P., Sirisathitkul, Y., & Intakosum, S. (2022). Comparative analysis of Tesseract and Google Cloud Vision for Thai vehicle registration certificate. International Journal of Electrical and Computer Engineering (IJECE), 12(2), 1849–1858. https://doi.org/10.11591/ijece.v12i2.pp1849-1858

Westerman, G., Bonnet, D., & McAfee, A. (2014). Leading digital: Turning technology into business transformation. Harvard Business Review Press.

Yang, Y., Wu, Z., Yang, Y., Lian, S., Guo, F., & Wang, Z. (2022). A survey of information extraction based on deep learning. Applied Sciences, 12(19), 9691. http://doi.org/10.3390/app12199691

Downloads

Published

2026-04-21

How to Cite

Jeemsuwan, M., Jinarat, S., & Chinpanthana, N. (2026). Information System Development through Image Data Extraction using Optical Character Recognition (OCR) Technology. EAU Heritage Journal Science and Technology (online), 20(1), 96–109. retrieved from https://he01.tci-thaijo.org/index.php/EAUHJSci/article/view/282062

Issue

Section

Research Articles