Information System Development through Image Data Extraction using Optical Character Recognition (OCR) Technology
Keywords:
data extraction, Optical Character Recognition (OCR), Large Language Model (LLM), digital data storageAbstract
This research aims to develop an information system for managing and storing fuel receipt data within organizations by integrating Optical Character Recognition: OCR technology with Large Language Models : LLMs to improve data accuracy and reduce manual processing efforts. OCR is used to detect and extract textual information from receipt documents, while LLMs are applied to correct extraction errors such as misspellings and formatting inconsistencies and to generate a structured, ready-to-use database. The system development process consisted of three main stages: (1) analysis of background, significance, and problem context; (2) system design and development; and (3) system testing under real-world conditions. The experimental evaluation assessed data extraction accuracy across receipts with different formats. The samples were divided into two groups for comparative analysis. The results indicate that, in both groups, the system successfully extracted key information—including transaction date, fuel volume, license plate number, and total amount—with an average accuracy rate of 95.98%. Furthermore, the system reduced the time required to compile fuel receipt records from three days to one day and lowered monthly operational costs by 50.35% compared to the traditional manual process. These findings demonstrate that the proposed information system significantly enhances the accuracy, reliability, and efficiency of fuel receipt data management while providing cost-effective advantages for practical organizational and commercial implementation. Additionally, the system facilitates the transition from conventional document-based storage to efficient and sustainable digital data management.
References
Aayush, N., Aayush, L., Ankit, P., & Amam S. (2025). Structured information extraction from Nepali scanned documents using layout transformer and LLMs. In K. Sarveswaran, A. Vaidya, B. Bal, S. Shams & S. Thapa (Eds.), Proceedings of the First Workshop on Challenges in Processing South Asian Languages (CHIPSAL 2025) (pp. 100–110). International Committee on Computational Linguistics.
Abdalla, M., Kasem, M. S., Mahmoud, M., Yagoub, B., Senussi, M. F., Abdallah, A., Kang, S. H., & Kang, H. S. (2025). ReceiptQA: A question-answering dataset for receipt understanding. Mathematics, 13(11), 1760. https://doi.org/10.3390/math13111760.
Anakpluek, N., Pasanta, W., Chantharasukha, L., Chokratansombat, P., Kanjanakaew, P., & Siriborvornratanakul, T. (2025). Improved tesseract optical character recognition performance on Thai document datasets. Big Data Research, 39, 100508. https://doi.org/10.1016/j.bdr.2025.100508
Bharadwaj, A., El Sawy, O. A., Pavlou, P. A., & Venkatraman, N. (2013). Digital business strategy: Toward a next generation of insights. MIS Quarterly, 37(2), 471–482.https://ssrn.com/abstract=2742300
Chompunut, A., & Rajalida, L. (2024). Menu item extraction from Thai receipt images using deep learning and template-based information extraction. In H. Shen, S. C. Tan, X. Jiang, X. Li & R. Latip (Eds.), Proceedings of the 6th International Conference on Information Technology and Computer Communications (ITCC 2024) (pp. 107-113). ACM.
Do, T., Tran, D. P., Vo, A., & Kim, D. (2025). Reference-Based post-OCR processing with LLM for precise diacritic text in historical document recognition. Proceedings of the AAAI Conference on Artificial Intelligence, 39(27), 27951-27959. https://doi.org/10.1609/aaai.v39i27.35012
Google Cloud. (2025). Vision AI: Image and visual AI tools. https://cloud.google.com/vision
Kumar, S. (2024). Autonomous document processing in the business sector using artificial intelligence. International Journal of Technoinformatics Engineering, 1(2), 33-40. https://aimbell.com/wp-content/uploads/2025/08/6-IJTE.pdf
Lefferts, S., & Kozenieski, D. (2025). Preprocessing images to improve OCR & DarkShield results. IRI. https://www.iri.com/blog/data-protection/preprocessing-images-for-ocr-darkshield/
Mankiw, N. G. (2016). Principles of economics (8th ed.). Cengage Learning.
Marangon, J. D. (2025). Google Cloud Vision API for image handling and OCR. Medium. https://medium.com/@johnidouglasmarangon/google-cloud-vision-api-for-image-handling-and-ocr-a6763969a2e6
Martinez, J. (2025). OCR preprocessing: How to improve your OCR extraction outcome. https://www.docuclipper.com/blog/ocr-preprocessing/
Patil, S., & Yadav, S. (2025). Automated expense tracking with OCR. International Advanced Research Journal in Science, Engineering and Technology (IARJSET), 12(1), 209–212. https://iarjset.com/wp-content/uploads/2025/02/IARJSET.2025.12142.pdf
Ramsey, S. (2025). Improving document content extraction with multi-modal LLM. Storytell. https://web.storytell.ai/blog/improving-document-content-extraction-with-multi-modal-llm
Smith, R. (2007). An overview of the Tesseract OCR engine. In F. Bortolozzi and R. Sabourin (Eds.), Ninth International Conference on Document Analysis and Recognition (ICDAR) (pp. 629–633). The Institute of Electrical and Electronics Engineers.
Thammarak, K., Kongkla, P., Sirisathitkul, Y., & Intakosum, S. (2022). Comparative analysis of Tesseract and Google Cloud Vision for Thai vehicle registration certificate. International Journal of Electrical and Computer Engineering (IJECE), 12(2), 1849–1858. https://doi.org/10.11591/ijece.v12i2.pp1849-1858
Westerman, G., Bonnet, D., & McAfee, A. (2014). Leading digital: Turning technology into business transformation. Harvard Business Review Press.
Yang, Y., Wu, Z., Yang, Y., Lian, S., Guo, F., & Wang, Z. (2022). A survey of information extraction based on deep learning. Applied Sciences, 12(19), 9691. http://doi.org/10.3390/app12199691
