Evaluation of ChatGPT Response Quality for Drug Information Queries: A Study of Medications Selected from National List of Essential Medicines

Main Article Content

Sommai Pimoaub
Thitima Tungsuay

Abstract

Objective: To evaluate the quality of responses generated by ChatGPT regarding medications listed in the National List of Essential Medicines (NLEM), focusing on accuracy, appropriateness, completeness, consistency, readability, and safety, and to analyze factors associated with meeting the quality threshold. Methods: This cross-sectional analytical study evaluated 960 ChatGPT responses concerning indications, dosages, administration, and adverse drug reactions of both modern and herbal medicines in the NLEM. Questions included both open-ended and closed-ended formats, administered across four time points: days 0, 7, 14, and 28. On day 0, three consecutive queries were performed, whereas subsequent days involved a single query. Two expert pharmacists independently assessed the responses using a 5-point Likert scale across six dimensions. Inter-rater reliability was analyzed using weighted Cohen’s kappa and intraclass correlation coefficient (ICC), while factors associated with passing the quality threshold were analyzed using generalized estimating equations (GEE). Results: Responses generated by ChatGPT achieved mean scores exceeding 4.0 out of 5 across all evaluation dimensions. The highest mean score was observed in accuracy (4.60 ± 0.67), followed by safety (4.43 ± 0.69) and appropriateness (4.40 ± 0.69). More than 92% of responses in all dimensions met the quality threshold (score ≥ 4). Inter-rater agreement was low across all domains. GEE analysis revealed no statistically significant associations between the likelihood of meeting the quality threshold and question category, drug class, question type, or assessment time point. Conclusion: ChatGPT demonstrates significant potential in providing information on NLEM medications, with high-quality responses across all evaluated dimensions. This study found no significant variations in response quality over the assessment period. However, ChatGPT may still encounter limitations regarding completeness and consistency in certain contexts. Therefore, it should not replace clinical decision-making or professional counseling by healthcare providers. Further studies in more diverse and complex clinical contexts are warranted to support the appropriate and safe application of artificial intelligence for medication-related inquiries.

Article Details

Section
Research Articles

References

National Statistical Office. The 2021 health and welfare survey [online]. 2022 [cited May 26, 2026]. Available from: www.nso.go.th/nsoweb/storage/sur vey_detail/2023/20230505170315_44270.pdf

Aekplakorn W. Thailand's national health examination survey VI 2019–2020 [online]. 2021 [cited May 26, 2026]. Available from: www.hiso.or.th/hiso/picture/re portHealth/report/sreport6/sreport6_full.pdf

National Drug System Development Committee. National list of essential medicines B.E. 2565 [online]. 2022 [cited May 27, 2026]. Available from: ndi.fda.moph.go.th/uploads/file_news/20220808893215585.PDF

National Drug System Development Committee. National list of herbal medicinal products B.E. 2566 [online]. 2023 [cited May 27, 2026]. Available from: herbal.fda.moph.go.th/media.php?id=865141728649289728&name=Main medicine-66.pdf

Sumpradit N, Chongtrakul P, Anuwong K, Pumtong S, Kongsomboon K, Butdeemee P, et al. Antibiotics smart use: a workable model for promoting the rational use of medicines in Thailand. Bull World Health Organ. 2012; 90: 905-13.

Yun HS, Bickmore T. Online health information–seeking in the era of large language models: cross-sectional web-based survey study. J Med Internet Res. 2025; 27: e68560.

Fridman I, Johnson S, Elston Lafata J. Health information and misinformation: a framework to guide research and practice. JMIR Med Educ. 2023; 9: e38687.

Sallam M. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare (Basel). 2023; 11: 887.

Walker HL, Ghani S, Kuemmerli C, Nebiker CA, Müller BP, Raptis DA, et al. Reliability of medical information provided by ChatGPT: assessment against clinical guidelines and patient information quality instrument. J Med Internet Res. 2023; 25: e47479.

Pornwattanakavee S, Leelakanok N, Todsarot T, Guinto GAT, Takun R, Sumativit A, et al. Effectiveness of ChatGPT, Google Gemini, and Microsoft Copilot in answering Thai drug information queries: cross-sectional study. JMIR AI 2025;4: e79751.

Aydin S, Karabacak M, Vlachos V, Margetis K. Navigating the potential and pitfalls of large language models in patient-centered medication guidance and self-decision support. Front Med (Lausanne). 2025; 12: 1527864.

Wei Q, Yao Z, Cui Y, Wei B, Jin Z, Xu X. Evaluation of ChatGPT-generated medical responses: a systematic review and meta-analysis. J Biomed Inform 2024; 151: 104620.

Wang RY, Strong DM. Beyond accuracy: what data quality means to data consumers. J Manag Inf Syst. 1996; 12: 5-33.

Wolters Kluwer. UpToDate Lexidrug (formerly Lexicomp) [online]. 2026 [cited Jun 13, 2026]. Available from: online.lexi.com

Merative. DRUGDEX® System [online]. 2026 [cited Jun 13, 2026]. Available from: www.micromedex solutions.com

Lexicomp. Adult drug information handbook 2025-2026. 33rd ed. Hudson (OH): Wolters Kluwer Clinical Drug Information, Inc.; 2024.

Meyer A, Schömig E, Streichert T. ChatGPT and reference intervals: a comparative analysis of repeatability in GPT-3.5 Turbo, GPT-4, and GPT-4o. Front Artif Intell. 2025; 8: 1681979.

Funk PF, Hoch CC, Knoedler S, Knoedler L, Cotofana S, Sofo G, et al. ChatGPT's response consistency: a study on repeated queries of medical examination questions. Eur J Investig Health Psychol Educ. 2024; 14: 657-68.

Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977; 33 :159-74.

Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016; 15: 155-63.

Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepaño C, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLoS Digit Health. 2023; 2: e0000198.

Morath B, Chiriac U, Jaszkowski E, Deiß C, Nürnberg H, Hörth K, et al. Performance and risks of ChatGPT used in drug information: an exploratory real-world analysis. Eur J Hosp Pharm. 2024; 31: 491-7.

Roosan D, Padua P, Khan R, Khan H, Verzosa C, Wu Y. Effectiveness of ChatGPT in clinical pharmacy and the role of artificial intelligence in medication therapy management. J Am Pharm Assoc (2003). 2024; 64: 422-28.

Huang X, Estau D, Liu X, Yu Y, Qin J, Li Z. Evaluating the performance of ChatGPT in clinical pharmacy: a comparative study of ChatGPT and clinical pharmacists. Br J Clin Pharmacol. 2024; 90: 232-8.

Terwee CB, Bot SDM, de Boer MR, van der Windt DAWM, Knol DL, Dekker J, et al. Quality criteria were proposed for measurement properties of health status questionnaire. J Clin Epidemiol. 2007; 60: 34-42.

Limsuwanchote S, Sakunphueak A, Boonrit N, Hopkins AM, Ruanglertboon W. Analysing credibility of information on Thai herbs generated by ChatGPT from pharmacists' perspectives. Thai Journal of Pharmacy Practice 2024; 16: 1257-76.

Huang L, Yu W, Ma W, Zhong W, Feng Z, Wang H, et al. A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Trans Inf Syst [online]. 2023 [cited May 26, 2026]. doi.org/10.1145/3703155

Kelly CJ, Karthikesalingam A, Suleyman M, Corrado G, King D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019; 17: 195.