การประเมินคุณภาพของคำตอบจาก ChatGPT สำหรับคำถามด้านข้อมูลยา : การศึกษาในรายการยาที่คัดเลือกจากบัญชียาหลักแห่งชาติ
Main Article Content
บทคัดย่อ
วัตถุประสงค์: เพื่อประเมินคุณภาพคำตอบที่ได้จาก ChatGPT เมื่อสอบถามข้อมูลเกี่ยวกับยาในบัญชียาหลักแห่งชาติ โดยประเมินในด้านความถูกต้อง ความเหมาะสม ความครบถ้วน ความสม่ำเสมอ ความเข้าใจง่ายของภาษา และความปลอดภัย รวมทั้งวิเคราะห์ปัจจัยที่สัมพันธ์กับการผ่านเกณฑ์คุณภาพของคำตอบ วิธีการ: การวิจัยเชิงวิเคราะห์แบบภาคตัดขวางครั้งนี้ประเมินคำตอบจาก ChatGPT จำนวน 960 คำตอบเมื่อถามด้วยคำถามที่เกี่ยวกับข้อบ่งใช้ ขนาดยา วิธีใช้ยา และอาการไม่พึงประสงค์ของยาในบัญชียาหลักแห่งชาติ ทั้งยาแผนปัจจุบันและยาสมุนไพร โดยถามด้วยคำถามแบบปลายเปิดและปลายปิดใน 4 ช่วงเวลา คือ วันที่ 0, 7, 14 และ 28 โดยการถามในวันที่ 0 จะสอบถาม 3 ครั้งแบบต่อเนื่อง แต่การสอบถามในวันอื่น ๆ เป็นการสอบถามเพียงครั้งเดียว เภสัชกรผู้เชี่ยวชาญ 2 คนประเมินคำตอบอย่างเป็นอิสระต่อกันโดยใช้มาตรวัดลิเคิร์ต 5 ระดับใน 6 มิติของการประเมิน การวิเคราะห์ความสอดคล้องระหว่างผู้ประเมินใช้ weighted Cohen’s kappa และ intraclass correlation coefficient (ICC) การวิเคราะห์ปัจจัยที่สัมพันธ์กับการผ่านเกณฑ์คุณภาพของคำตอบใช้ generalized estimating equations (GEE) ผลการวิจัย: คำตอบจาก ChatGPT มีคะแนนเฉลี่ยมากกว่า 4.0 จากคะแนนเต็ม 5 ในทุกมิติการประเมิน โดยด้านความถูกต้องมีคะแนนเฉลี่ยสูงสุด (4.60±0.67) รองลงมาคือความปลอดภัย (4.43±0.69) และความเหมาะสม (4.40±0.69) ในทุกมิติมีสัดส่วนของคำตอบที่ผ่านเกณฑ์การประเมิน (ได้คะแนน ≥4) มากกว่าร้อยละ 92 ความสอดคล้องระหว่างผู้ประเมินอยู่ในระดับต่ำในทุกด้านของการประเมิน ผลการวิเคราะห์ด้วย GEE ไม่พบความสัมพันธ์อย่างมีนัยสำคัญทางสถิติระหว่างโอกาสการผ่านเกณฑ์คุณภาพของคำตอบ กับหมวดคำถาม ประเภทยา ประเภทคำถาม หรือช่วงเวลาการประเมิน สรุป: ChatGPT มีศักยภาพในการให้ข้อมูลยาในบัญชียาหลักแห่งชาติ โดยคุณภาพของคำตอบอยู่ในระดับสูงในทุกมิติของการประเมิน การศึกษาไม่พบความแตกต่างของคุณภาพคำตอบตามช่วงเวลาการประเมินในการศึกษานี้ อย่างไรก็ตาม ChatGPT ยังอาจมีข้อจำกัดด้านความครบถ้วนและความสม่ำเสมอในบางบริบท และไม่ควรใช้ทดแทนการตัดสินใจทางคลินิกหรือการให้คำแนะนำโดยบุคลากรทางการแพทย์และสาธารณสุข ทั้งนี้ ควรมีการศึกษาเพิ่มเติมในบริบททางคลินิกที่หลากหลายและซับซ้อนมากขึ้นเพื่อสนับสนุนการประยุกต์ใช้ปัญญาประดิษฐ์เพื่อตอบคำถามเกี่ยวกับข้อมูลยาอย่างเหมาะสมและปลอดภัย
Article Details

อนุญาตภายใต้เงื่อนไข Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
ผลการวิจัยและความคิดเห็นที่ปรากฏในบทความถือเป็นความคิดเห็นและอยู่ในความรับผิดชอบของผู้นิพนธ์ มิใช่ความเห็นหรือความรับผิดชอบของกองบรรณาธิการ หรือคณะเภสัชศาสตร์ มหาวิทยาลัยสงขลานครินทร์ ทั้งนี้ไม่รวมความผิดพลาดอันเกิดจากการพิมพ์ บทความที่ได้รับการเผยแพร่โดยวารสารเภสัชกรรมไทยถือเป็นสิทธิ์ของวารสารฯ
เอกสารอ้างอิง
National Statistical Office. The 2021 health and welfare survey [online]. 2022 [cited May 26, 2026]. Available from: www.nso.go.th/nsoweb/storage/sur vey_detail/2023/20230505170315_44270.pdf
Aekplakorn W. Thailand's national health examination survey VI 2019–2020 [online]. 2021 [cited May 26, 2026]. Available from: www.hiso.or.th/hiso/picture/re portHealth/report/sreport6/sreport6_full.pdf
National Drug System Development Committee. National list of essential medicines B.E. 2565 [online]. 2022 [cited May 27, 2026]. Available from: ndi.fda.moph.go.th/uploads/file_news/20220808893215585.PDF
National Drug System Development Committee. National list of herbal medicinal products B.E. 2566 [online]. 2023 [cited May 27, 2026]. Available from: herbal.fda.moph.go.th/media.php?id=865141728649289728&name=Main medicine-66.pdf
Sumpradit N, Chongtrakul P, Anuwong K, Pumtong S, Kongsomboon K, Butdeemee P, et al. Antibiotics smart use: a workable model for promoting the rational use of medicines in Thailand. Bull World Health Organ. 2012; 90: 905-13.
Yun HS, Bickmore T. Online health information–seeking in the era of large language models: cross-sectional web-based survey study. J Med Internet Res. 2025; 27: e68560.
Fridman I, Johnson S, Elston Lafata J. Health information and misinformation: a framework to guide research and practice. JMIR Med Educ. 2023; 9: e38687.
Sallam M. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare (Basel). 2023; 11: 887.
Walker HL, Ghani S, Kuemmerli C, Nebiker CA, Müller BP, Raptis DA, et al. Reliability of medical information provided by ChatGPT: assessment against clinical guidelines and patient information quality instrument. J Med Internet Res. 2023; 25: e47479.
Pornwattanakavee S, Leelakanok N, Todsarot T, Guinto GAT, Takun R, Sumativit A, et al. Effectiveness of ChatGPT, Google Gemini, and Microsoft Copilot in answering Thai drug information queries: cross-sectional study. JMIR AI 2025;4: e79751.
Aydin S, Karabacak M, Vlachos V, Margetis K. Navigating the potential and pitfalls of large language models in patient-centered medication guidance and self-decision support. Front Med (Lausanne). 2025; 12: 1527864.
Wei Q, Yao Z, Cui Y, Wei B, Jin Z, Xu X. Evaluation of ChatGPT-generated medical responses: a systematic review and meta-analysis. J Biomed Inform 2024; 151: 104620.
Wang RY, Strong DM. Beyond accuracy: what data quality means to data consumers. J Manag Inf Syst. 1996; 12: 5-33.
Wolters Kluwer. UpToDate Lexidrug (formerly Lexicomp) [online]. 2026 [cited Jun 13, 2026]. Available from: online.lexi.com
Merative. DRUGDEX® System [online]. 2026 [cited Jun 13, 2026]. Available from: www.micromedex solutions.com
Lexicomp. Adult drug information handbook 2025-2026. 33rd ed. Hudson (OH): Wolters Kluwer Clinical Drug Information, Inc.; 2024.
Meyer A, Schömig E, Streichert T. ChatGPT and reference intervals: a comparative analysis of repeatability in GPT-3.5 Turbo, GPT-4, and GPT-4o. Front Artif Intell. 2025; 8: 1681979.
Funk PF, Hoch CC, Knoedler S, Knoedler L, Cotofana S, Sofo G, et al. ChatGPT's response consistency: a study on repeated queries of medical examination questions. Eur J Investig Health Psychol Educ. 2024; 14: 657-68.
Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977; 33 :159-74.
Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016; 15: 155-63.
Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepaño C, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLoS Digit Health. 2023; 2: e0000198.
Morath B, Chiriac U, Jaszkowski E, Deiß C, Nürnberg H, Hörth K, et al. Performance and risks of ChatGPT used in drug information: an exploratory real-world analysis. Eur J Hosp Pharm. 2024; 31: 491-7.
Roosan D, Padua P, Khan R, Khan H, Verzosa C, Wu Y. Effectiveness of ChatGPT in clinical pharmacy and the role of artificial intelligence in medication therapy management. J Am Pharm Assoc (2003). 2024; 64: 422-28.
Huang X, Estau D, Liu X, Yu Y, Qin J, Li Z. Evaluating the performance of ChatGPT in clinical pharmacy: a comparative study of ChatGPT and clinical pharmacists. Br J Clin Pharmacol. 2024; 90: 232-8.
Terwee CB, Bot SDM, de Boer MR, van der Windt DAWM, Knol DL, Dekker J, et al. Quality criteria were proposed for measurement properties of health status questionnaire. J Clin Epidemiol. 2007; 60: 34-42.
Limsuwanchote S, Sakunphueak A, Boonrit N, Hopkins AM, Ruanglertboon W. Analysing credibility of information on Thai herbs generated by ChatGPT from pharmacists' perspectives. Thai Journal of Pharmacy Practice 2024; 16: 1257-76.
Huang L, Yu W, Ma W, Zhong W, Feng Z, Wang H, et al. A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Trans Inf Syst [online]. 2023 [cited May 26, 2026]. doi.org/10.1145/3703155
Kelly CJ, Karthikesalingam A, Suleyman M, Corrado G, King D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019; 17: 195.