A comparative analysis of large language models in digital content on painless labor: A systematic evaluation of readability, reliability, quality, and accuracy
Main Article Content
Abstract
Aim: Painless labor is an important concern in obstetric care, with both medical and psychosocial dimensions. With the increasing use of artificial intelligence (AI) in health communication, evaluating the quality of content generated by large language models (LLMs) is essential. This study aimed to compare AI-generated content on “painless labor” produced by three LLMs—ChatGPT (Chat Generative Pre-trained Transformer), Gemini, and DeepSeek—in terms of readability, reliability, content quality, and medical accuracy.
Materials and Methods: A total of 270 texts were generated using 30 frequently searched keywords on three dates in May 2025 via the free versions of the three LLMs. Readability was assessed with six validated indices. Reliability and quality were evaluated using the Modified DISCERN tool, the JAMA (Journal of the American Medical Association) Benchmark Criteria, the Ensuring Quality Information for Patients (EQIP)-36, and the Global Quality Score. The medical accuracy of 90 texts was assessed by an obstetric anesthesia expert.
Results: ChatGPT texts were the most readable and were written at lower grade levels, but scored lower in reliability and quality. Gemini ranked highest for both reliability and content quality, although it produced more complex language. DeepSeek showed variable performance. No significant differences were found in medical accuracy or content completeness. Gemini also demonstrated the most consistent performance across all time points.
Conclusion: LLMs vary substantially in how they present medical content. ChatGPT achieved higher readability scores, indicating simpler language structures, whereas Gemini demonstrated higher reliability and content quality metrics. However, the absence of source citations across all models raises concerns about the verifiability of content, highlighting the need for critical oversight in healthcare applications.
Downloads
Article Details
Section

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
CC Attribution-NonCommercial-NoDerivatives 4.0