Forschungsteam vor Bildschirmen mit Visualisierungen künstlicher neuronaler Netze

Projekt

Reply to Li and Kang

10.1055/a-2849-7511 We thank Li and Kang for their thoughtful engagement with our pilot evaluation of ChatGPT-4o for patient education on gastric cancer prevention [ 1 ]. Their commentary raises three interrelated issues – prompt-dependent variability, the limits of subjective accuracy ratings as a proxy for hallucina…

10.1055/a-2849-7511 We thank Li and Kang for their thoughtful engagement with our pilot evaluation of ChatGPT-4o for patient education on gastric cancer prevention [ 1 ]. Their commentary raises three interrelated issues – prompt-dependent variability, the limits of subjective accuracy ratings as a proxy for hallucination detection, and the absence of an independent reference standard – that we believe constructively delineate the next steps for this line of work. We fully agree that output variability is a defining feature of contemporary large language models and one that distinguishes them from static educational resources. Beyond the contextual effects emphasized by Li and Kang – prior prompts, conversational drift, and user-introduced inaccuracies – are differences in how source documents are tokenized or chunked at ingestion. Our decision to fix a single prompt, a single guideline, and a single model version was deliberate, intended to maximize internal validity in a pilot setting, and we explicitly identified this as a limitation while calling for systematic prompt-engineering studies and head-to-head comparisons across models and versions. The call by Li and Kang for standardized prompts and concise operational frameworks is the natural practical complement to such research: clinicians and patient organizations will ultimately need vetted, version-controlled prompt templates – analogous to validated clinical instruments – rather than ad hoc instructions, if artificial intelligence (AI)-assisted patient education is to be deployed reliably and equitably across users. On the question of hallucinations, we concur that the absence of obvious factual errors, inferred from expert accuracy scores, is an indirect and insufficient safeguard. Anchoring the model to the source guideline (“retrieval-grounded” generation) reduces, but does not eliminate, the risk of subtle distortions such as selective omission, miscalibrated emphasis, or unwarranted certainty around graded recommendations. We therefore endorse the recommendation of Li and Kang for structured hallucination audits with predefined error taxonomies and independent adjudication, and we anticipated this need explicitly in our discussion. Regarding the lack of an independent reference standard, the critique is fair. The Digestive Cancers Europe (DiCE) material represents current real-world practice rather than a validated gold standard, and concordance between the two summaries may reflect shared blind spots – most evidently in their joint failure to meet American Medical Association (AMA) readability targets. A more rigorous benchmark would, in our view, combine a structured rubric derived directly from the graded recommendations of the source guideline, automated faithfulness and readability scoring, and ultimately patient-centered end points such as comprehension, retention, and screening adherence. We thank Li and Kang for sharpening these considerations. Their commentary reinforces what we regard as the central message of our study: that the comparable subjective ratings of AI-generated and traditional materials should be interpreted with caution, and that the responsible integration of generative AI into patient education will require formal validation frameworks, transparent governance, and sustained human oversight before any patient-facing dissemination. Publication History Article published online: 22 July 2026 © 2026. Thieme. All rights reserved. Georg Thieme Verlag KG Oswald-Hesse-Straße 50, 70469 Stuttgart, Germany

Technologien

Themengebiete

Hochschulen