Forschungsteam vor Bildschirmen mit Visualisierungen künstlicher neuronaler Netze

Projekt

Human vs. machine - Evaluating GPT-4, GPT-4o and GPT-4o-mini in rating depression interviews in a convenience sample of the German population

Despite high prevalence, depression often remains undetected. Artificial Intelligence (AI) screening methods may help address this gap. AI models’ proficiency in rating depression interviews in comparison with human raters has been demonstrated for clinical populations. The capacity of three Large Language Models (LLM…

Despite high prevalence, depression often remains undetected. Artificial Intelligence (AI) screening methods may help address this gap. AI models’ proficiency in rating depression interviews in comparison with human raters has been demonstrated for clinical populations. The capacity of three Large Language Models (LLMs) for rating depression interviews was evaluated. The test dataset contained interviews from the general population to evaluate model capacity for screening. LLM ratings were compared with assessments provided by trained psychology students. GPT-4, GPT-4o and GPT-4o-mini were prompted to rate interview transcripts from N = 55 participants from a convenience sample of the German population on a depression scale. Models were used without task-specific fine-tuning. LLM performance was evaluated via intraclass correlations with trained psychology students’ assessments leveraging a scale explicitly designed to yield reliable ratings when applied by first-time raters. GPT-4 and GPT-4o achieved moderate to good interrater reliability with student raters (ICC 0.76 and 0.75) while ICC for GPT-4o-mini was moderate (0.55), with 98.88% CIs [0.38 to 0.70] indicating greater uncertainty. F1 accuracy scores ranged from 0.56 for GPT-4o-mini to 0.75 for GPT-4. The results suggest GPT-4’s and GPT-4o’s potential for scoring depression interviews from the general population, with GPT-4o meeting sensitivity and specificity values for case finding instruments. Our results require further validation with clinician raters in a larger and more diverse sample. Future research should explore fine-tuning and potential model bias. • We compared three LLMs with trained student raters scoring depression interviews • GPT-4 and GPT-4o estimations showed moderate to good reliability with trained student raters • GPT-4o-mini’s scores showed less than moderate reliability • GPT-4o met required sensitivity and specificity values for case finding instruments • The models were neither pretrained, nor fine-tuned for this task.

Technologien

Themengebiete

Hochschulen