Forschungsteam vor Bildschirmen mit Visualisierungen künstlicher neuronaler Netze

Projekt

Tone Is Not Judgment: An Empirical Case Study on Emotional Sycophancy in Child-Facing Commercial LLMs

Sycophancy — the tendency of RLHF-trained language models to prefer responses that align with user beliefs over responses that are truthful or corrective — is a documented property of commercial large language models. This paper documents an empirical observation that while frontier labs have begun addressing factual…

Sycophancy — the tendency of RLHF-trained language models to prefer responses that align with user beliefs over responses that are truthful or corrective — is a documented property of commercial large language models. This paper documents an empirical observation that while frontier labs have begun addressing factual sycophancy, emotional sycophancy remains largely unaddressed, with consequential impacts for child users. Grounded in five test conversations with a widely deployed commercial LLM (Gemini 3 Flash) conducted in July 2026, this study demonstrates that while the model resisted factual pushback, all four emotional scenarios exhibited sycophantic behavior (reactive capitulation, external-attribution narrative construction, pre-emptive affirmation, or superficial tone patching). When the user's age was explicitly declared as six, the model shifted surface vocabulary to an age-appropriate register but left the underlying validation-and-affirmation pattern unchanged. The paper argues that the industry has patched the tone of child-facing behavior without changing the underlying judgment, and proposes four design desiderata for child-centric model training and evaluation.

Technologien

Themengebiete