Real Users, Simulated Struggles: Expert and LIWC-Based Evaluation of LLMs for Behavior Change Conversations.
- B. Chen and N. Dethlefs
Hide/Show Full Abstract
Large language models (LLMs) are increasingly applied to mental health and behavior change, yet their ability to support users in accordance with established psychological frameworks remains underexplored. This study evaluates GPT-3.5 Turbo, LLaMA-3.2, Qwen-2.5, and SmolLM-1.7B for Motivational Interviewing (MI)-grounded behavior change conversations using expert evaluations and linguistic analyses. Qwen-2.5 and GPT-3.5 Turbo showed the strongest performance in promoting engagement and motivation, although limitations in personalization, contextual understanding, and empathy remained evident. LIWC analysis showed that GPT-simulated users captured reflective reasoning and goal-oriented language but underrepresented key motivational struggles such as ambivalence and self-doubt. SHAP analysis further corroborated the LIWC and effect size findings, showing that GPT-simulated users exhibit more polished, confident, and solution-oriented language, whereas human users express more authentic and personally grounded experiences. These findings suggest that LLM-based simulations are useful for early-stage evaluation but remain limited in modeling authentic motivational processes.- 2026 Proceedings of SemDial (LuffDial), Loughborough.