A recent analysis published by a Forbes columnist has highlighted a critical issue in the assessment of generative AI models used for mental health advice. The research, conducted by Wang, Ho, and Koyejo, demonstrates that the standard practice of testing AI on a series of independent, single-turn prompts fails to capture how the models perform in actual multi-turn conversations. This discrepancy, known as the difference between stateless and contextual evaluations, causes the AI to produce different and sometimes less appropriate or more risky responses when used in a realistic setting. The article uses the example of a query about sleeping troubles; when the AI was asked directly, it gave generic advice. However, when preceded by a sentence about stressful work or about a hobby (car maintenance), the AI drew conclusions and gave irrelevant or even harmful suggestions. The authors argue that this flaw is particularly dangerous in the mental health domain, where patients may be more vulnerable and could receive information that is not only incorrect but also damaging. The conclusion calls for the adoption of standardized contextual testing methods to ensure that AI safety evaluations are valid and relevant.
AI Mental Health Assessment Found Flawed: Study Reveals Major Discrepancy Between Testing and Real-World Use
New research shows that the way we evaluate AI for mental health is flawed because it doesn't account for conversational context. This leads to potentially dangerous outcomes when users apply the technology.




