AI-Generated Mental Health Advice Misjudged Due To Differences In Stateless Versus Contextual Evaluations
A research discussion on how AI evaluations can be misleading focuses on the difference between “offline” testing and real “contextual” use, arguing that many safety and utility benchmarks fail to reflect deployed behavior. The analysis cites a paper titled “The Inadequacy Of Offline Large Language Model Evaluations: A Need To Account For Personalization In Model Behavior” by Angelina Wang, Daniel E. Ho, and Sanmi Koyejo, published December 12, 2025. The paper explains that standard benchmarks typically ask models one question at a time with a stateless setup, labeling this as offline evaluation. It contrasts that with field evaluations using 800 real users interacting with ChatGPT and Gemini. The findings claim that identical prompts can produce different behaviors in offline versus field settings, which may cause safety evaluations to miss deployment risks. The discussion also notes that common testing methods compare restart/refresh responses, which can differ substantially across LLMs and dialogue conditions, including preset custom instructions.






