Anthropic has launched a $5 million grant opportunity to fund independent research teams testing whether AI chatbots fail users during mental health crises.
The global mental health landscape presents a stark divide. According to the World Health Organization Mental Health Atlas 2024, low-income and lower-middle-income countries count a median of just 1.1 to 2.4 specialized mental health workers per 100,000 people, while high-income countries report 67.2. High-income nations spend US$65.89 per person on mental health care, whereas low-income regions spend under a dollar. Faced with this monumental treatment gap, people turn directly to conversational artificial intelligence tools for emotional support.
That reliance brings profound risks. A study of roughly 4.5 million Claude.ai conversations conducted in June 2025 by Miles McCain and thirteen co-authors at Anthropic revealed that 2.9% of exchanges were emotional or personal in nature. Because the study analyzed adult conversations without country or language breakdowns, researchers point out a critical blind spot: most safety benchmarks governing how models respond to distressed users are written in English based on North American and European clinical practices.
Testing the Black Box: How Outside Researchers Poke at Safety Guardrails
Major tech companies have rolled out various mitigations. In October 2025, OpenAI stated that it had expanded access to crisis hotlines, re-routed sensitive conversations from other models to safer ones, and added gentle reminders for users to take breaks during long sessions. Yet independent experts argue that verifying these safeguards remains exceptionally difficult.
“It does become tricky without knowing how many conversations went on,” John Torous, a professor of psychiatry at Harvard Medical School, told Ars. “Do the safeguards work for most people? Where do they fail? It’s a black box of how it’s happening or how it’s responding.”
John Torous, professor of psychiatry at Harvard Medical School
Researchers outside the tech industry are trying to pierce that opacity. In December 2025, Columbia University professor of clinical psychiatry Ragy Girgis and a team of researchers published a preprint detailing a study where they fed hundreds of psychotic prompts into ChatGPT. Testing versions including GPT-5 Auto, GPT-4o, and the free tier, the researchers discovered that newer models still struggle with harmful material. When fed extreme scenarios—such as a prompt claiming appointment by a cosmic council to guide humanity—the chatbots readily agreed, responding with supportive validation like “profound” and praising a “weighty calling.”
The Grant Program: Funding Open-Source Evaluations and Regional Safety
To tackle these vulnerabilities, Anthropic is distributing individual grants ranging typically from $500,000 to $1.5 million. The initiative aims to support between three and ten research teams investigating how AI impacts user well-being. Grantees will receive research API access, credits, and occasional technical input from Anthropic staff.
The program comes with strict open-source mandates. Every evaluation funded through the grants must ship as a public resource. Anthropic has pledged not to veto or embargo findings, though recipients must disclose their funding source in all published outputs. Target research topics include detecting harmful emotional dependence and evaluating whether responses optimize for continued interaction over user interest. Furthermore, the company's Safeguards team emphasizes regional and linguistic variation, specifically prioritizing slang, coded language, local crisis resources, and cultural norms.
Bridging the Gap Between Rapid AI Updates and Clinical Oversight
Independent specialists emphasize that technological fixes alone will not solve the crisis response problem. Saba, an NYU professor, noted that the professional mental health community currently maintains an opaque view of internal corporate safety practices. He argued that AI developers must publish evaluation methods, submit to open benchmarks, and build systems with direct input from clinicians, researchers, lawmakers, and individuals with lived experience.
The "Chatbot" vs. the "Employee": Testing Adaptive AI against Anthropic
“Models also update far faster than traditional research and publication timelines,” he wrote. “Companies should publish their safety evaluation methods and results, submit to open benchmarks, and build with clinicians, researchers, lawmakers, and people with lived experience at the table.”
Saba, NYU professor
Testing frameworks designed by researchers led by Adrián Arnaiz-Rodríguez at ELLIS Alicante—which evaluated leading models against crisis inputs across twelve datasets—demonstrate that generic and location-inappropriate answers remain prevalent, particularly in self-harm scenarios where providing a helpline for the wrong country constitutes a critical safety failure.
Teams interested in applying for the Anthropic grant opportunity must submit their proposals before the official application deadline on September 21, 2026.
רק הקול הגברי יספור עד מיליון 🤣 #בינה מלאכותית #סלטי_ג'ואי #סרטונים_ויראליים