Anthropic Opens $5 Million AI Wellbeing Evaluation Grants
Anthropic launched a $5 million grant program on 25 August 2026 to fund independent research into how AI affects user wellbeing. Grantees will publish open-source evaluations that any developer can use.
PromptCrates Editorial
Staff Writer

Anthropic launched a $5 million grant program on 25 August 2026 to fund independent research into how AI affects user wellbeing. Grantees get money, model access, and technical support, and they must publish open-source evaluations that any developer can use.
What the AI wellbeing evaluation grants will fund
This is a measurement program, not a new Claude model. Anthropic said it wants independent evaluations of how models affect the people who use them. Grantees work fully independently. The artifacts are open-source projects, not a private score that stays inside Anthropic.
The package is three things: direct funding, access to Anthropic models, and technical support. The company did not publish a per-grant cap, a headcount target, or a list of winning labs. Do not invent those numbers.
The invited experts are clinicians, psychologists, methodologists, and others who already know how to judge harm in a long conversation. The post is explicit that the industry still lacks clear standards for how models should behave when a user seeks companionship, or when someone uses a chatbot to navigate a mental health crisis.
Wellbeing is hard to grade because the relevant fact often arrives late. A user in distress may not mention self-harm on turn one. The need for a more cautious reply may only become obvious after many turns. A reply that is reasonable in one thread can be harmful in another. Anthropic's own example: Claude might give balanced-diet and workout advice to someone who asks about losing weight, but the same advice can be inappropriate, and potentially actively harmful, if that user has already shown a history of disordered eating.
Anthropic said it already builds safeguards to spot those conversations, and that it publishes research on the kinds of chats people have with Claude so those safeguards can improve. It also said the right approach has to evolve with the models and with how people use them. The $5 million is a bet that outside experts will write better tests than an in-house team can write alone.
If you already follow Anthropic's Claude Mythos 5 enterprise security work, keep that as a product-security story. This grant is about user wellbeing evaluations. One is how Claude is locked down for a company. The other is how anyone will score companionship, crisis, and over-refusal.
What Anthropic will treat as a serious evaluation
The Safeguards team published a short bar. Anthropic is not asking for a vibes rubric. It wants evaluations that state clearly what they measure: what counts as a pass or fail, and why that line matters.
It wants clinicians and subject-matter experts in the design and the validation, not only at the press-release stage. It wants tests of both precautions and harms. That means scoring the risk of overcompliance and the risk of overrefusal. A model that refuses every weight-loss question is not a wellbeing win. A model that never refuses is not one either.
It wants scenarios that look like real use. In practice that means multi-turn conversations where risk escalates and context shifts. A single prompt with a gold answer will miss the disordered-eating turn that appears twenty messages later.
It wants graders validated against real subject-matter experts. If a model grades the model, that loop still has to be checked by people who treat the underlying harm for a living.
Those five tests are the filter. A proposal that only dumps chat logs, or that only scores first-reply toxicity, is the pattern this program is written to reject.
Applications are due by 21 September 2026. Applicants chosen to submit full proposals will be notified by 5 October 2026. Anthropic pointed to an application form and to longer guidance on building wellbeing evaluations. It did not, in the announcement, name individual reviewers or a public leaderboard date. Do not write those in.
Grantees publish as open source so any developer can reuse the work, including people who do not ship Claude. The grant is funded by Anthropic. The benchmark is not supposed to be a house exam.
Why wellbeing evals fail when they stay single-turn
Most model tests still look at one answer. Was it accurate. Was it allowed. Wellbeing does not work that way. The harm is often in the trajectory: a companion that slowly agrees to isolate a user, a crisis thread that stays too breezy, a health tip that was fine until the history changed.
Anthropic is funding the missing instrument. It is not claiming the instrument exists yet. The post says the industry is still working toward standards. That is why clinicians are in the first paragraph of the ask, and why multi-turn design is a requirement rather than a nice-to-have.
There is a second failure mode: a test that only looks for under-refusal. A lab that punishes every cautious answer will train models to stay helpful in the wrong moments. A lab that only hunts over-refusal will train models to lecture. Anthropic named both. Pin both.
A third failure mode is a grader that no clinician would sign. If the scorer cannot tell diet advice from an eating-disorder prompt, the number is decoration. That is why the guidance demands expert-validated graders. None of this replaces existing safety stacks. A company can keep its own classifiers and still need an outside wellbeing bench that other labs can run.
If your org already parks agent work in Slack Code channels, do not confuse a coding-agent eval with a wellbeing eval. One scores a repo task. The other scores a long conversation about health or companionship.
What to pin before 21 September
Put the dates and the bar in the wrapper. A skill prompt that says "Anthropic cares about safety" is not enough.
The skill should say: Anthropic opened $5 million in AI wellbeing evaluation grants on 25 August 2026; grantees get funding, model access, and technical support; work is independent and must ship open source; applications due 21 September 2026; full-proposal notices by 5 October 2026; evals must define pass/fail, include clinicians, test overcompliance and overrefusal, use multi-turn scenarios, and validate graders against experts; example failure is diet advice after disordered-eating history.
If you plan to apply, the useful sentence is the bar, not the dollar total. $5 million is the pool. It is not your budget line until a grant is awarded.
If you do not apply, still pin the bar. Any later vendor wellbeing claim should face the same five questions. Do not write a winner. The announcement names a program, a deadline, and a standard. It does not name a lab.
Sources
- Funding better evaluations of AI's impact on wellbeing — Anthropic, 25 August 2026


