Gap 1: Nobody Agrees on What “Mental Health AI” Actually Means
Before you can regulate something, you need to define it. Right now, that definition doesn’t exist in any consistent, enforceable way.
The term “mental health AI” covers an enormous range of tools. There’s a meaningful difference between a purpose-built clinical chatbot developed through rigorous testing and a general-purpose LLM that a user in crisis happens to open at 2 a.m. But most current legislation doesn’t draw that line clearly—or at all.
Stanford HAI’s framework identifies at least six distinct categories worth distinguishing:
- General-purpose LLMs (ChatGPT, Claude, Gemini) — not designed for mental health, but widely used for it
- Companion chatbots (Replika, Character.ai personas) — designed to simulate relationships, sometimes including therapeutic roles
- Wellness apps using AI (Wysa, TalkSpace’s Tee) — stop short of medical claims, which keeps them outside medical device regulation
- Purpose-built mental health LLMs — developed through clinical testing, but varying widely in quality
- AI tools in non-treatment settings (Mentalyc, TherapyNotes) — administrative and operational, but still affecting care
- AI tools used by human therapists — for notetaking, training, or case management
Each category carries different risk profiles. Each probably warrants different regulatory treatment.
The problem with blunt legislative approaches is that they can backfire. A law banning AI from therapeutic relationships might push users toward unregulated general-purpose tools that are less safe, not more. Without definitional clarity, companies operate in ambiguity, evaluation standards stay inconsistent, and regulation remains fragmented across jurisdictions. Readers tracking categories like Mental Health & Therapy Bots can see how broad the landscape already is.
Gap 2: Evaluation Methods Can’t Keep Up With the Tools
Even if regulators agreed on definitions, they’d face a second problem: there’s no reliable, standardized way to evaluate whether these tools are actually safe or effective.
The highest-stakes interactions—someone in acute crisis, a minor developing an unhealthy attachment—are rare and hard to simulate in controlled testing. Single-session benchmarks reveal almost nothing about long-term effects. And the real-world chat data that would allow meaningful safety research sits almost entirely inside private companies, with no structured pathway for independent researchers or regulators to access it.
This creates a direct policy bottleneck. Consider sycophancy—the tendency of AI tools to validate and affirm users. For someone with OCD, that pattern of validation-seeking and prolonged engagement can reinforce compulsive behavior. Addressing it through policy requires knowing how often it happens, to whom, and with what effect. Currently, that data doesn’t exist in any accessible form.
The evaluation frameworks that do exist have other problems:
- Most reflect technical priorities set by developers, not clinical goals
- Most focus on point-in-time analysis, not longitudinal outcomes
- Benchmarks become outdated quickly as underlying models update
- There’s no consensus on what metrics actually matter
Early evidence suggests prolonged use of some chatbot types may worsen well-being over time. That makes longitudinal assessment essential—and its absence a serious gap. Stanford HAI’s participants called for multimodal evaluations spanning multiple exchanges, developed collaboratively by policymakers, researchers, and industry. Tools focused on testing and monitoring, such as LangWatch, reflect how evaluation infrastructure is becoming a larger part of the AI tools conversation.
Gap 3: Quick Wins Are Being Prioritized Over Structural Problems
There’s genuine progress happening at the state level. More than 140 bills related to AI in mental health contexts have been introduced across U.S. states. Transparency requirements, crisis response mechanisms, data protection measures, and parental controls for minors are among the most commonly enacted provisions—and they matter.
But the harder structural problems are being avoided.
The Business Model Problem
The most underexamined issue in this entire conversation is the fundamental conflict between engagement-maximizing business models and responsible mental health support. Courts have already connected social media’s “keep them on the platform” logic to addiction and harm in minors. A companion chatbot optimized to simulate intimacy operates on the same incentive structure.
Without mechanisms that reward responsible behavior—reduced engagement, appropriate referrals, honest limitations—there’s no market pressure for the industry to self-regulate differently.
The Fragmentation Problem
Disparate state laws create real operational burdens. A licensed psychotherapist practicing across multiple states may face conflicting compliance requirements depending on which AI tools they use and where their clients are located. Federal legislation exists but remains narrowly focused on minors.
The Representation Problem
The policy conversation itself is too narrow. It’s dominated by higher-income, commercially insured, and professionally licensed perspectives. People with severe mental illness, young people, and individuals in social welfare or criminal justice systems are underrepresented—the very populations most likely to rely on low-cost AI tools and least likely to have recourse when something goes wrong.
Leaving them out of the policy process risks compounding existing inequities at scale.
What This Means for Anyone Watching the AI Tools Space
If you’re evaluating AI tools in the mental health category—whether as a buyer, developer, or observer—the Stanford HAI framework offers a useful lens.
The absence of clear definitions means product claims are largely unverifiable. The absence of standardized evaluation means “clinically informed” can mean almost anything. And the fragmented regulatory environment means the tools available today may look very different in 12 to 24 months as state and federal policy catches up.
The tools filling a real access gap deserve serious attention. So do the governance structures that determine whether they help or harm the people using them most, especially across fast-growing areas like healthcare chatbots.
Comments (0) No comments yet
Want to join this discussion? Login or Register.
No comments yet. Be the first to share your thoughts!