In short
Some studied mental-health chatbots have reduced depression or anxiety symptoms over short follow-up periods. Results depend on the specific tool, participants, comparison group, and outcome measured. Evidence for a research chatbot cannot establish that a general assistant works equally well. Treat consumer AI as self-help support and involve a professional for clinical care. Early structured-app trials and newer generative-AI trials need separate interpretation, including how researchers monitored safety and whether benefits lasted after use ended.
The short answer: it works in a narrow way
Some mental-health chatbots work in a limited, specific sense: controlled studies have reported improvements in symptoms using particular tools. Early evidence includes structured cognitive behavioral self-help, or CBT, for low mood and anxiety. Newer generative-AI research also reports benefits, so describing every effect as small or confined to mild symptoms is too broad. Results still cannot establish that every app helps every problem.
The stress level test can help you describe current strain, and a mood tracker can record changes in daily functioning alongside any app use. These are self-reflection aids. A screening score or an improving chart cannot show that a chatbot caused the change or establish a diagnosis.
Three caveats matter from the start. First, most of the positive evidence comes from a small number of named apps that were built and tested by research teams, not from the wider field of companion or wellness chatbots. Second, the trials are short, such as the brief follow-up in the Fitzpatrick trial, so we know little about whether benefits last. Third, a large share of apps on the market have never been studied at all, which means their marketing claims are unverified.
If you are in crisis or thinking about suicide, do not rely on an app. In the US, call or text 988 to reach the Suicide and Crisis Lifeline, available 24 hours a day. AI therapy tools are not designed for emergencies and should never be used as a crisis service.
In my own testing, the purpose-built tools that follow a structured method helped me more than a general chatbot that simply agreed with me. The moment an AI starts telling you what you want to hear, it has stopped helping and started flattering you. That lines up with the research: the boring, structured apps are the ones with actual trials behind them.
What the research actually shows
A handful of peer-reviewed studies anchor the case that AI therapy can help. In a 2017 randomized trial published in JMIR Mental Health, Fitzpatrick, Darcy, and Vierhile tested Woebot, a CBT-based chatbot, against an information-only control in young adults. The chatbot group showed a meaningful reduction in symptoms of depression over two weeks. In 2018, in a real-world evaluation rather than a controlled trial, Inkster and colleagues studied Wysa, another CBT-informed chatbot, and reported improvements in self-reported mood among more engaged users, published in JMIR mHealth and uHealth. The Wysa evaluation compared more-engaged and less-engaged users without random assignment. Differences between those groups could reflect motivation or other circumstances, so the association does not establish that app use caused improvement.
Reviews that pool many studies reach a cautious conclusion. A 2020 systematic review by Abd-Alrazaq and colleagues in JMIR looked at chatbots for mental health and found promising but limited evidence: some trials showed reductions in distress, depression, and anxiety, but study quality varied, sample sizes were often small, and follow-up was short. The overall signal is positive but weak, and the authors were clear that more rigorous, longer trials are needed.
Effect sizes vary by study, symptom, and comparator. Early trials include subclinical symptoms; newer trials also include clinically significant symptoms. And the results that exist belong to specific, named apps tested under research conditions, which is not the same as the experience of downloading a random chatbot from an app store. Woebot's consumer app was discontinued in mid-2025, which underlines how young and unstable this market is: an extensively studied app can become unavailable. Our walkthrough of the major AI therapy studies covers each trial in more detail.
The newest wave of tools, built on large language models, now has its own review literature. Scoping reviews by Hua and colleagues that mapped the research literature describe potential for support tasks such as psychoeducation and structured exercises, alongside a consistent set of problems: study designs vary widely, evaluation is rarely standardized, safety and fairness remain underexplored, and none of the evidence supports using these tools as standalone treatment.
How AI therapy works mechanistically
Many studied mental-health chatbots deliver established self-help techniques through a conversational interface. The chatbot guides you through structured CBT exercises: noticing an unhelpful thought, examining the evidence for and against it, reframing it, and practicing a small behavioral step. The conversational format makes those exercises feel lighter and more personal than reading a worksheet.
Several ordinary mechanisms plausibly explain the benefit. The app prompts you to name and rate your feelings, which is a form of emotional labeling that can lower their intensity. It nudges you to practice skills regularly, and consistency is much of what makes CBT work. It may be available outside appointment hours, creating opportunities to practice between scheduled sessions. And the simple act of writing out a worry, even to software, can create useful distance from it.
Newer apps built on large language models can hold a more fluent, free-flowing conversation, which feels more natural. But fluency is not the same as clinical effectiveness, and trials of a particular system cannot isolate conversational smoothness as the reason for improvement. It can even introduce risks, such as confidently wrong advice or responses that miss signs of serious distress. Evidence should match the specific system and use case, including structured and generative tools. Fluency also cannot supply the accountable care relationship in human therapy. A 2018 meta-analysis of 295 studies and more than 30,000 patients found the therapeutic alliance to be one of the most robust predictors of outcome, and an app can imitate the language of that bond without being able to enter into it.
Who it tends to help, and who it does not
AI self-help may suit people with mild to moderate symptoms who want to build coping skills, track their mood, or have support available between or before formal care. It can suit someone on a waitlist, someone who wants to practice CBT techniques between sessions, or someone testing the waters before committing to a human therapist. People who engage consistently appear to get more out of these tools than people who open them once and drift away.
It is a poor fit, and sometimes an unsafe one, for serious or complex conditions. That includes active suicidal thoughts, psychosis, severe depression, trauma that needs specialized treatment, eating disorders, and substance use disorders. In those situations an app is not enough, and relying on one can delay real care. Consumer AI support also falls short when a problem needs diagnosis, medication, or the judgment and relationship that a trained clinician provides.
There is also wide individual variation. Some people find a chatbot supportive and motivating. Others find it repetitive, scripted, or hollow, and disengage quickly. Comfort with technology, the nature of the problem, and personal preference all shape whether a given tool helps a given person.
Is AI therapy legit, or just marketing?
Both, depending on the app. The category is legitimate in the sense that real research supports specific, structured tools for specific uses and populations. It is also crowded with products that borrow the language of therapy without any of the evidence. A chatbot that calls itself an AI therapist is not regulated the way a licensed clinician is, and most are not cleared as medical devices.
A few honest signals separate evidence-based AI therapy from marketing. Look for a named clinical approach such as CBT or DBT rather than vague talk of wellness. Look for published, independent studies of that actual app, not a single company-funded survey. Look for clear, prominent crisis guidance and pointers to resources like 988. And look for a transparent privacy policy, because these tools collect sensitive emotional data and how they handle it matters. Apps that follow best practices for AI chatbots in therapy make these signals easy to find. Read AI therapist reviews with the same skepticism you bring to marketing copy.
When those signals are missing, treat the claims with skepticism. An app that promises to replace your therapist, cure anxiety, or diagnose your condition is overpromising. The credible tools describe themselves more modestly, as self-help aids that support, not substitute for, professional care.
How to set realistic expectations
Approach consumer AI as an on-demand self-help aid for noticing everyday stress and practicing coping exercises. Used that way, with realistic expectations, it can be a reasonable first step or a useful supplement. Expect modest help with noticing thoughts, building a habit, and tracking how you feel, rather than a transformation or a cure.
Give it a fair but bounded trial. Pick a tool with a recognized therapeutic approach and some published evidence, use it consistently for a few weeks, and check whether your mood or coping actually improves. A practical guide on how to use AI as a therapist can help you structure that trial. If it helps, keep using it as a supplement. If it does not, or if your symptoms are worsening, that is a signal to seek a human professional rather than to keep trying apps.
Keep the limits in view at all times. Consumer AI tools cannot provide a clinical diagnosis or take responsibility for treatment decisions, ongoing care, or crisis response. If you are in danger or thinking about suicide, call or text 988 in the US. If you want to understand the safety trade-offs more deeply, read about whether AI therapy is safe, and if you would rather work with a person, browse licensed therapists in our directory.
How to read newer generative-AI trial results
The Heinz and colleagues Therabot trial in NEJM AI randomly assigned 106 participants to the chatbot and 104 to a waitlist, with an intervention lasting 4 weeks. It reported greater symptom improvement with the studied system. The trial tested a purpose-built research chatbot; the result cannot establish effectiveness for ChatGPT, Claude, or another consumer app.
A waitlist comparison asks whether offering the studied system helps more than waiting under those trial conditions. It cannot establish equivalence to sessions with a licensed therapist. Research oversight and safety procedures also matter: downloading an app does not reproduce a monitored study environment.
Before relying on an effectiveness claim, check whether participants resemble you, which symptoms were measured, who dropped out, and whether harms were actively recorded. Look for follow-up after the intervention ends and replication by independent researchers. An improvement in a questionnaire score and a lasting change in everyday functioning answer different questions.
Key takeaways
- Some studied chatbots improve symptoms over short follow-up periods. Effects vary by tool, condition, and comparison group; a single size-of-benefit claim cannot describe the whole field.
- The evidence comes mostly from a few specific, studied apps such as Woebot and Wysa, not the whole market, and many popular apps have no published research.
- Short follow-up limits what trials can establish about lasting benefit. Fitzpatrick and colleagues tested a brief intervention, while newer studies require their own assessment of duration and safety.
- Studied chatbot approaches include CBT exercises, emotional labeling, and repeated practice. These features can support self-help while professional care supplies clinical judgment and accountability.
- It helps people with mild to moderate symptoms who engage consistently, and is a poor or unsafe fit for crisis, severe, or complex conditions.
- Consumer AI cannot provide a clinical diagnosis or take responsibility for treatment, ongoing care, or crisis response. In the US, call or text 988 for a mental-health crisis; use emergency services for immediate physical danger.
- Fact: Fitzpatrick and colleagues randomized 70 young adults to Woebot or an information-only control for 2 weeks. Depression improved more in the chatbot group; anxiety improved in both groups among completers. Source: JMIR Mental Health, the Woebot randomized trial.
- Fact: The Abd-Alrazaq review included 12 studies, with safety assessed in only 2. Absence of reported harm in those studies cannot establish broad safety. Source: JMIR, Effectiveness and Safety of Using Chatbots to Improve Mental Health.
- Fact: The Therabot trial compared a 4-week chatbot intervention with a waitlist, using 106 and 104 participants respectively. It did not directly compare the chatbot with a licensed therapist. Source: Heinz and colleagues, NEJM AI.
Want evidence-based care?
Browse licensed therapists in our directory.
Frequently asked questions
Does AI therapy actually work?
Some studied mental-health chatbots improve symptoms under research conditions. The Fitzpatrick Woebot trial found greater improvement in depression than an information-only control, while the cited Wysa evaluation was observational. Newer Therabot research compared a purpose-built system with a waitlist. These findings support investigating specific tools; they cannot establish that every consumer app works or that a chatbot replaces professional care.
How effective is AI therapy?
Effectiveness varies by app, symptoms, comparison group, and follow-up. Early chatbot reviews report limited and mixed evidence, while newer purpose-built systems have also shown symptom improvements in randomized trials. A favorable result against a waitlist cannot establish equivalence to a licensed therapist. Check the actual study and its safety procedures before applying its conclusions to your own situation.
Is AI therapy legit?
It is legitimate for a narrow use. Specific, evidence-based tools built around CBT have real research behind them for milder symptoms. The category is also full of apps that use therapy language without any evidence and are not regulated as medical devices. Look for a named clinical approach, published independent studies of that app, clear crisis guidance, and a transparent privacy policy. Verify that the current app version and intended users match the research being advertised.
Is there evidence-based AI therapy?
Published research supports specific chatbot systems and uses. The cited Woebot study was randomized; the cited Wysa evaluation compared engagement groups without random assignment. The Therabot trial adds evidence for a purpose-built generative system under research conditions. Ask whether the exact app has controlled studies, independent replication, appropriate safety monitoring, and follow-up beyond the period of use.
How does AI assisted therapy work?
Most evidence-based tools offer CBT-informed self-help through a conversation. The chatbot guides you to notice an unhelpful thought, weigh the evidence, reframe it, and try a small behavioral step. It also prompts you to label and rate emotions, encourages regular practice, and is available any time. Newer language-model apps generate more flexible responses; trials must evaluate the complete system, including its safety procedures.
Can AI therapy replace a real therapist?
No. Consumer AI can provide self-help and support. It does not diagnose, treat, or cure mental-health conditions and is not a crisis service. It can complement professional care or serve as a low-cost starting point for mild symptoms, but it cannot replace the judgment, relationship, and accountability of a licensed clinician. In a crisis, call or text 988 in the US.
Related AI therapy guides
References
- https://mental.jmir.org/2017/2/e19/ mental.jmir.org
- https://mhealth.jmir.org/2018/11/e12106/ mhealth.jmir.org
- https://www.jmir.org/2020/7/e16021/ jmir.org
- https://doi.org/10.1037/pst0000172 doi.org
- https://arxiv.org/abs/2401.02984 arxiv.org
- https://arxiv.org/abs/2408.11288 arxiv.org
- https://doi.org/10.1056/AIoa2400802 doi.org
- https://pmc.ncbi.nlm.nih.gov/articles/PMC12955234/ pmc.ncbi.nlm.nih.gov
- https://www.samhsa.gov/mental-health/988 samhsa.gov
Cite this source
Fontane Pennock, S. (2026, September 15). Does AI Therapy Work? What the Evidence Says. Psychology.com. https://psychology.com/ai-therapy/does-ai-therapy-work
