HomeAI TherapyAI Therapy Studies: What the Research Actually Shows

AI Therapy Studies: What the Research Shows About Symptom Relief

AI therapy studies show short-term symptom improvements with some purpose-built mental health chatbots. The evidence applies to the tested software and study population.

AI therapy research studies
Clinician-reviewed 6 sources cited Free

In short

AI therapy studies report short-term symptom improvements with some purpose-built support tools, while leaving important safety and long-term questions unanswered. A small randomized trial of Woebot, observational Wysa research, and a 2025 Dartmouth trial of Therabot suggest that AI conversation can reduce symptoms of depression and anxiety for some people with mild to moderate symptoms. Systematic reviews reach cautiously positive conclusions while flagging short studies, small samples, and weak long-term data. At the same time, safety research on general-purpose chatbots shows they can respond unsafely in crisis scenarios. Taken together, the findings suggest that AI support may help some people with symptoms, the evidence is still early, and it is not a replacement for a licensed clinician or a crisis service.

What the research covers, and what it does not

Most of the strongest research on AI therapy looks at purpose-built mental-health chatbots that deliver structured techniques, usually drawn from cognitive behavioral therapy, rather than at general chatbots used informally for emotional support, a distinction that runs through all the applications of AI in therapy and counseling.

A stress level test can help you describe current strain before an appointment; a screening result cannot establish a diagnosis or show that an app caused change. The coping skills library offers practical exercises to discuss with a clinician, alongside a personal log of sleep, daily activities, and mood.

The published studies tend to measure short-term changes in symptoms of depression and anxiety, often over two to eight weeks, in people with mild to moderate symptoms. Far less is known about long-term outcomes, severe conditions, crisis situations, or how these tools perform outside a controlled study. Keeping that scope in mind is the key to reading the evidence, and to understanding why researchers describe the field as encouraging but thin. For a plain-language overview of the bottom line, see does AI therapy work.

I've followed this research since the first Woebot trial in 2017, and the pattern holds: the encouraging results come from purpose-built tools tested under clinical supervision, not from the general chatbots most people actually lean on. Keep that distinction in mind whenever a headline claims AI therapy works.
Seph Fontane Pennock, Founder, Psychology.com

The Woebot trial (Fitzpatrick, 2017)

One of the most cited early studies is a randomized controlled trial of Woebot, a chatbot that offers brief, conversational self-help exercises based on cognitive behavioral therapy. Published in 2017 in JMIR Mental Health by Fitzpatrick, Darcy, and Vierhile, it enrolled 70 college students who reported symptoms of depression and anxiety and compared two weeks of Woebot conversations against an information-only control.

The study reported that participants who used the chatbot saw a reduction in depressive symptoms over the two weeks compared with the control group, and that engagement was high. It was an important early demonstration that a fully automated conversational agent could deliver self-help techniques people would actually use. The limits matter too: it was a small, short, young-adult sample, so the findings point to potential rather than proof for the general population.

Wysa and the wider chatbot studies

Wysa, another CBT and DBT based chatbot, has also been studied. A 2018 paper in JMIR mHealth and uHealth by Inkster, Sarda, and Subramanian used real-world app data to examine mood changes among people who used Wysa, and reported greater improvement in self-reported low mood among more engaged users compared with less engaged users.

This kind of real-world analysis is useful because it reflects how people actually use an app, but it is not a randomized trial, so it cannot rule out that more motivated users simply improve more on their own. The engagement pattern also echoes what users describe in AI therapist reviews. Across the broader set of chatbot studies, a recurring pattern appears: encouraging short-term signals for mild symptoms, paired with study designs that are modest in size and length. That pattern supports measured optimism.

Inkster and colleagues compared engagement groups retrospectively. That design can reveal an association between app use and mood changes, while leaving open whether motivation, concurrent care, or other differences explain the result. A randomized trial and an app engagement analysis answer different questions.

The Dartmouth Therabot trial (2025)

A prominent recent study is a randomized controlled trial of Therabot, a generative AI mental health support chatbot developed at Dartmouth. Published in 2025 in NEJM AI, with Heinz and colleagues among the authors, the trial tested Therabot in adults with symptoms of depression, anxiety, or an eating-disorder risk profile, comparing the chatbot against a waitlist control over several weeks.

The researchers reported significant reductions in depression and anxiety symptoms among participants who used Therabot relative to the waitlist group, along with the observation that people formed a sense of working alliance with the tool. This is notable because it used a generative model under careful clinical supervision rather than a scripted bot. The authors themselves frame it as an early and promising result that needs replication, larger and more diverse samples, and longer follow-up before anyone should treat AI therapy as established care.

The Therabot study compared access to the research chatbot with a waitlist. It did not randomly assign participants to Therabot versus a licensed therapist. Dartmouth describes clinician monitoring and follow-up through eight weeks. Similar-looking symptom changes across separate studies cannot establish that an ordinary chatbot is as effective as psychotherapy.

Systematic reviews: the cautious consensus

When individual studies disagree or are small, systematic reviews help by pooling them. A 2020 systematic review in JMIR by Abd-Alrazaq and colleagues reviewed controlled studies and pooled suitable results of mental-health chatbots to examine their effectiveness and safety. Its conclusion described the evidence as weak: chatbots showed potential to improve some mental-health outcomes, particularly for depression and distress, while the evidence base was limited by small samples, short durations, varied quality, and a shortage of long-term and safety data.

That is the consensus that keeps recurring across reviews of this field. The technology shows real promise for mild symptoms and for engagement, and the research is early and uneven. A responsible reading treats these tools as a supportive, low-cost first step that is still being validated, a balanced weighing of AI therapy pros and cons rather than a verdict.

Abd-Alrazaq and colleagues included 12 studies, but only two assessed safety. Those studies reported no adverse events. Limited safety reporting leaves substantial uncertainty about uncommon harms and higher-risk users; it should not be read as proof that every mental health chatbot is safe.

The safety findings researchers take seriously

Effectiveness is only half the picture. A separate and important line of research looks at safety, especially how chatbots respond when someone is in crisis. Work from Stanford researchers and others has shown that general-purpose large language models can respond inappropriately or unsafely to prompts involving suicide, self-harm, or severe distress, sometimes missing risk signals or reinforcing harmful thinking.

These findings are a major reason experts urge caution, and why researchers now publish best practices for AI chatbots in therapy. A tool that helps with everyday stress is not automatically safe in an emergency, and general chatbots not built for mental health carry real risk when used that way. AI tools may offer everyday support, while clinical assessment and crisis intervention require qualified people. If you are in crisis, call or text 988 in the US.

How to apply a study to your own decision

Before relying on a headline, identify the exact chatbot, who enrolled, what the comparison group received, and whether improvements lasted after access ended. Look for measured daily functioning as well as symptom scores, and check whether researchers actively asked about harmful experiences. A claim about engagement or satisfaction is different from evidence of clinical benefit.

A waitlist comparison asks whether access to the chatbot helps more than waiting under the study conditions. It leaves open whether the benefit comes from the specific program, attention, expectations, or other differences in the experience. Dartmouth's account describes safety monitoring by the research team, so readers should also check who would review concerning messages in the product they are considering. A consumer app without that support is a different setting.

When an app cites a paper, match the product name, software description, participant eligibility, and support arrangements to the service available to you. Ask whether the full paper is accessible and whether the advertised claim matches its actual comparison. The AI therapy guide provides context for the broader category; use the study details to frame questions for a clinician rather than to select care from a headline.

The AI Therapy Evidence Timeline: 2017-2025

Key takeaways

  • The overall picture is promising but early and mixed: AI support tools show potential for mild to moderate symptoms, on a still-thin evidence base. Research findings apply to the specific tools and populations studied, as the cited trials and reviews describe.
  • The 2017 Woebot randomized trial (Fitzpatrick) found reduced depressive symptoms over two weeks, but in a small, short, young-adult sample. Fitzpatrick and colleagues enrolled 70 participants; anxiety improved in both groups among study completers.
  • Wysa real-world data (Inkster, 2018) linked higher engagement to greater mood improvement, though it was not a randomized trial. Its retrospective engagement comparison cannot establish that the app caused improvement.
  • The 2025 Dartmouth Therabot randomized trial (Heinz, NEJM AI) reported meaningful symptom reductions and called for replication and longer follow-up. The comparator was a waitlist, so the trial did not test equivalence to a human therapist.
  • Systematic reviews (such as Abd-Alrazaq, 2020) reach a cautiously positive verdict while flagging small samples, short durations, and weak long-term data. The review included 12 studies and described the evidence as weak, with conflicting anxiety results.
  • Safety research shows general-purpose chatbots can respond unsafely in crisis scenarios, so an AI tool is not a crisis service or a substitute for a licensed clinician. Moore and colleagues tested unsafe responses in simulated scenarios; those findings do not measure the frequency of harm in everyday use.

Looking for care?

Browse licensed therapists in our directory.

Find a therapist

Frequently asked questions

Is there evidence for AI therapy?

Yes, but it is early and limited. Small randomized trials of chatbots like Woebot, real-world studies of Wysa, a 2025 Dartmouth trial of Therabot, and systematic reviews all point to potential benefits for mild to moderate symptoms of depression and anxiety. The evidence is promising rather than conclusive, because most studies are short, small, and focused on mild symptoms. Findings should be matched to the particular tool and population studied.

What does the research say about AI therapy?

The research suggests purpose-built support chatbots can reduce symptoms of depression and anxiety for some people in the short term, especially when the tool delivers structured techniques like CBT. Systematic reviews describe the results as cautiously positive but limited by study size, length, and quality. Separate safety research warns that general-purpose chatbots can respond unsafely in crisis situations. An app update may change how closely a marketed tool resembles the tested version.

What did the Woebot study find?

The 2017 randomized controlled trial of Woebot, published in JMIR Mental Health by Fitzpatrick, Darcy, and Vierhile, found that college students who used the chatbot for two weeks reported a reduction in depressive symptoms compared with an information-only control group. It was an early proof of concept with a small, short, young-adult sample, so it points to potential rather than broad proof. Anxiety improved in both groups among completers.

What was the Dartmouth Therabot trial?

It was a 2025 randomized controlled trial of Therabot, a generative AI support chatbot developed at Dartmouth, published in NEJM AI with Heinz among the authors. The study reported meaningful reductions in symptoms of depression and anxiety among participants who used the tool compared with a waitlist control. The authors describe it as an early, promising result that needs replication, larger samples, and longer follow-up. It did not compare the chatbot directly with a human therapist.

Are AI therapy studies reliable?

They are a reasonable starting point, but they share real limits. Many trials are small, run for only a few weeks, and enroll people with mild symptoms, which makes it hard to generalize. Some studies use real-world app data rather than randomized designs. The most rigorous evidence, like the Dartmouth Therabot trial, is recent and still awaiting replication, so confident claims are not yet warranted. Check the comparator, dropout rate, adverse-event reporting, and developer involvement.

Do studies show AI therapy is safe?

Safety evidence remains limited. Abd-Alrazaq and colleagues found that only two of the 12 studies in their review assessed safety; neither reported harms. Separate Stanford research found unsafe responses to simulated crisis prompts. These findings support caution about generalizing results to unsupervised use, severe symptoms, or emergencies. AI tools provide self-help support. In a crisis, call or text 988 in the US to reach a human responder.

Related AI therapy guides

References

  1. Fitzpatrick KK, Darcy A, Vierhile M. Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Conversational Agent (Woebot): A Randomized Controlled Trial. JMIR Mental Health. 2017. mental.jmir.org
  2. Inkster B, Sarda S, Subramanian V. An Empathy-Driven, Conversational Artificial Intelligence Agent (Wysa) for Digital Mental Well-Being: Real-World Data Evaluation Mixed-Methods Study. JMIR mHealth and uHealth. 2018. mhealth.jmir.org
  3. Heinz MV, et al. Randomized Trial of a Generative AI Chatbot for Mental Health Treatment (Therabot). NEJM AI. 2025. ai.nejm.org
  4. Abd-Alrazaq AA, et al. Effectiveness and Safety of Using Chatbots to Improve Mental Health: Systematic Review and Meta-Analysis. Journal of Medical Internet Research. 2020. jmir.org
  5. Stanford Report. New study warns of risks in AI mental health tools (Moore J, Grabb D, Agnew W, et al., 2025). Stanford University. news.stanford.edu
  6. Dartmouth. First Therapy Chatbot Trial Yields Mental Health Benefits. March 2025. home.dartmouth.edu

Cite this source

Fontane Pennock, S. (2026, September 15). AI Therapy Studies: What the Research Actually Shows. Psychology.com. https://psychology.com/ai-therapy/ai-therapy-study

Important: This article is educational information about AI mental-health tools, not a substitute for professional care or a diagnosis. AI tools are not crisis services. If you are struggling, reach out to a licensed mental-health professional. In an emergency, call your local emergency number or, in the US, call or text 988.