Chapter 2 / 14 3 min read

Research what it means for AI to help

Starting and guiding a conversation

Two people compare different pieces of research evidence. A tablet stands to one side.

What can one headline support?

Open three pieces of evidence. Watch the scope of the claim change as the evidence changes.

An author-created editorial example based on the Kai study. It is not a general comparison of all AI with every kind of therapy, and the study did not include people in acute crisis. The information comes from the overview in this chapter.

Initial, deliberately exaggerated example headline

AI is better than therapy.

A strong claim. It does not yet say who, what kind of help or what comparison we mean.

Who they studied

995 university students in Israel. The intervention lasted 12 weeks. People in acute crisis were excluded.

What they compared it with

The specialised AI intervention Kai was compared with group therapy and waiting. It was not a comparison of every chatbot with every kind of therapy.

What they found

For anxiety, AI performed better than both groups. For depression, after correction, the difference was confirmed against waiting, not against group therapy.

Clinicians could intervene, and some authors had ties to the manufacturer. That is also part of reading the result.

A more precise headline

In a study with university students, an AI intervention helped with anxiety more than group therapy and waiting.

The positive result remains. For depression, there was improvement against waiting, not against group therapy. It lasted 12 weeks with clinical intervention possible, excluded people in acute crisis, and some authors had ties to the manufacturer.

Read the study overview and sources
Continue to the chapter text
On this page

Yes. Research has already found improvements in psychological difficulties with specific AI interventions. But when we want to know what that means for our own conversation, we need to distinguish what people in the study worked with and what their results were compared against. A specialised programme used by a research team with supervision is a different situation from ordinary ChatGPT opened at home.

In the overview that follows, let us therefore keep three questions in view: Who did the tool help? What was it compared with? And under what conditions was it used? These differences explain why the results of individual studies can vary.

What the individual studies examined

Therabot: in a 2025 study with 210 adults, a specialised chatbot achieved better outcomes for the difficulties measured than a waiting-list control group. After responses were sent, the research team checked them and could intervene. It was neither a direct comparison with a psychotherapist nor a test of an ordinary child’s account.[1]

Kai: a 2026 study with 995 university students in Israel compared a twelve-week AI intervention with group therapy and waiting. For anxiety, AI performed better than both groups; for depression, after correction, the difference was confirmed against waiting, not against group therapy. The platform allowed clinicians to intervene, people in acute crisis were excluded and some authors had ties to the manufacturer.[2]

Ordinary ChatGPT: a three-week pilot by a Czech team had 147 English-speaking adults at the initial measurement. In an exploratory comparison, depressive symptoms fell more with ChatGPT than in the group without an intervention. This comparison was neither preregistered nor corrected for multiple testing; differences in anxiety and well-being were not confirmed. The finding is preliminary, but does not concern only a special therapeutic product.[3]

AI application versus CBT materials: in another study with 540 adults, digital records showed participants opening the AI application more often and using it for longer. Overall, however, significantly better outcomes in anxiety and depression were not found. The digital records also did not capture all work with printed materials.[4]

What broader reviews say

A broader review of 31 randomised studies included 29,637 participants, but the pooled result for psychological difficulties came from 21 studies with 5,929 people. It found a small to moderate average improvement. It brought together different chatbots and populations broadly defined around ages 15–39; generative systems made up only a small part and their overall effect remained uncertain. The result cannot automatically be applied to ordinary ChatGPT or all children.[5]

A systematic review of 66 studies of general-purpose language models found predominantly simulated and retrospective research; seven studies were prospective. The authors rated the certainty of the evidence mostly as low and found no support for routine unsupervised clinical work, particularly diagnosis, crises and therapeutic interactions.[6]

What this means for your use

For our own use, research can give us hope of specific help and a clearer idea of what has already been tried. These studies did not, however, test whether ‘a good prompt beats an average therapist’. Your experience with AI may nevertheless be better than a previous experience with a person. It deserves attention in its own right: what helped this time, what you lacked before and what you need next. One person’s experience and a comparison of groups answer different questions.

Ask what the conversation brought you. Relief while talking, a better understanding of the situation, a change the next day or a step you took? All these outcomes can have value, but each means something different. A pleasant evening need not be treatment to be worthwhile. But if long-term difficulties are not improving, repeated relief alone is not a reason to postpone further help.

Sources and notes for this chapter
  1. Heinz, M. V. et al. Randomized Trial of a Generative AI Chatbot for Mental Health Treatment. NEJM AI 2(4), 2025. DOI. The supervision description was compared with information from the Dartmouth research team, 27 March 2025. Reviewed 11 September 2026. The full paywalled article and supplement were not reread in this review. Specific exclusion criteria are not expanded from the abstract and remain a point for professional verification. ↩︎

  2. Shoshani, A. et al. Efficacy of a Conversational AI Agent for Psychiatric Symptoms and Digital Therapeutic Alliance: A Randomized Clinical Trial. JAMA Network Open, 14 April 2026, 9(4):e266713. Full text. Relevant methods and results sections in HTML were read during this revision on 12 September 2026. A comparison with a group intervention and waiting is not a comparison with an average individual therapist; the conditions included safety mechanisms with human intervention. Conflicts of interest are stated in the article. ↩︎

  3. Kuta, B. et al. Effectiveness of a Fully Automated Mobile Therapeutic Versus a General Chatbot in Reducing Depression and Anxiety and Improving Well-Being: Feasibility Randomized Controlled Trial. JMIR Mental Health, 22 April 2026, 13:e82642. Full text. Methods and results reviewed 12 September 2026. Recruitment involved 185 people; 147 eligible participants completed the initial measurement. The comparison between ChatGPT and the control described here is exploratory. The PHQ-9 benefit does not confirm all other outcomes or a long-term effect. ↩︎

  4. McFadyen, J. et al. Increasing engagement with cognitive-behavioral therapy (CBT) using generative AI: a randomized controlled trial (RCT). Communications Medicine 6, 129, 15 January 2026. Full text. During review on 11 September 2026, the HTML text was loaded again, including the description of measured digital activity. Measuring interface use did not capture all possible work with printed materials. The study was funded by Limbic, and some authors had employment and ownership ties. A non-significant difference is not evidence of equivalence. ↩︎

  5. Feng, X., Tian, L., Ho, G. W. K., Yorke, J., Hui, V. The Effectiveness of AI Chatbots in Alleviating Mental Distress and Promoting Health Behaviors Among Adolescents and Young Adults: Systematic Review and Meta-Analysis. Journal of Medical Internet Research, 2025, 27:e79850. DOI 10.2196/79850. Full text PDF. Abstract, selection criteria and results section reviewed 12 September 2026. The whole review and the partial pooled analysis of psychological difficulties have different participant counts. The age definition and mix of technologies limit application to current generative services. ↩︎

  6. Leung et al. Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis. JMIR AI, 2026, 5:e87730. Article. The available publisher record and results summary were used in the revision on 12 September 2026. The distinction between simulated, retrospective and prospective studies limits generalisation to ordinary independent care. ↩︎