Researchers write a position paper outlining a roadmap for evaluating large language models for mental health. Will the AI companies listen? A conversation with Charlotte Blease, researcher and author of “Dr Bot: Why Doctors Can Fail Us and How AI Could Save Lives,” on how AI is a mass population experiment.
Transcript
Charlotte Blease: What we’ve got on our hands, whether we like it or not, is a mass population experiment. We don’t know if people are, in some circumstances, faring better or doing a lot worse as a result of these tools. We just simply don’t have the data on it.
Stephanie Hepburn: This is CrisisTalk. I’m your host, Stephanie Hepburn. Today I’m speaking with researcher Charlotte Blease. She’s the co-author of a recent position paper published in The Lancet that outlines a roadmap for evaluating and implementing AI and mental health. We discuss not only the paper, but also how AI has become a shadow healthcare system and what happens now. Let’s jump in.
Charlotte Blease: My name is Charlotte Blease. I’m an associate professor of health informatics at Uppsala University in Sweden. And I’m a research affiliate at Beth Israel Deaconess Medical Center in Boston. And I’m also the author of a book called Dr. Bot, Why Doctors Can Fail Us and How AI Could Save Lives.
Stephanie Hepburn: So, Charlotte, we’re we’re talking today because you are the author of a paper that was published in The Lancet. It is a position paper titled Multidisciplinary Research Priorities for Artificial Intelligence and Mental Health: A Call to Action. What was the impetus for writing this paper with your co-authors?
Charlotte Blease: Yeah, so this paper was actually led by my colleague Jake Linarden, who’s in Australia, together with Joseph Firth and John Torres and myself, but together with a bunch of other people. And the idea was really to launch some research priorities in relation to AI and mental health care. We recognize that to tackle some of the urgencies in relation to patients using these tools in mental health care, just the incursions of AI within mental health care, that we need to set out an agenda. And that basically falls across four major priorities: getting better evidence, putting patients’ ethics and equity at the center of this agenda, working out thirdly, what happens to clinicians, sort of the evolving professional and the expertise needed to use these tools. And finally, actually getting evidence about what works within healthcare. Because if these tools don’t integrate successfully, they don’t fit workflows or they add additional time, they’re not going to be particularly successful potentially. In a nutshell, we’re basically laying down what we think needs to happen next. So it’s a very broad paper, but it’s a call to action.
Stephanie Hepburn: Is this specific to generative LLM models as opposed to rule-based systems? Is it specific to those that are very much tailored to mental health? Or does it also apply to the broad use, large language models like Claude and Gemini and ChatGPT?
Charlotte Blease: The urge was really with large language models. And there’s a recognition that consumer chatbots are increasingly used, both by clinicians and by patients, really the population at large. I mean, I work in Sweden. The Swedish Prime Minister last year said that he consulted ChatGPT for policymaking. So everybody is using these tools and they’re often using them in a surreptitious manner. And we need more research. Preface that by saying, John Torres and myself, we did the very first survey of clinicians’ use of these tools. And we did this in 2023. So it was almost one year after ChatGPT had been unleashed, if you like. It was a small convenient sample survey, but it was American Psychiatric Association physicians. And we found that about a third of them said that they were using Chat GPT. So that was sort of a clarion call, a very early clarion call. And since then, I mean I’ve conducted other studies, the UK, Sweden, recently a global psychotherapy study. It’s a preprint, it’s not yet been peer reviewed, but we conducted it earlier this year, between January and March this year, 2026, across 30 countries of mental health clinicians, predominantly clinical psychologists, psychotherapists. 55% told us they were using generative AI tools to assist with psychotherapy practice. And of them, 85% said they used ChatGPT. So we have a big issue on our hands with the adoption of these tools. Incidentally, most of those people who used the tools, eight and ten, said they’d have no training. And they weren’t just using them for documentation or administrative tasks. And documentation, incidentally, of course, can be problematic if you’re giving away sensitive information, but they were using them for treatment recommendations for patients and for generating psychoeducational material for patients and also for their own reflective practice as well. So really core clinical areas. And then on the patient side, we’re seeing an uptick in adoption might not yet be a full tipping point among the public, but we’re certainly seeing people turning to these tools. They’re readily available, ease of access, lack of stigmatization. People fear stigmatization, particularly in mental health contexts, means that people are driven to sort of pour their hearts out to these tools. What we’re trying to say is we need to know in much more fine-grained detail what are the short-term, the long-term consequences of adopting these tools? How can we best have clinicians using these tools? And how do we make these tools work? Because they’re not going to go away.
Stephanie Hepburn: This umbrella that you mentioned is rather large because you’re mentioning consumer-facing mental health apps and chatbots, there are electronic health records that are integrating AI, but you’re also mentioning just how broad use, general use AI chatbots are being used, both by clinicians and by consumers. Can one umbrella really encompass the standardization and regulatory aspect, or does that need to be tiered out? It’s great to have an idea of what needs to happen. How would you get the buy-in, for example, of Anthropic or Google or OpenAI, wouldn’t that be different than a product that is the wrapper on top of OpenAI that is tailored specifically for mental health?
Charlotte Blease: Yeah, it’s a complicated question. Where we previously had the sort of hegemony and the monopoly, if you like, of medical professionals or healthcare professionals, include more broadly clinical psychology and psychotherapy dominating the mediation of healthcare. Now we’ve got this very awkward, from a regulatory perspective, hybrid approach because behind the scenes people are availing of these consumer tools. Do we want Google to be our healthcare organization or anthropic? How do we regulate that? My preoccupation has always been AI doesn’t enter into existence as sort of a neutral bubble or even against what we could romanticize as optimal care. And in fact, what I find really interesting about the debate about AI is it tends inadvertently to shine a light on the problems with, if you like, human-mediated care. So take the issue of serious, severe mental health, the example of suicidal ideation. When we focus on the outcomes rather than overly romanticize what the clinical judgment looks like or human judgment looks like, we have some really stark facts. If you look at electronic health records, you tend to see, and there’s been work on this at Harvard and MIT, some of it’s a few years old, but if you look at, there’s a validated study that across diverse health systems that basically use huge data sets to create prediction models that found that there was a sort of guestalt of features within an electronic health record based on people who do end up taking their own lives. And for about 40% of patients, they were able to estimate or predict that they would do this two years in advance. So you’ve got all kinds of ways in which you can use electronic health records, data that can shed light on patterns that are just not visible to human beings. I think it’s great that we are asking serious questions about the harms of AI, but let’s also ask the very same questions about what we’ve currently got. The whole purpose of my book was to ask the question: what are the problems to which AI could be a solution? And if you think about it from that perspective, there’s quite a lot of problems with care, as we know it, to which AI might potentially offer some solutions.
Stephanie Hepburn: There’s some differentiation there between taking huge data sets and interpreting versus dialogue, for example, with the large language models. With generative AI, it’s pretty inconsistent. You’re not always going to get the same answer. It may help identify what gaps are existing in our current system. So, for example, if people are turning to AI, you know, at three o’clock in the morning because they’re struggling, or they’re more apt to turn to an AI chatbot to share information about their suicidal ideation than they are with their therapist. I think there needs to be an examination of why, what is not being served by existing systems, or are they underfunded, or are they missing key components that need to be there to ensure that people can get the care that they need?
Charlotte Blease: I think that’s absolutely right. Aside from cost of healthcare coverage, there are other kinds of costs associated with actually taking time out of a gig economy job, transportation to visit a professional. And if you do actually manage to get to see a healthcare professional, say it’s a psychotherapist or a clinical psychologist, and that’s scheduled once a week, you know, the average patient in the States, there’s a study of a sort of time of motion study that for a 20-minute visit, and this was for a primary care visit, people took an average of two hours out of their day. And that was higher if you’re lower income. So you’ve all these barriers. And again, if you get to see that clinician for, say, an hour a week or whatever it is, you’ve also got the issue of you’ve got to store up your problems, and there’s a performative aspect. I’ve got to discuss the problems that I’ve been experiencing in this concentrated one hour per week versus the accessibility of, you know, most people have smartphones now. That’s increasing the world over. But it invites a whole range of new challenges and problems as well. Is the endless availability going to lead to, aside from even the leakage of privacy and sensitive information, the seductions of doing that, does that lead to dependency? We might complain that people are dependent on their therapists when they do get to see them, and that that’s sometimes not a good thing. And we may be entering a phase where people overly turn to AI. We could be on the cusp of something even greater. I think a large part of it depends on how these tools are designed. The other issue is it’s very hard to tell people not to do things when they see a use. People are pretty imaginative on how they’re using these tools in all kinds of ways: work and leisure and health support. And it reaches a point when one must also ask when is it paternalistic to say don’t use these tools, particularly if clinicians are indulging in them themselves?
Stephanie Hepburn: I think sometimes users have an assumption that AI is without biases or can’t be overconfident, which they often illustrate themselves to be. They’re often sycophantic. Because they’re generative, they can give inconsistent answers. All of the data and information to train AI has been scraped from humans. And so that inherently, as we know, is filled with biases. There’s also this assumption that the AI chatbot has the answer, right? Because it can do a deep dive, that it’s going to arrive at the right answer. So I think there’s a concern there that people are turning to chatbots, not only because you mentioned about accessibility and also feeling like you’re not going to be judged, but also an assumption that the query response and the information that you’re getting from the chatbot is accurate.
Charlotte Blease: Yeah, I think it’s a great point. And it’s not an uncommon problem in the sense that people voice this problem a lot algorithmic bias and the equity issue here. There are real challenges here. So, in a sense, the AI, the machine tools really reflect society at large, but we’ve also got problems with care as we know it. And biases can kick in in a face-to-face visit, where these could potentially be subverted via not having a face-to-face visit. Incidentally, the Kaiser Family Foundation surveys found that African Americans and Hispanic Americans were more likely to have confidence in the use of these tools. And that’s been quite consistent over surveys that KFF have conducted from 2024-2026. There’s a higher up adoption, and more favorable adoption among minorities in the States. And I think that’s interesting because it might be a reflective of accessibility, it could be reflective of legacy structures in care as we know it, and fears about stigmatization or prejudice in face-to-face visits. And this is where I think AI, if we were to take very seriously the opportunities for equity, and I’m certainly not alone in this perspective in taking this view, it may be easier to de-bias technology than it is to de-bias humans. It’s quite difficult to de-bias a human being. People tend to spot the prejudice in other people, but they’re quite bad at spotting it in themselves. So there you have certain opportunities with these tools if we really take seriously the opportunity and we take the need for equity seriously here. But that’s going to be a big challenge if we’re leaving, if we’re going to end up outsourcing healthcare to big tech companies who may have no interest in sharing the same values as health organizations.
Stephanie Hepburn: So in the paper, you mention some global standards. You mentioned spirit AI and consort AI. Is that more about guardrails?
Charlotte Blease: Those are largely to do consort in particular is thinking about the role of clinical trials here. So I see you’ve hit another really important point that we do need to discuss. And part of the article was focusing on control groups in AI. And actually, when you look at the evidence, there has been, if you consider a drug that’s going to market, you’ve got to have a control group. For the studies that compare AI tools, they tend to just compare them to weightless controls, which is a very inadequate way of assessing the potential benefits and risks and the effectiveness of these tools. So consort AI is thinking about what are the reporting guidelines when you do a clinical trial into AI and use it, for example, as a decision-making tool or thinking about an actual treatment in itself. So, for example, New England Journal of Medicine AI, which is a very prestigious, relatively new journal, it published a study on a tool called Therobot for depression, and it reported that this tool, this AI chatbot, was highly effective for depression. But what did they compare it to? They didn’t have a control, they just had a wait list. Now, when you do that, you end up inflating the effect size because of all kinds of reporting biases and the noise that can exist in a clinical trial. That’s why you need a placebo control group. So what we’ve advocated here is we need to have more controls. We need adequate and robust controls. That requires a great deal of ingenuity in how you design a control in this scenario. But nonetheless, we need to have the appetite to start doing that. Otherwise, we’re just going to give a free pass to technology. And again, that leads to ethical consequences whereby people may be harmed because we’re telling them to use tools that haven’t been properly assessed, and we’re missing opportunities to help people because we’re giving them inadequate lines of treatment. So there’s a lot of hard work to be done in assessing the evidence base for these tools. And that’s where the need for more stringent reporting guidelines needs to be in place.
Stephanie Hepburn: If I’m understanding correctly, the idea is that there need to be very clear protocols during the trial, after the trial, to be able to examine the efficacy.
Charlotte Blease: Absolutely. Yeah, that’s the point. And need for replication as well. Again, there’s real serious potential here because, for example, if we’re trying to assess the effectiveness of an AI chatbot, that in some ways might be easier than doing it with human beings and clinical psychology. It’s really challenging. In another life, I look very seriously at the evidence based in clinical psychology, did a lot of research in it. And it’s really a big challenge designing controls for psychological treatments that are administered by humans. There’s so many variables. In the case of AI chatbots, it could be easier to design these controls, but we need to design controls. We can’t just have it as a case again of a waitlist control because people tend to do worse when it is, particularly if they have enrolled in something and then they’re just allocated to the wait list. And then you tend to get this huge disparity between people who are offered the treatment under scrutiny, the chatbot under scrutiny, they tend to be more inflated in their responses. And we’ve got to work harder there to screen out, as I say, some of that noise.
Stephanie Hepburn: One challenge as well is that LLMs are updating all the time. So if we’re looking at it from a study perspective, that would make it almost immediately outdated.
Charlotte Blease: It’s a problem with all kinds of survey research, incidentally, as well, where by the time you’ve surveyed a population and it’s actually published in a peer-reviewed journal, it’s very often anything between 12 to 18 months out of date. So we’re always catching up when it comes to AI, but there are also opportunities to leverage AI to move more rapidly here as well in design and assessment of these tools. You know, what’s interesting because I was recently at a conference that was all about clinical trials and the use of what do we do about assessing AI interventions and moving the needle in all of this and progressing. But what was interesting there too was clinical trialists saying that they too, sort of behind the scenes, were availing of commercial AI tools to cut corners in all kinds of regards. So, really, in every corner of healthcare, we’ve got to be more transparent too about the use of all of these tools and how we’re using them, because I think there’s a great deal of camouflage and cover-up that’s happening right now for all kinds of status reasons as well. People are afraid to sort of admit that they indulge in these tools.
Stephanie Hepburn: Yeah, that’s a great point as well. I think it would be interesting if there was more openness with the large language models where they could give and would give and would be willing to give more information about because right now they’re kind of a black box of data. So occasionally, I notice OpenAI will give kind of a summary report, put out a blog, and it’ll talk about mental health and mental health crisis. But it would be interesting and I think incredibly helpful to have that be more open and have that be a data source. I feel like there could be some more symbiosis between the researchers and these large language model companies. I don’t know if there’s any incentive on their part to do that. But, you know, maybe that’s something that could end up being compelled by legislation. I don’t know. But it’s such a wealth of data, similar to how I feel 911 would just be such a wealth of data as well, like opening up that black box just to be able to identify what people are really struggling with, what the gaps are in the system. So I feel like there’s room for partnership, but that would require a certain degree of buy-in from these large language model companies. I think the companies that are maybe not designating themselves as wellness, but designating themselves as mental health, they’re more apt to buy in. So I was curious as to what your thoughts are there. I think this conversation is so important. I think the paper is so important. I’m curious what your thoughts are, though, in terms of the audience. Who’s going to really take this feedback and say, okay, we need to do something about this? And what kind of buy-in can we have from the large language model companies to partner on this? It’s such an important question.
Charlotte Blease: And it ultimately, I think, gets to the crux of the matter. And in fact, it moves it beyond the paper itself. But I believe that politically there has to be more of a reckoning about exactly what you say. Open AI is sitting in an absolute gold mine of information. It knows the scale and the extent of people using these tools for all kinds of health queries. Was it in October 2025 they said 1.2 million people every week were discussing suicidal ideation with ChatGPT? If anything, those figures will have gone up. And the sort of get out of jail clause, quite literally, is the this is not a replacement for care. But the reality is a system that is wholly responsible will recognize people are using these tools. They’re using them exactly for mental health support, crisis support. And what we’ve got on our hands, whether we like it or not, is a mass population experiment. We don’t know if people are, in some circumstances, faring better or doing a lot worse as a result of these tools. We just simply don’t have the data on it. And it’s very hard with small funding pots and limited capabilities in academic institutions to conduct this kind of research, even in international consortia. It needs a much more robust national and international effort and a recognition that our health systems, whether private or public sphere, are being unseated in some sense. That we do have a shadow health system with these kinds of AI tools. And the experimentation is extending to issues of justice. Some people are going to be higher adopters of these tools. To what consequence? What scale of misinformation or benefit? We just don’t know the answers. And the only way to achieve that is to say, is what you’ve said is also that we’ve got to commit sort of serious funding and serious coalitions to investigate this. One can only tinker in the ivory tower. It becomes somewhat ridiculous. You’re commenting and stuff to which very often you have limited access.
Stephanie Hepburn: Yeah. Your point about it being a shadow healthcare system is so poignant because I think that’s something I haven’t thought of it. And that’s exactly accurate if there are individuals who are turning to AI chatbots. Maybe they’re doing it in a supplementary way, but I think when people don’t have access or are fearful of stigma repercussions, they may be turning to it and it alone. So I think that was such a profound point. What are you working on next?
Charlotte Blease: I’m returning to my interest in psychological treatments, specifically considering the regulations and the effectiveness and the challenges associated with psychological treatments, but also what’s happening with AI. And I’ve always been interested in psychotherapy, the evidence base, how it works, informed consent to it. But part of the reason I’m returning to this interest is because of AI and because of a recognition that there’s almost a banality to it. I mean, I hear so often from people they consult with AI over all kinds of things, but the use of it within mental health support is really interesting. An elderly gentleman came up to me at the end of a presentation recently and he said, What do you do when you’re in grief? But an AI chatbot knows you better than any human. And I thought, my God. And I thought he will not be alone in this. And knowing that younger people are faster adopters of these tools and they may always have this chatbot counsellor on hand, what is that going to do to human relationships? What are the consequences for real-world relationships for mental health, both for good and for ill?
Stephanie Hepburn: That was researcher Charlotte Blease on how large language models have become a shadow healthcare system. While we can’t necessarily put the genie back in the bottle, she says we need clear, enforceable, evidentiary standards to make AI safer. If you enjoyed this episode, please subscribe and leave us a review wherever you listen to the podcast. It helps others find the show. Thanks for listening. I’m your host and producer, our associate producer’s Rin Koenig, Audio Engineering by Michael Keene. Music is Vinyl Couch by Blue Dot Sessions.
References
Multidisciplinary research priorities for artificial intelligence in mental health: a call to action
Strengthening ChatGPT’s Responses In Sensitive Conversations
Dr Bot: Why Doctors Can Fail Us and How AI Could Save Lives
Randomized Trial of a Generative AI Chatbot for Mental Health Treatment
Where to listen and subscribe
Apple | Spotify | Amazon | iHeart | YouTube
We want to hear from you
Have you turned on “Trusted Contact” in ChatGPT? If so, we’d love to chat with you.
What questions do you want answered? Have you turned to an AI chatbot to discuss a personal or mental health issue? Work at an AI company? Are you a researcher studying AI and mental health? We want to hear from you. Reach us at editor@crisisnow.com
What to listen to next
- More Young People Are Turning to AI Chatbots for Mental Health Advice. Now What? – Ep 16
- Canary in the Coal Mine: What is the Impact of AI on Tech Workers? Part 2 – Ep 15
- Canary in the Coal Mine: What is the Impact of AI on Tech Workers? Part 1 – Ep 14
- What’s Happening With All Your Therapy App Messages? — Ep 13
- Can an AI Chatbot Connect Student Survivors of Sexual Violence to Resources? — Ep 12
Credits
CrisisTalk is hosted and produced by Stephanie Hepburn. Our associate producer is Rin Koenig. Audio engineering by Michael Keene. Music is Vinyl Couch by Blue Dot Sessions.
