AI Detects Depression with Real-World Conversations

This title was summarized by AI from the post below.

Most AI that "detects depression" is quietly cheating. It's trained on clinical interviews, where a doctor asks the scripted PHQ-9 questions — and the model learns to read the interviewer's prompts, not the patient. Great benchmark scores, little real-world value. If a structured assessment is already happening, you didn't need the model. A new paper from the Ash by Slingshot AI team (the Ash app) goes after the harder problem. They fine-tune a 27B LLM to predict someone's PHQ-9 score from nothing but the free text of their first week of conversations with an AI therapy app. No questionnaire, no prompts — just how people actually talk when they're seeking support. And the results hold up: MAE ~2.6 and a 0.80 correlation with the real score, with strong discrimination (AUC > 0.87) across the entire severity range, not just at a single cutoff. What I respected most, as an engineer: → They chose the honest, harder setting — naturalistic dialogue instead of clinician-elicited speech. That distinction is the whole game, and most papers blur it. → They kept the full 9-item PHQ-9, including the suicidal-ideation item that most studies quietly drop to make the task easier. → Smart data work: with few labels, they used a strong reasoning model to rebalance the dataset, then iteratively self-trained. Clean engineering for a low-label clinical problem. They were candid about the model's weak spots instead of hiding them. Worth being clear-eyed about the limits, though: → It's a single snapshot — one stretch of conversation mapped to one score. The real prize is tracking the same person over weeks and catching the moment they start to slip. That's a much harder problem, and the mountain still to climb. → The model drifts toward the average and is weakest at the low end — exactly where general-population screening matters most. → It's one platform, with users who self-selected into both chatting and filling out the questionnaire, in a group where ~80% were already depressed. Impressive in that setting; an open question how it travels to the wider world. → And the architecture has a cost we don't talk about enough: a 27B model is too heavy to run on a phone, so deeply personal mental-health conversations have to leave the device and be processed in the cloud. For psychiatric data, that's not a detail — it's a core design constraint. The most sensitive signals are exactly the ones you'd most want to keep on-device. But the direction is right. The future of mental health measurement isn't another form to fill out. It's measurement that happens in the background of care — understanding how someone's doing without adding burden. Ideally without that data ever having to leave the patient's phone. Good to see a serious, honest step toward it.

  • chart, scatter chart

Important paper, but let's not over-frame it. Predicting a PHQ-9 score from free text is not 'detecting depression'—it's detecting linguistic correlates of a self-report screener in a population where ~80% were already depressed. That performance won't travel well to the general public. More critically, this model should never be used alone clinically. It relies entirely on the practitioner to ground-truth it against affect, behavior, and life context—things no LLM sees. It must be overridden when it misses subtle risk, especially around suicidality. If a clinician uses this score as a proxy for their own judgment, we've just automated clinical laziness. Good engineering. But the real mountain isn't just tracking over time—it's integrating these tools without eroding the human judgment that keeps care safe. Let's not confuse a strong correlation with clinical readiness. Kintsugi's tech was clinically valid but died from the VC-FDA timeline mismatch. Ash is promising but unproven. Both need rigorous validation and a clear support role—not replacement—in care. Governance. Healthy skepticism warranted.

I agree Tanel Petelot. the naturalistic vs. clinical-interview distinction is real and important. I am curious though if it’s it's worth asking how much of the signal is explicit symptom disclosure elicited by the CBT framing vs. something more latent 🤷♀️ because the answer shapes how well this could travel to other platforms and other interaction styles…

IMO, not many AI mental health apps lead to meaningful, lasting change. However, this tool seems like it could provide clinicians with valuable longitudinal data from everyday conversations, helping them identify concerning shifts in mood, cognition, or behavior earlier and track changes over time.

I can see this becoming an advantageous tool if implemented such that privacy is protected. Validated scales pose a major burden. Prediction from words and records is preferable as a route if validated.

Try the SMST instead. Pre-screens for depression, AUC-derived cut-off at 10.5 out of 20 possible points. https://europepmc.org/article/MED/31303020

Tanel Petelot - I agree that this is the direction of travel which is most aligned with a natural humane milieu. I wonder if you may find this site interesting? https://embodied-cognition-psych-rzxpdb6.gamma.site/

  • No alternative text description for this image

Now combine everyday conversations, facial expressions and gestures. Smiles

Like
Reply
See more comments

To view or add a comment, sign in

Explore content categories