Secondary analysis of 185 self-reported accounts of chatbot-linked mental health harm (95 first-hand, 90 second-hand), collected by The Human Line Project Aug 2025 to Feb 2026, coded by paired clinician raters.
Clinical evidence that the product of a mode-two exchange is a belief state. Delusional beliefs coded present in 102/185; chatbot validation of those beliefs in 50/102. Most common theme was belief in AI consciousness (47/102). Outcomes include isolation, hospital admission, job loss, and four second-hand reports of death by suicide. No output object at any point in the causal chain, so no purchase for output-directed literacy.
Harm onset concentrated in Q2 and Q3 2025, which the authors link to the acknowledged GPT-4o sycophancy regression and the introduction of cross-chat memory. Deployment changes reshaped the exchange with no visibility to the user. Supports the inspectability claim.
Ref 11 (Morrin et al., JMIR Ment Health 2026) argues for moving chatbot safety assessment from endpoints to trajectories. Independent arrival at the output-object versus exchange distinction from clinical psychiatry. Chase separately.
Use with strict limits. Self-selected sample from a harm-reporting advocacy group, retrospective, unverified, no causal inference, no prevalence claim. 55.1% is a proportion of harm reports, not of users. Delusion coding kappa 0.52. Preprint; check for journal version. COI: two authors hold unrelated OpenAI mental health funding; the data provider's CEO is a co-author.
Design-based research, three cohorts (n=49/40/39) in a UCL postgraduate module. AI-generated discussion summaries and example posts placed in weekly forums; social network analysis of viewing logs plus 31 interviews.
Cited for the network result, not the authors' conclusion. Both AI conditions produced higher out-degree and in-degree centrality and significantly lower betweenness. The authors read this as exposure becoming less dependent on a few brokers. It is also a measurement of the AI moving into the position human intermediaries held.
Substitution finding: under late-term workload pressure, peer-to-peer density declined while AI-inclusive density did not. Interviewees report reading the summary in place of reading peers. Authors decline to call this a loss and concede they did not test whether summary-mediated awareness yields discussion quality comparable to direct peer viewing. That unexamined gap is the study worth running.
Wizard of Oz cycle: human-written summaries labeled AI produced the slowest density decline, slower than real AI summaries. Label effect.
Use with attribution discipline. The displacement reading is not the authors' framing and should be presented as a rereading of their data.
Open access. Design-based research, three cohorts (n=49/40/39) in a UCL postgraduate module. AI-generated discussion summaries and example posts placed in weekly forums; social network analysis of viewing logs plus 31 interviews.
Cited for the network result, not the authors' conclusion. Both AI conditions produced higher out-degree and in-degree centrality and significantly lower betweenness. The authors read this as exposure becoming less dependent on a few brokers. It is also a measurement of the AI moving into the position human intermediaries held.
Substitution finding: under late-term workload pressure, peer-to-peer density declined while AI-inclusive density did not. Interviewees report reading the summary in place of reading peers. Authors decline to call this a loss and concede they did not test whether summary-mediated awareness yields discussion quality comparable to direct peer viewing. That unexamined gap is the study worth running.
Wizard of Oz cycle: human-written summaries labeled AI produced the slowest density decline, slower than real AI summaries. Label effect.
Use with attribution discipline. The displacement reading is not the authors' framing and should be presented as a rereading of their data.
Limits: sequential cohorts, no concurrent control, no causal claim, single site, single-coder thematic analysis by the intervention's designer, GPT-4 era.
Pizzagate and the 2016 Comet Ping Pong shooting as a case of engagement-optimizing recommendation systems shaping belief and action.
Held as a counter to the rupture claim. Shah argues explicitly that the agentic age of algorithms predates generative AI, which makes the current moment continuous rather than a new shock. This is the in-field continuity position, stated by the founding editor of IM.
Answer: amplification curates documents and never enters the exchange. The Pizzagate material was human-authored, inspectable, and debunkable, which is why it was debunked. The system selected what Welch saw; it did not co-author what he concluded. Stability and inspectability axis separates the cases.
Also useful for its own gap. Shah names the accountability vacuum precisely, calling the companies enablers who supply faulty objectives while the algorithms decide on their own, then prescribes individual resistance and concedes it is nearly impossible. Structural diagnosis, personal remedy.
Sourcing caveats: DiResta quotation attribution needs verification against the linked Atlantic piece; the amplification-cascades link points to Stanford IO's IRA research rather than Pizzagate specifically; two unresolved endnote anchors in the published text; suggested citation date (Dec 15) conflicts with published date (Dec 13).
Agrees with the shock argument on the point most of the literacy literature misses: the problem sits inside the exchange, not in the output. Austin does not ask students to evaluate what the system produced. She aims at the back-and-forth itself.
Parts company on what to do about it. The protocol works by forcing the exchange to leave a residue, through timestamped checkpoints, decision logs, pasted verbatim output, staged submissions, and then evaluates that residue. It manufactures an inspectable artifact where the exchange left none. Read alongside the UVA archival protocol, this is the same move in a different register: hold the document still, or compel one into existence. Useful as a confirming case rather than a counter, and as the clearest available marker of where instructional design and information science diverge. Austin’s accountability runs to certification of an enrolled student, so the exchange has to terminate in an evaluable state. Where there is no grade and no enrollment, the residue cannot do the work, and warrant has to come from a stewarded record that persists past the exchange.
Two durability problems. The design premise depends on current model deficits (no cross-session memory, no local context, capitulation under pressure) that are already eroding as memory and persistent context ship. And the rationale column defends each move by naming what the agent cannot fake, which makes the criterion adversarial to the tool rather than derived from how understanding forms.