Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Randomized trial in Nature Health: a GPT-4o respiratory…

Randomized trial in Nature Health: a GPT-4o respiratory chatbot in WeChat beats web search for lay diagnosis (70.0% vs 55.4%)

★★after cutoffresearchNature Healthconfidence: medium

A preregistered randomized study published in Nature Health on Oct 5, 2026 tested LungDiag, a GPT-4o-based respiratory chatbot inside WeChat, against ordinary mobile web search (with AI search features disabled). Among 2,400 Chinese adults without medical training who worked through respiratory illness vignettes, the chatbot group reached the correct diagnosis 70.0% of the time, against 55.4% for search, a 14.6-point difference.

Key facts

What happened

The study compares a task-specific, retrieval-augmented chatbot built on GPT-4o and embedded in WeChat with what lay people usually do, which is to search the web. Search-engine AI answers were switched off in the control arm, so the comparison is with plain search.

Why it matters

It is one of the larger prospective randomized comparisons of a consumer health chatbot against web search, and it is peer-reviewed. The outcome is accuracy on case vignettes, not patient outcomes. The model, GPT-4o, was released in May 2024, more than two years before publication. The numbers come from the Bioengineer.org summary. The Nature Health page was confirmed to exist, but its full text was not read, hence confidence: medium.

Changelog

  • 2026-10-05: created (science sweep, 21:30 run)

Sources (2)

id: 2026-10-05-lungdiag-chatbot-vs-web-search-rct-nature-health · updated 2026-10-05 · open in the interactive timeline