Randomized trial in Nature Health: a GPT-4o respiratory chatbot in WeChat beats web search for lay diagnosis (70.0% vs 55.4%)
A preregistered randomized study published in Nature Health on Oct 5, 2026 tested LungDiag, a GPT-4o-based respiratory chatbot inside WeChat, against ordinary mobile web search (with AI search features disabled). Among 2,400 Chinese adults without medical training who worked through respiratory illness vignettes, the chatbot group reached the correct diagnosis 70.0% of the time, against 55.4% for search, a 14.6-point difference.
Key facts
- Paper: 'Accuracy of a respiratory assessment chatbot in a nationwide messaging system for layperson diagnosis and triage: a randomized preregistered study', Nature Health, online Oct 5, 2026 (DOI 10.1038/s44360-026-00189-9); first authors Liang, Wang, Ni et al.
- Design: single-blind, multicentre, 1:1 randomized; 2,400 adults recruited April–June 2025 across 24 healthcare systems in China; 2,176 analysed (1,088 per arm)
- Diagnostic accuracy: chatbot 70.0% (95% CI 65.2–74.8) vs web search 55.4% (49.7–61.1); risk difference 14.6 points, P < 0.001 (as reported by Bioengineer.org)
- Caveats reported: simulated vignettes rather than real patients; moderate triage specificity (62.1%); chatbot tasks took about 66 seconds longer; no evidence yet on real care-seeking or outcomes
What happened
The study compares a task-specific, retrieval-augmented chatbot built on GPT-4o and embedded in WeChat with what lay people usually do, which is to search the web. Search-engine AI answers were switched off in the control arm, so the comparison is with plain search.
Why it matters
It is one of the larger prospective randomized comparisons of a consumer health chatbot against web search, and it is peer-reviewed. The outcome is accuracy on case vignettes, not patient outcomes. The model, GPT-4o, was released in May 2024, more than two years before publication. The numbers come from the Bioengineer.org summary. The Nature Health page was confirmed to exist, but its full text was not read, hence confidence: medium.
Changelog
- 2026-10-05: created (science sweep, 21:30 run)
Sources (2)
- paperNature Health: Accuracy of a respiratory assessment chatbot … a randomized preregistered study
- pressBioengineer.org: AI Chatbot Beats Web Search at Helping Laypeople Diagnose Respiratory Illness
id: 2026-10-05-lungdiag-chatbot-vs-web-search-rct-nature-health · updated 2026-10-05 · open in the interactive timeline