C5R opens Facility-0, an AI-run wet lab, and the SciUniverse benchmark: Claude Fable 5.1 leads with 45.3%
On Sept 24, 2026 San Francisco startup C5R (founder Michael Akilian) launched Facility-0, a physical lab where frontier models plan experiments and run them by controlling instruments and instructing human operators. It also launched SciUniverse Level 1, a benchmark of 92 tasks in 17 families across chemistry, biology and materials science, some done in the real lab and some in a digital twin. Claude Fable 5.1 scored 45.3% pass@1, ahead of GPT-6 Astra (32.5%) and Claude Opus 5 (30.5%). C5R says models "repeatedly miss critical details in laboratory work".
Key facts
- SciUniverse Level 1: 92 tasks, 17 task families (chemistry 7, biology 5, materials 4, cross-domain 1); tasks 'take no more than a few hours' for a scientist; family-weighted pass@1
- Leaderboard (pass@1, model inference cost per task): Claude Fable 5.1 xhigh 45.3% ($40.61); GPT-6 Astra xhigh 32.5% ($52.37); Claude Opus 5 xhigh 30.5% ($46.31); Grok 4.6 xhigh 26.2% ($13.41); Gemini 3.8 Flash high 14.6% ($4.55); GPT-5.6 Sol xhigh 9.4% ($16.53)
- Real-world tasks include synthesising and detecting an amide (N-benzyl-4-methylbenzamide) by LC–MS, cell-free sfGFP expression, PCR and DNA recovery, alpha-alumina synthesis and pressing BaTiO₃ pellets; simulated tasks include a two-week automated PCR lab with supply shortages and equipment failures
- Failure modes named by C5R: pipetting frozen samples, contaminating DNA, ignoring evaporating solvents, vortexing open containers
- Facility-0 was built in 12 weeks; press coverage reports more than 40 integrated instruments
- Launch post on X: 787k views (account with about 3k followers) when checked Oct 5; covered in Import AI 475 (Oct 5)
What happened
C5R says its goal is to give models a physical workspace: "Getting AI to do useful scientific work requires a physical workspace connected to instruments and experiments, rather than another system that operates only on digital data" (Akilian, per Runtime Wire). In Facility-0 the model drives a harness that operates machines (pipettes, LC–MS, X-ray, presses, a Hamilton liquid handler in simulation) and gives instructions to human operators. The website shows recorded footage of GPT-6 Astra synthesising an amide.
SciUniverse Level 1 is deliberately easy work for a trained scientist. The best model still fails more than half of it. C5R notes that real-world tasks have few rollouts, so those scores are noisy.
Why it matters
Most AI-for-science benchmarks test reasoning on text or data. This one tests whether a model's plan works with real reagents and instruments, and the low scores show how far "AI scientist" claims are from autonomous wet-lab work. It is a new benchmark from a startup and has not been independently reproduced. Funding was not disclosed.
Changelog
- 2026-10-05: created (found via Import AI 475 in the science sweep, 21:30 run)
People
Related posts (2)
- michael original ↗ michael @akilian · x · 2026-09-24
Cited as a source by: 2026-09-24-c5r-facility-0-sciuniverse-benchmark - C5R: 'In 12 weeks, we built a research facility that is run entirely by AI' (SciUniverse benchmark launch) original ↗ C5R CORP @c5rcorp · x · 2026-09-24
Launch post for Facility-0 and the SciUniverse real-lab benchmark; 787k views from an account with about 3k followers.
Sources (6)
- officialC5R: SciUniverse benchmark and results
- officialC5R launch post on X
- discussionMichael Akilian launch thread
- pressRuntime Wire: C5R is building physical labs for AI models that still struggle with experiments
- pressGlitchwire: C5R built a research lab run by AI in 12 weeks
- discussionImport AI 475 (Jack Clark, Oct 5, 2026)
id: 2026-09-24-c5r-facility-0-sciuniverse-benchmark · updated 2026-10-05 · open in the interactive timeline