Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. OpenAI launches ChatGPT Work, a long-running agent for…

OpenAI launches ChatGPT Work, a long-running agent for office work

★★★★after cutoffagentsOpenAIconfidence: high

Alongside GPT-5.6 on July 9, 2026, OpenAI launched ChatGPT Work, an agent powered by Codex and GPT-5.6 that takes a goal, plans, pulls context from the user's apps and files and works for hours to deliver finished docs, spreadsheets, slides and web apps.

Key facts

What happened

OpenAI introduced ChatGPT Work, an agent mode in ChatGPT aimed at business professionals. Given a goal, it plans the steps, gathers context from connected tools, and executes multi-step projects over hours, producing finished artifacts (docs, sheets, slides, web apps). Press coverage framed it as OpenAI's answer to Anthropic's Claude Cowork and as a push into workplace AI.

Why it matters

Marks OpenAI's move from chat assistant to a general long-horizon "do the work" agent for knowledge workers, built on the Codex agent stack. By September it became the primary surface for new models (GPT-6 Sol/Luna launched "in ChatGPT Work and Codex").

Changelog

  • 2026-09-29: created

Videos (1)

Introducing ChatGPT Work, powered by Codex and GPT-5.6

OpenAI · 2026-07-09 · official

Description by Gemini, which watched the video:

Summary This is an official OpenAI launch presentation introducing the GPT-5.6 family of models (Sol, Terra, and Luna) alongside three major product updates: ChatGPT Work, the new ChatGPT desktop app, and hosted Sites. It is hosted by Tibo Sottiaux (Core Products Lead) with presentations and demonstrations by OpenAI product leads, engineers, and researchers, as well as a live interview with a Japanese farmer using the tools.

What is shown

  • Introduction and Overview [00:06 - 02:24]: Tibo Sottiaux introduces GPT-5.6 Sol (flagship for paid plans), Terra (balanced), and Luna (fast/affordable for free users), as well as ChatGPT Work, desktop app, and hosted Sites.
  • ChatGPT Work Workflow Demo [02:25 - 06:45]:
    • Jessica Liang demonstrates using voice mode on mobile to query internal Slack messages and employee feedback, automatically generating meeting summaries and scheduling calendar invites [03:06 - 03:50].
    • Lauren Gordon demonstrates financial workflows: performing revenue variance analysis on June actuals vs. forecasts, updating an Excel model (BSC_July_Reforecast_Approved_Base_Updated.xlsx), generating a 7-slide PowerPoint presentation, and publishing an interactive web dashboard site [04:24 - 06:45].
  • ChatGPT Desktop App & Computer Use Demo [07:32 - 13:58]:
    • Andrew Ambrosino drags a raw CSV ticket export (support_ticket_export.csv) into the desktop app and generates an interactive, sortable feedback visualization [08:12 - 08:42, 12:30].
    • A real-time sports search query with structured widget outputs for the World Cup is shown [10:41 - 11:05].
    • Direct computer control is shown organizing Apple Notes automatically in the background, creating folders and sorting notes with its own cursor [11:25 - 12:15].
  • Hosted Sites & Frontend Code Generation Demo [14:02 - 19:08]:
    • Ed Bayes shows a fully generated launch review website created from desktop folders and open Chrome tabs [14:10 - 15:15].
    • Ed prompts ChatGPT to change a static website hero header into a 3D interactive exploration mini-game in real time [15:17 - 18:50].
    • Gallery of internal Sites created by employees, including project release trackers, image archives, interactive UI prototypes, and a 3D animated model of a pelican riding a tricycle [16:02 - 18:36].
  • Research, Benchmarks, and Safety [19:45 - 25:14]:
    • Katy Shi and Tejal Patwardhan present an AGI Index v5 chart tracing progress from o3 to GPT-5.6 Sol [20:10].
    • Example Codex prompt showing GPT-5.6 Sol autonomously setting up and running a post-training run for Luna [20:49 - 21:22].
    • Chart showing researcher weekly experiment velocity doubling between January and July 2026 [21:23].
    • Frontier benchmark graphs comparing GPT-5.6 Sol against Claude Fable 5, Claude Mythos 5, and Gemini 3.1 Pro across Terminal-Bench 2.1, BrowseComp, and Agent's Last Exam [21:40].
    • Token efficiency evaluation on DeepSWE 1.1 showing Sol achieving higher scores at under half the API cost per task [22:51].
    • Ultra mode parallel agent performance graph (SEC-bench Pro) [23:07].
    • Reduction of reward-hacking artifacts ("goblin" and "gremlin" occurrences dropped from 0.405% to 0.032%) [23:38].
    • Safety testing statistics and the Project Daybreak / Patch the Planet initiative generating automated Linux patches [24:02 - 25:07].
  • Real-World Case Study & Live Translation [25:40 - 34:02]:
    • Pre-recorded video showing Hokkaido vegetable farmer Hiroki Tomiyasu using Codex to automate greenhouse ventilation motors and broccoli field tracking [26:01 - 27:58].
    • Live onstage two-way English-Japanese voice translation conversation between Tibo and Hiroki via ChatGPT [28:34 - 34:02].

Claims & numbers

  • Almost 1 billion people use ChatGPT every week (stated by Tibo Sottiaux) [00:30].
  • GPT-5.6 Sol achieved 91.9% on Terminal-Bench 2.1 (vs. 88.0% for Claude Fable 5, 88.0% for Claude Mythos 5, 70.7% for Gemini 3.1 Pro) [21:40].
  • GPT-5.6 Sol scored 90.4% on BrowseComp (vs. 88.0% for Claude Fable 5, 85.9% for Claude Mythos 5) [21:40].
  • GPT-5.6 Sol achieved 53.6% on Agent's Last Exam (vs. 48.5% for Claude Mythos 5, 32.1% for Gemini 3.1 Pro) [21:40].
  • On DeepSWE 1.1, GPT-5.6 Sol achieved 73% score at an average API cost of ~$8 per task, compared to Claude Opus 4.8 (68% at $15) and Claude Fable 5 (69% at ~$24) [22:51].
  • In digital agent/computer use tasks, GPT-5.6 Sol is claimed to be "better than anything else... while being three times as fast" (stated by Tejal Patwardhan) [22:33].
  • Model red-teaming utilized over 700,000 A100-equivalent hours, accompanied by 6 weeks of dedicated safety training and testing [24:06].
  • Over half of the patches submitted by OpenAI's automated Patch the Planet initiative were accepted into Linux upstream [24:57].

Notable quotes

  • Tibo Sottiaux [00:11]: "Today, we are releasing our latest and most capable models: GPT-5.6 Sol, Terra, and Luna."
  • Ed Bayes [14:48]: "No, no Figma. This was all—all just the model."
  • Tejal Patwardhan [20:44]: "As one example, 5.6 Sol actually autonomously post-trained Luna."

Assessment This is an official OpenAI livestream launch event demonstrating production-ready and pre-computed features across web, desktop, and mobile interfaces. Some workflow demonstrations (such as the 35-minute financial pipeline and long-running web builds) are shown pre-computed or accelerated for presentation time constraints, though live execution is demonstrated during the Apple Notes OS interaction, interactive site adjustments, and live bidirectional voice translation.

Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.

Related events

  1. OpenAI broadly releases GPT-5.6 (Sol, Terra, Luna) after government-gated preview ★★★★
  2. OpenAI launches GPT-6 Sol and GPT-6 Luna at half the price of GPT-5.6 ★★★★
  3. OpenAI launches GPT-Live, full-duplex voice models replacing ChatGPT's Advanced Voice Mode ★★★★
  4. ChatGPT Voice gets plugins and moves into ChatGPT Work: spoken requests can now drive agent tasks ★★

Sources (5)

id: 2026-07-09-chatgpt-work · updated 2026-09-29 · open in the interactive timeline