The Actual Bet
Coding agents already changed how software gets built. The question OpenAI is now asking is whether the same shift can happen for everyone else—the accountants, the ops managers, the VCs assembling investment memos at midnight.
ChatGPT Work is essentially a version of Codex rebuilt for people who have never typed a command into a terminal and have no plans to start. It connects to your email, calendar, browser, and a stack of SaaS tools, then runs multi-step tasks on your behalf without you having to babysit every step.
The commercial logic is straightforward: agents burn more tokens than chatbots, and reaching new professions is how AI labs justify the eye-watering cost of training frontier models.
Why Non-Engineers Are a Hard Unlock
The challenge isn’t the model. It’s the harness—the software layer that decides what the model sees, what tools it can use, and how it presents results back to you.
For developers, a command-line interface was enough to spark a productivity revolution. Most people, though, don’t use CLIs. There’s a reason Windows replaced DOS.
OpenAI’s own non-engineering teams ran into this early. When communications and finance staff tried Codex, it greeted them with empty diffs and questions about code. Not ideal. The team spent months making it more general-purpose, and ChatGPT Work is the result.
The design philosophy borrows from an unlikely place: skeuomorphism. Those early calculator apps that looked like physical calculators weren’t just dated aesthetics—they helped people make the cognitive leap to a new tool. ChatGPT Work uses a similar logic, adding buttons and project selectors that power users don’t need but newcomers do.
“Discoverability matters in this phase,” one OpenAI engineer noted. “At some point we won’t have the button.”
What It Actually Does Well
The use cases getting traction are mostly data-heavy coordination tasks:
- Weekly metrics reports that run themselves
- Spreadsheets that become live planning tools
- Investment memos assembled from scattered communications
- Auto-updating dashboards for publicly traded companies
- Queryable databases built without writing a single line of Python
One test case that worked cleanly: extracting a weirdly-formatted preschool calendar from email and pushing it into Google Calendar. Tedious, repetitive, exactly the kind of task that makes a strong case for agents.
The less glamorous version of the pitch is that most knowledge workers are drowning in information they can’t act on fast enough. ChatGPT Work’s value proposition is that it doesn’t just give you access to your tools—it actually uses them, which is why the broader conversation around AI automation keeps expanding.
Where It Gets Wobbly
Setting up permissions is confusing. Giving an agent read-only access to a cloud drive produced circular error messages until a mobile dialog box finally explained that only full access would work. Many settings only exist on the web app, which means bouncing between devices mid-task.
There are also odd capability gaps—link it to Google Calendar and it can create events, but not new calendars. And effort level settings, which determine how hard the agent tries, aren’t intuitive for new users yet.
The deeper challenge is evaluation. Code either works or it doesn’t. A business strategy, a sales pitch, or a presentation doesn’t come with a pass/fail test. OpenAI uses an internal benchmark called GDPval, drawn from 44 occupations and hundreds of knowledge work tests, but the honest answer is that this is still being figured out.
The Claude Problem Nobody Wants to Talk About
OpenAI’s engineers were notably reluctant to discuss how ChatGPT Work compares to Claude Cowork. The “Mad Men ‘I don’t think about you at all’” deflection came up more than once.
The history here is worth knowing. OpenAI built Codex first, but bet too heavily on the model handling tasks autonomously with minimal user input. Anthropic’s Claude Code took a different approach—checking in with users frequently, offering options, running A/B comparisons. More work for the user, fewer catastrophic errors from the model.
That approach proved more effective. OpenAI eventually followed suit, adding more interaction points to Codex. ChatGPT Work continues that evolution.
Download data suggests Codex has recently edged ahead of Claude Code in demand, though the gap is narrow and partly explained by reported compute constraints on Anthropic’s side.
The Token Math Nobody Mentions
One detail worth flagging: four days of casual use on a $20/month subscription consumed over 80 million tokens—roughly $65 worth of compute by the model’s own estimate. That’s a 3x subsidy on four days of light experimentation.
OpenAI points to an 80% price cut on its Luna model as evidence that efficiency is improving. The bet is that costs fall fast enough to make the economics work before the subsidy becomes a problem. It’s a reasonable bet, but it’s still a bet.
The Actual Takeaway
ChatGPT Work is a serious product aimed at a genuinely hard problem. The model quality is real. The ambition to move beyond software engineers is necessary for the entire AI industry, not just OpenAI.
But the gap between “impressive demo” and “tool that a non-technical person sets up and trusts with their inbox” is still wide. If you’re evaluating it now, the honest framing is: high ceiling, rough edges, best suited for users willing to invest setup time on data-heavy, repeatable tasks.
If you’re waiting for the version that works out of the box for everyone—that’s the product OpenAI is still building.
Comments (0) No comments yet
Want to join this discussion? Login or Register.
No comments yet. Be the first to share your thoughts!