I should probably be more excited about OpenAI DevDay.
I have been an OpenAI fanboy for basically this entire ride. I have paid for the expensive plans, run stupid amounts of work through Codex, built products on OpenAI models, and defended Codex as one of the best values in AI development even during periods when Anthropic clearly had the stronger model for certain kinds of work. So when I say I am not particularly excited about DevDay this year, it feels a little strange coming out of my mouth.
There will almost certainly be new agents, new platform capabilities, faster inference, better tooling, maybe some new ways to let models run around doing things without asking permission every five minutes. All of that is potentially useful. I will absolutely pay attention to what gets announced.
I am just having a hard time getting excited about it.
And the reason is not that OpenAI's models are dumb. Quite the opposite. The problem is that they are incredibly smart, right up until the moment when they suddenly make a decision so stupid that you start wondering whether somebody secretly swapped the model out for a Roomba.
That distinction matters a lot more than I think most of our current AI benchmarks capture.
The Ceiling Is Not the Problem
GPT-6 Sol can be brilliant. Astra can be absolutely fucking spectacular when it is on. Google's models can demonstrate levels of intelligence that would have sounded ridiculous a few years ago. The ceilings are incredibly high.
The problem is the floor.
You can give one of these models a reasonably complicated problem, watch it understand all of the nuance, produce something genuinely impressive, and then ten minutes later watch the same model make a decision that seems almost inexplicably bad. That inconsistency is much more frustrating than a model that is simply less capable overall, because the better result has already established what you know the model is capable of doing.
I have seen this exact pattern before in an entirely different field.
Years ago, when I was doing performance management work, one of the strongest indicators that users were going to start complaining about an application was not simply that the application was slow. People can tolerate slow surprisingly well when it is predictable.
What they absolutely hate is variance.
I believed this so much that it became part of how I pitched the product I was selling at the time. Yes, we track standard deviation. Or, as I liked to put it, the pissed-off user metric. I said that to some of the most powerful people in the Fortune 500. Not one of them batted an eye. Every one of them agreed.
If an application consistently takes fifteen seconds to respond, users may complain about it, but eventually they understand what fifteen seconds feels like. If an application averages two seconds but sometimes responds in half a second and sometimes takes six, people lose their fucking minds. The system feels unreliable even though its average performance is objectively much better.
There was a relationship I used to watch pretty closely: once the standard deviation started getting to more than roughly twice the average response time, complaints tended to start showing up. The upper end did not even have to be catastrophically slow. The inconsistency itself was enough to make the application feel broken.
AI models have essentially the same problem.
If I use a smaller model like Luna and it makes a weird choice occasionally, fine. Luna is cheap, fast, and remarkably capable for what it costs. My expectations are calibrated appropriately. Its peak capability is not supposed to be the best in the world.
But when a frontier model demonstrates that it understands an incredibly complicated architectural problem and then screws up something painfully obvious, that failure feels much worse. I already know what the model can do. I just watched it choose not to do it.
Reasoning levels make this even weirder. Luna on low is basically useless, but it gets better almost linearly as you give it more time to think. Astra does great on x-high and then gets worse when you turn on Max. Fable did the same thing in a lot of my runs. More time to think, worse decisions. That is the floor problem showing up exactly where you would expect it to go away. These days I do not run Max on anything except Luna, because Luna is the only one where I can predict what the extra thinking buys me.
Anthropic Raised the Floor
This is where Opus 5.5 has changed the equation for me.
I am not even convinced that Opus 5.5 has some dramatically higher theoretical intelligence ceiling than everything else. That is not the part that has impressed me most. Anthropic seems to have raised the floor, and they raised it by a lot.
The difference shows up in things like taste, judgment, organization, restraint, design decisions, information architecture, and the general ability to avoid doing stupid shit. Those sound like soft qualities until you actually build software for a living. Then you realize those qualities are what separates technically correct output from something you would actually ship.
I recently had an idea for gamifying structured business networking. A lot of networking organizations still run their meetings from a PowerPoint. Members trade referrals, schedule one-on-ones, give testimonials, and report closed business. Some organizations are still doing portions of this on paper because apparently 1997 remains undefeated.
The idea was to digitize the whole thing. Members log activity from their phones. Referrals get tracked through their lifecycle. Chapters can finally see who is participating and what the network is actually producing.
Then I wanted to make the meeting itself alive. Somebody logs a one-on-one during the meeting, and a Twitch-style notification pops up over the presentation. Somebody reports that a referral turned into $25,000 in business, and it shows up. Leaderboards and points move in real time. Instead of a slide showing last quarter, the presentation shows what the chapter is doing right now.
That means an Electron overlay app for Windows and macOS, a mobile-friendly app, backend services, real-time messaging, analytics, demo data with fake people, and branding. It is not the Manhattan Project, but it is definitely bigger than a bread box.
I gave Opus 5.5 a reasonably mature specification, and about four hours later the thing was basically there. Not "there" in the AI demo sense where somebody proudly shows you a login screen and three buttons that do nothing. I mean there. Coherent product, coherent design, working systems, reasonable architecture, all of it.
Could Astra build it? Probably. I suspect Astra could get me something passable in about the same four hours.
And then I would probably spend the next three weeks refining it.
That difference is enormous.
The Refinement Phase Is Where AI Gets Expensive
The expensive part of AI-assisted development is increasingly not generating code. Code is cheap. The expensive part is everything that happens after the model technically satisfies the requirement but somehow still misses the point.
"No, that interaction makes no sense."
"Why did you put that there?"
"Those two things are obviously related."
"Stop changing unrelated files."
"This technically works, but nobody would ever want to use it."
"Why is this component 700 lines long?"
"That animation looks like shit."
"You solved the requirement and completely missed the product."
That is where human time disappears.
A model that can collapse the refinement phase has a massive economic advantage over a model that simply produces more tokens.
I built reporting into Faultline specifically to measure this across real incidents. For each one, I can see which model and reasoning level did the fix, how complicated the change was, how many follow-ups it took in the same thread before the pull request was merged, and what the inference cost was.
That gives me three useful dimensions: change complexity, follow-up, and cost. There is a very obvious corner of that chart where you want your models to live. High complexity, low follow-up, low cost.
The early numbers point the same direction as everything else in this post. I am not publishing them yet. The harness only recently got consistent about how reviews and follow-ups happen, and I want a clean window before I put numbers behind my name.
And hilariously, building the dashboard that measures this became its own demonstration of the problem.
Astra built the first version. It was fine. Technically competent. It used a 3D visualization, the filters worked, and all of the information was there. But it just did not feel right.
One small example explains the problem pretty well. Astra separated the model selector and the reasoning-effort selector, then put another unrelated filter between them. There is nothing technically invalid about that. It just makes no sense from the perspective of somebody actually using the thing.
Reasoning effort is directly associated with the model. Those concepts belong together.
Opus 5.5 took another pass and combined the model and the reasoning levels available for each model into a single hierarchical control. Of course it did, because that is how a human actually thinks about the data. The visual itself also went from "this is technically a chart" to something striking, obvious, and immediately useful.
Astra implemented the fields.
Opus understood the information.
That is the difference I keep seeing.
Tokens Are Not the Product
For a long time, one of my strongest arguments for Codex was simple economics. The subscription was absurdly subsidized. I could generate an enormous amount of API-equivalent inference for the subscription price, and even when Anthropic occasionally had the stronger model, OpenAI was often still the better overall deal.
I have actual data across my fleet now. The amount of subsidized inference available from both major providers has roughly been cut in half from the crazier early days. OpenAI still arguably lets me squeeze more raw inference out of the subscription.
I just no longer think raw inference is the metric that matters.
Tokens are not the product. Completed work is.
What matters is shippable software. What matters is billable output. What matters is how much useful work gets completed before I have to intervene and start correcting the model.
The formula I care about now is much closer to shippable work divided by subscription cost, or maybe shippable work divided by subscription cost plus the value of my time spent correcting the model.
That second version gets ugly very quickly for inconsistent models.
I track the cost of every incident in Faultline. I have watched Sol take a swing at a fix for two to four dollars, miss, try again, miss again, and miss a third time. Then Fable picks it up, spends twelve dollars, and gets it right. I never touch it again. That has happened more than once. The cheap runs were not cheap. Three misses plus the real fix cost more than starting with Fable would have, and that is before you count my time. My time is considerably more expensive than the tokens.
That changes how I think about the economics of Codex versus Anthropic pretty dramatically. OpenAI may still be giving me more subsidized inference, but if Anthropic gives me substantially more work that I can actually ship and bill for, Anthropic is the cheaper platform in the only way that really matters.
This Is Why DevDay Feels Weird This Year
So now we come back to DevDay.
OpenAI may announce a persistent agent. Great. Everybody has an agent. They may announce faster inference, better application hosting, managed runtimes, sandboxes, tools, skills, computer use, background execution, and a partridge in a fucking pear tree. All of that may be genuinely useful.
But all of it sits downstream from the thing I currently care about most, which is the quality and consistency of the thing doing the thinking.
If the model occasionally makes terrible decisions, giving it more autonomy does not solve the problem. It gives the problem more leverage. The more agentic these systems become, the more important consistency becomes, because one bad judgment can compound across an entire workflow.
A bad response in ChatGPT is annoying. A bad judgment at step fourteen of a thirty-step autonomous workflow can poison everything that follows it.
That is why I do not need OpenAI to show me another model that scores three points higher on somebody's benchmark. I need OpenAI to raise the floor. I need fewer bizarre tool choices, fewer inexplicable instruction misses, fewer unnecessary rewrites, better information architecture, better design judgment, better taste, and better restraint.
I need a model where the bad days are no longer very bad.
Because Anthropic has apparently figured out how to do a lot of that with Opus 5.5, and for the first time since I started seriously using these products, I am considering changing how I allocate subscriptions across my team.
I may downgrade a Codex Pro account and put that money toward another Anthropic Max subscription. That sentence feels bizarre coming out of my mouth because I have never really wavered on Codex being the better value. Even when Anthropic had a temporary lead here or there, OpenAI was close enough, and Codex was useful enough, that the overall argument was still easy to make.
Now I am not so sure.
Codex was unavailable for a while recently, and I barely noticed. I was already using Opus for most of the work that mattered, and frankly, I was having a better time.
That should probably scare OpenAI more than any benchmark.
The problem is not that a competitor temporarily has a smarter model. Models leapfrog each other constantly. The real problem is preference migration. It is when developers stop reaching for your product first, when the important work starts going somewhere else automatically, and when "give it to Opus" becomes muscle memory.
OpenAI is absolutely capable of swinging back. They have the infrastructure, money, talent, distribution, ecosystem, and product capability to do it. I am certainly not writing them off.
But AI time moves really fast.
Fable was not released that long ago, and it already feels ancient. Anthropic has now had multiple generations to reinforce the same advantage in taste, judgment, and consistency, especially on front-end and design work, while OpenAI has not really had its counterpunch yet.
Maybe tomorrow is the start of it.
If OpenAI shows something that materially improves how its models understand the work around them, make decisions, maintain context, and behave consistently over long-running tasks, I will be paying very close attention. I would love that. I genuinely want OpenAI to make this difficult for me again.
But if tomorrow is mostly agents, hosting, faster inference, and orchestration, then cool. I will read about it. I will probably use some of it.
I am just not particularly excited.
OpenAI is down, but it is absolutely not out.
Right now, though, it is down pretty fucking hard.
