Gemini 3.8 is... Good?
September 5, 2026·13 min read·Matthew Bradford

Gemini 3.8 is... Good?

Gemini 3.8 FlashDeepSWELLM codingAI benchmarkssoftware engineering

In the last article, I said that complimenting Gemini's ability to produce valid tool calls might be the last time I complimented Gemini. Perhaps ever.

Well, shit. Now I have to write a follow-up.

Gemini 3.8 Flash is better. Substantially better. I put it back into the same difficult repositories where 3.5 Flash had fallen apart, and it solved problems the old model could not solve. It understood relationships across files. It recovered from mistakes. Some of the work was excellent.

The people at DeepMind built a much better model. They should feel good about it.

But… this is a Google model. I also watched it work… and well… yeah.

3.5 is based on a VERY smart model. That hasn’t changed. 3.8 is also based on a VERY smart model. The problem with 3.5 was that even when it solved problems it was less like a skilled engineer and more like a marble falling down a funnel. Gravity brought it to the result, each action on its own was perhaps fine. In the context of the previous actions and what should come after… well, the story changed. There was little evidence it propelled itself to the answer… it just sort of found itself there and if it was lucky it looked up and saw it succeeded and stopped.

3.8 is much better. It propels itself towards a solution. It can explain the wall, identify what is behind the wall, and write a surprisingly good test for the wall.

Then it spends another twenty calls checking the wall.

If GPT-5.6-Sol is good because it has a brand of autism that works well for long horizon coding tasks, Gemini 3.8 is a super smart developer with the hyperfixation of someone with ADD who forgot to take their meds. And look... that focus can be a superpower. Watch these runs and you can see how it helps Gemini solve difficult problems, then keeps it working long after another test is useful. Google got it to the top of DeepSWE. Now I think they have some tuning left to make it faster and more pleasant to use.

I brought back the problems it hated

I did not repeat the entire experiment from June. That one involved hundreds of trials, controlled tasks, recovery tests, and was generally annoying. This time I reran six official DeepSWE tasks where Gemini 3.5 Flash had recorded zero strict successes.

Everything else was the same though... Same repositories. Same tasks. The same Pier and mini-swe-agent route. One fresh attempt at each with Gemini 3.8 Flash through OpenRouter, using its medium reasoning default and a 200-action cap.

These were selected failures, so this is not a general pass rate. DeepSWE provides the broader comparison. I deliberately took the thing back to the places where it had embarrassed itself because I care less about the pass rate and more about how it gets there.

Same six tasks Gemini 3.5 Flash Gemini 3.8 Flash
Artifacts that passed the official verifier 1 3
Strict passes, including the execution protocol 0 2
Tool calls 678 1,066
Total tokens 52.3 million 137.1 million
Recorded cost $20.43 $18.31

The Cattrs task passed every official test. Its run also contained command timeouts, which fail my stricter protocol check. I am counting the correct code as correct code. Calling that an outright coding failure would hide part of the improvement.

And three verifier passes with two strict passes happens to match what GPT-5.5 managed on this same subset in the previous experiment. Gemini earning those passes means something. (See? I am at least trying to be nice here.)

Adaptix is the clearest win. The task involved alternate names for fields, including collisions and the unpleasant little edge cases that appear when several individually reasonable rules interact. Version 3.5 missed six feature tests. Version 3.8 passed all 44, along with all 2,738 tests checking preserved behavior.

That requires understanding how the feature fits together. You don’t get there by writing a function that looks plausible and hoping nobody asks about the rest of the library.

Anko was another real win. Default function arguments sound simple until you have to make the parser, variable binding, evaluation order, and virtual machine agree about what they mean. The old run hit its cutoff. The new one got through the official verifier cleanly.

It broke some code along the way and repaired it. I am fine with that. Debugging involves discovering that your last idea was wrong. The improvement is that 3.8 can often use that discovery to do something better next.

That was exactly the faculty I found so lacking last time.

The trip is still weird… even if it eventually arrives.

On Adaptix, the model had the full local suite passing at tool call 158. All 2,845 tests.

It submitted at 174. In between, it ran the full suite twice more.

On Anko, the feature test was green at 173 and the full suite was green at 175. It then went looking for more reassurance. Race tests. An unrelated race failure. A stash cycle to investigate that failure. More tests. More inspection. Commit ceremony. Submission at 194.

There is an explanation for each of those actions. That is what makes this version interesting. It is much easier to defend the next individual move than the whole sequence.

It is like it knows we’ve been giving it so much shit and now it has a complex about how much its older siblings sucked. But really though… you can always run another test. You can always inspect another diff. There is always some uncertainty left in a codebase. An engineer has to decide which uncertainty matters to the job in front of them, and whether the next check is worth the time and the risk.

Gemini is getting better at answering the question it just asked. It is still bad at deciding whether it needed to ask.

OPA was the painful example. Its new syntax-tree tests passed at call 187. Its new partial-evaluation tests passed at 195. With almost no budget left, it decided to add another test at the command layer.

That test contained an unterminated Go string.

It ran out of actions looking at the failure it had just created. The official verifier never ran.

I cannot tell you that the code was correct before this happened. Its own tests passing would not establish that. What I can tell you is that it spent its last few actions creating a new problem instead of submitting its best available answer.

The marble has learned to steer. Sometimes it is afraid of the bottom of the funnel.

It is a giant nerd, it loves tests a little too much

The new model writes a lot of tests. Across the five tasks where we captured a final patch, it added 2,145 lines to test files.

Some of those tests were useful. The problem is what happens when the model's tests inherit the model's misunderstanding.

On Textual, it added a 366-line test file and passed all six of its cases. The official verifier then failed five cases involving the public fields of a change message and whether that message should be posted when nothing had actually changed.

On Mashumaro, it passed 32 local cases and still missed a collision between two prefixed child objects. It fixed the old model's alias problem and preserved more existing behavior, but left a different hole.

Writing more tests does not make your interpretation independent of itself. You can build an elaborate test suite around the same wrong assumption that produced the code.

This is where the book-smart description fits. It knows what responsible engineering looks like. Responsible engineers test their work. Responsible engineers investigate failures. Responsible engineers check for regressions.

It can perform all of those activities while losing track of what they are supposed to establish.

Let's chat a bit about the costs

The new run cost less. It also used 2.62 times as many total tokens and made 57% more tool calls. But it did cost about 10% less than that same subset on 3.5-flash.

This is why cost per token tells you so little on its own. On the DeepSWE snapshot below, Gemini averages 166 steps per task. Astra averages 29, with a slightly higher underlying score that rounds to the same 74%. Gemini is much cheaper per task, yes. But everything it does between receiving the problem and handing back the answer matters too. If Google can keep this level of capability while getting those working habits under control, this becomes a much more interesting model. Right now, the cheap tokens are paying for a very long trip.

Sonnet 5 shows where that road ends. On the same leaderboard, at max reasoning, it scores 54%, averages 268 steps and 214,000 output tokens per task, and costs $26.40 per task. Astra gets 74% in 29 steps for $6.52. The model marketed as small and cheap costs four times more per task than the frontier model and solves fewer problems. I do not see a valid use case for it anymore. It is too expensive to be as dumb as it is. Gemini is not there. Google gets away with it by making Gemini smarter and evidently FAR cheaper per token... But it is on the same road, and it needs to solve the efficiency problem if it wants the crown.

The lower bill is welcome. I like spending less money. But a lower price per token does not remove the work of processing tokens, and it certainly does not run the next shell command before the model has decided to issue it.

The six new jobs took a little over an hour, running in three parallel lanes. That includes tools and verification, so it is not a measurement of Google's raw inference speed. It is the time the work took, which is the clock I care about when I am waiting for the work.

Google even acknowledges occasional slowness, timeouts, and increased token consumption in the 3.8 model card. I am not telling you anything new. Just agreeing that "yeah, I saw that too."

Flash is cheap. Flash is not fast.

DeepSWE deserves the credit too

In the previous article, I was hard on benchmarks that give a model a finish line it cannot reliably recognize in ordinary work. DeepSWE is the exception. Its tasks demand substantial changes in real repositories, and its checks exercise behavior. I used those tasks precisely because passing them requires more than looking competent.

On the September 3 DeepSWE v1.1 leaderboard, Gemini 3.8 Flash scores 74%, tied at the displayed score with GPT-6 Astra and Claude Opus 5. That is a fantastic result. The listed configurations differ: Gemini uses high reasoning, Astra xhigh, and Opus max. My six-task rerun used Gemini's medium default. The leaderboard records the settings and results.

My little collection of previously failed tasks points in the same direction. More correct code. Better understanding. Real problems solved. I would be moving the goalposts if I used DeepSWE to criticize Google in June and decided it was meaningless when Google did well in September.

But look at the other columns on that same leaderboard from DeepSWE.

DeepSWE v1.1 snapshot Gemini 3.8 Flash, high GPT-6 Astra, xhigh
Score 74% 74%
Average output tokens 143,000 30,000
Average agent steps 166 29
Average cost per task $2.36 $6.52

The score says Google built something capable. The step and token counts say it takes a long route to demonstrate that capability. The price says Google is selling that route cheaply.

All of those things can be true.

That is why DeepSWE should be taken seriously here. The larger evaluation and the individual traces are telling compatible stories. The benchmark found the improvement. Watching the work explains why the improved model can still be frustrating to use.

Google arrived just in time to see the party switch venues… maybe.

The timing is rough. Google released 3.8 Flash on September 2. By the time I am writing this, Fable 5.1 and GPT-6 Astra are already part of the conversation about what comes next. Anthropic announced Fable 5.1 on September 1. Google barely had time to enjoy its release week.

They got there on the benchmark that, to me, matters most. But I also use these models. Astra is the best publicly available model for computer use, and in my experience it is not a close contest. Gemini matching its rounded coding score does not make the two equally good to work with. UI, computer use, and the general experience of getting something done tip the scales meaningfully away from Gemini still.

What already feels a generation behind is the way Gemini spends its effort. Arriving at the answer is becoming a less complete description of what a good agent does. I also want it to make sensible decisions about the work, including the decision to stop doing it.

For DeepMind, this should be a big win. I think it is one. The researchers made a model that can solve difficult problems at a price where using it extensively might make sense.

For Google, I still find it disappointing. This is the company with the research talent, the hardware, the distribution, and products people already use throughout their day. I want to see that combination produce something the rest of the industry has to chase. Instead, I am relieved that its newest model might be worth using and is keeping up in at least one or two categories.

That is a strange place for Google to be.

It is also a better place than it was in the last article. A cheap model that can produce strong work, given enough time and a bounded task, has a purpose. I can imagine handing this one a job I do not need back immediately and reviewing the result. That is a meaningful change from deciding the free allowance was not worth the aggravation.

I am not giving it my car keys based on a coding benchmark. But I might give it some code.

And look, I did not build this experiment because I needed Google to keep losing. I built it because I wanted to understand why a company that should be making my life easier kept making it harder.

So Google… much better. Would have been amazing in June. Worth talking about now… but it is like you won “most improved” not “best in class.” You can do better. And honestly, I am kinda pulling for you now.

All posts