← Dispatch archive
September 26, 2026

Claude's New Model Rescued a Failed Coding Job. Then Someone Checked How Much It Eats.

Sources:Grok (X search)Perplexity APIClaude Opus 5.5 (writer pass)

Theo's code rewrite stalled with two rival AI models, then Anthropic's Opus 5.5 finished it in 10 hours. A benchmark crew says there's a catch.

Anthropic released Claude Opus 5.5, its new top-tier coding model, on September 22. Four days later, builders are loving it and the skeptics have brought charts. The rescue story is the best part, so we'll start there.

▸

Two Rival AIs Stalled on His Project. Claude Finished It in 10 Hours.

Theo was moving a big codebase from one programming language to another. OpenAI's GPT-5.6 Sol stalled at about 35% of tests passing, and OpenAI's newer GPT-6 Astra stalled at about 85%. Opus 5.5 got the whole thing working in 10 hours, then spent the next day making it faster. He says usage felt "practically unlimited," which is why he finally gave Claude a job this size. see Theo's post

▸

The Benchmark Crew Showed Up With a Bill

Bridgebench tested how many tokens (the units AI companies charge by) each model uses to finish the same coding task. Opus 5.5 was the hungriest model on their chart, at over 4x what GPT-6 Astra used. A cheaper price per word doesn't help much if the model writes four times as many words. see the chart

4x+how many more tokens Opus 5.5 used than OpenAI's GPT-6 Astra on the same task, per Bridgebench
▸

One Reviewer Wrote 600,000 Lines in a Day, and Hit Several Weekly Limits Doing It

Jack Vinijtrongjit gave it about 24 hours and says he can trust Claude with code again for the first time since Opus 4.6. His praise: it runs things to check them instead of guessing, and it can stay on a long task. His complaints: it burned through several weeks of usage limits in one day, it's weak at big-picture design, and it still ignores his written instructions. read his full review

▸

Great at Finding Bugs, Worse at Avoiding Them

Bubo shared results from Deloitte's code-review tests. On its lowest effort setting, Opus 5.5 caught 72% of the planted bugs, and the old Opus 5 caught 56% even on its highest. The same tests found 44% more concurrency bugs (the errors that appear when many tasks run at once) in the code Opus 5.5 writes itself. Good reviewer, sloppy author. read more ↗

▸

Eight Copies Running at Once, Because Why Not

0xSero ran eight Claude sessions side by side for two days because it was cheap enough to do it. Then he dropped his local AI models (the free ones that run on your own computer) and switched mostly to Claude. He calls this a bigger jump than the one from Opus to Anthropic's top model, Fable. read more ↗

▸

A Pelican on a Bicycle, a Shooter Game, and a 3D Kart Racer

On launch day, riba2534 posted the classic AI test of drawing a pelican riding a bicycle, and it got 189,000 views. Then they open-sourced it along with two more games Opus built in one try each: a CrossFire-style shooter and a 3D copy of the kart racer QQ Speed. play with the code

▸

Think It Got Lazy? Check One Setting.

Anthropic quietly moved the default effort setting from high to medium. Many people who called it "lazier" were still on medium. @notomarsol and @pingrishabh say medium or high is where the gains show up, at roughly a quarter of the old cost. Also, stop telling it to "think step by step." Users report that old habit now makes it worse. read more ↗

▸

20% Cheaper on Paper

The list price fell from $5 to $4 per million words of input and from $25 to $20 for output, and repeat reads got much cheaper. See the Bridgebench chart above before you celebrate. Anthropic published its own math on what a typical task costs. Anthropic's cost breakdown

Featured Tweets

00xSero @0xSeroLocal AI is still 6 months behind. Opus-5.5 singlehandedly moved it back up from 3 months.view post ↗AAddy Osmani @addyosmaniOpus 5.5 is my new daily driver. Fable-level on most work, cheaper and faster than Opus 5.view post ↗AAddy Osmani @addyosmaniA 40-second "how browsers work" animation, with every frame drawn in JavaScript by Opus 5.5.view post ↗TTheo @theoMy TypeScript-to-Rust port stalled at ~35% with GPT-5.6 Sol and ~85% with GPT-6 Astra. Opus 5.5 got it working in 10 hours. The usage is practically unlimited.view post ↗MMichael Fenech @Michael_Fenech_Less than a day in, the thing I noticed wasn't the benchmark. It seems to track the actual objective, not just the prompt. I have to explain less.view post ↗

The big unanswered question is whether the token appetite cancels out the price cut. More people are running full workloads on it this week, so check back tomorrow for the first real bills.

Sources