TL;DR

The short version

GPT-5.6 Sol isn't the smartest model available. It's the first one Dan Shipper trusts to run loops of knowledge work on its own — and that reliability, not raw intelligence, is what changes how you actually work.

The move isn't chasing the smartest model; it's managing a system that does the work and hands you decisions. From Every.

Built on the 9 Jul 2026 Every piece by Dan Shipper, who ran GPT-5.6 Sol for a month across coding and knowledge work.

The old move was chasing the smartest model

The instinct is to reach for the most powerful model and eat the cost. Shipper's own numbers show why that breaks down. Fable runs $10 per million input tokens and $50 per output; Sol runs $5 and $30 — cheaper across the board. And Fable is slow. "With Opus or Fable, you're just going to be waiting for a long time."

More power also means more waiting and more spend on tasks that never needed it. Shipper describes telling Fable to do one thing and watching it spin up a fleet of agents until he's suddenly out of credits. Smartest-model-always is the expensive habit. On coding he rates Sol A-tier, not S-tier: for big, complicated tasks he still flips to Fable, which scored higher on Every's internal rewrite benchmark. Sol does the job — it just builds something more complicated than it needs to.

The better move is managing a system that does the work

Shipper's real point is a shift in altitude. Sol is "smart enough, fast enough, and reliable enough" that you can stop doing rote knowledge work yourself and instead tune a system that does it and hands you decisions.

Coders have worked this way for a while: report an issue, an agent does the work, reviews it, ships the PR. Sol is the first general model Shipper trusts to run that loop for non-coding work. An email app turns his inbox into cards, each with a suggested action; he approves or rejects, and it improves over time. The same setup reads meeting transcripts he missed and tells him what he needs to know.

I'm the one who's tuning 5.6 inside of Codex to tell it what to pay attention to, what the next actions are, and then when it presents me with decisions to make, I just make decisions.

Dan Shipper, Every

That's the whole shift: from doing the work to managing the thing that does it.

You don't have to choose one model

The sharpest tactic in the piece is combining them. For a hard task, Shipper drives Fable and has it call Sol as a sub-agent — Fable's reasoning, Sol's speed and cheap tokens, without burning through Fable credits. "They're kind of a match made in heaven."

That reframes the whole "which model" question. The smartest model plans; the fast, cheap one executes. You stop picking a winner and start assigning roles.

Where Sol still loses

Keep the boundary honest. Sol is the better writer of the current field — clearer and more direct than Opus 4.8 or Fable, which tend to over-explain — and it one-shots emails and marketing copy without the usual AI tells. But on design it's a step up from its predecessor and still clearly behind Opus 4.8 and Fable. And on the hardest coding, Fable stays ahead.

Two cautions on the evidence. The benchmark scores are Every's own internal metric — not externally reproducible, and the article and the video don't even agree on Fable's number. And this is a same-day launch reaction from a business that sells early model access, so read the tier grades as enthusiast opinion. The part worth keeping is the lived experience, not the scores.

The bar that matters now

Related Articles

Sources