TL;DR
The short version
In Claire Vo's product-work benchmark, OpenAI's GPT-5.6 Sol beat Anthropic's Fable 5 on the thing that actually matters — getting to shippable user value — at half the price, even though Fable is the more technically precise model.
The durable takeaway isn't "switch to Sol." It's the evaluation itself: rank models by whether they reach shippable value in your workflow, not by raw intelligence. From How I AI.
Built on the 9 Jul 2026 How I AI episode with Claire Vo, who benchmarked GPT-5.6 Sol against Fable 5 and Sonnet 5 across real product work.
The benchmark most leaders miss
Most model comparisons rank raw intelligence. Claire Vo, host of How I AI, ran a different test: which model does she actually want to work with all day.
She built her own eval across the tasks a product leader lives in — writing PRDs, prototyping, coding, and talking to an agent. She ran it across Anthropic's Fable 5, Sonnet 5, and OpenAI's three new GPT-5.6 tiers (Sol, Terra, Luna), had a GPT-5.5 model grade the outputs, then overruled the machine with her own taste on a 70/30 split. GPT-5.6 Sol won by a wide margin.
Theoretically intelligent colleagues don't ship
Her verdict is the line worth keeping:
Fable is theoretically hyper-intelligent and Sol is practically effective.
Claire Vo, How I AI
The distinction maps straight onto managing people. "I really struggle working with theoretically intelligent colleagues who can't get anything done," Vo says. Fable, in her testing, is the brilliant engineer who over-thinks the problem. It fans out, solves genuinely hard things, and scores every risk — then hardens the architecture until it breaks itself. On two of her projects, Fable built tool-calling loops so rigid that only one specific model could run them; Sol looked at the same code, loosened the constraints, and got it working in one shot.
Her reframe cuts against the instinct to buy the highest benchmark: "When you're building products, exact precision is neither helpful nor possible." Precision is a feature for a security researcher scoring exploits. It's a liability when the job is intuition, design, and knowing which constraints to drop to reach a user.
Communication is a capability, not a nicety
The other gap is how the model talks. Sol writes like a person. Fable, in Vo's words, "talks to me like an engineer that has never met a human before. It's like its first day on Earth." That isn't an aesthetic complaint — it's a collaboration cost. A model you can't read is a model you can't work with, no matter how capable it is underneath.
The nuance: Sol isn't the best writer for everything. On agentic voice — the assistant that answers "can you move my meeting?" — Anthropic's Sonnet 5 still sounds the most human. Match the model to the seat.
The price makes it easy
Sol is cheaper than the model it beats — exactly half, on input tokens.
Public API pricing, verified
Cheaper and more effective for product work removes the usual tradeoff. It also lands as Anthropic pulls Fable from Claude subscriptions into metered usage credits, which sharpens the cost gap further.
What to copy
The takeaway isn't "switch to Sol." Benchmarks move monthly, and this one is a single practitioner's taste-weighted eval, heavily tilted toward front-end prototyping — read it as one operator's read, not a verdict.