TL;DR

The short version

An Every team ran the same UI prompt through six frontier models. Five converged on the same dark-mode wireframe. One built for the human reading it.

The interesting part isn't which model won — it's that models now have distinguishable taste and risk appetite, and you can select for it. Model choice has quietly become a taste decision, not only a capability one.

Built on a design-review clip from Every — one in-house team's subjective read, labeled as such throughout.

The tell: they could name the model by its look

An Every panel put GPT-5.6 up against Fable, Claude, Sonnet 5, Tara, and Luna on one design brief. Fable, Sonnet, Tara, and Luna all shipped dark mode — functional, complete, and, in the reviewer's words, "actually pretty hard to parse."

Then a smaller detail did more work than the verdict: the reviewer said they can spot "Claude burnt orange glowy suns a mile away." Not good or badrecognizable, blind, by its default palette and layout habits.

That's the real signal. We've crossed from "which model is more capable" to "which model has a house style." Output from different models is starting to carry a fingerprint, the way a designer's work does.

Convergence is the default; distinctiveness is the exception

Five of six models landed in roughly the same place. That isn't coincidence — it's what optimization toward a safe average produces. Tuned to satisfy the most raters, models regress to the competent middle: dark mode, dense, technically correct, forgettable.

GPT-5.6 was "the only one that ever did a different design, like took the risk of doing something different." It used color that carried meaning — red for error, yellow for warning, no legend required — and more white space, which the reviewer found far more usable. Same functionality score as Fable. The preference came down to one model being willing to be distinctive instead of safe.

For an operator, that's the useful frame. Most models hand you the median answer. If you want a non-obvious one, you now have to select for it — because the default setting is convergence, and convergence is invisible until you put six outputs side by side.

Human-to-model vs model-to-model

The sharpest line in the review was about who the output is for. Fable, the panel said, is "exceptional at managing agent-to-agent work" — a clean handoff that expects you to stay out of its way. GPT-5.6 "feels like a power teammate."

It feels like it has optimized the human-to-model interface, whereas Fable has optimized the model-to-model interface.

Every — GPT-5.6 vs Fable design review

That distinction may matter more than raw capability for a while. Some work is agent-to-agent: one model's output is another model's input, no human ever reads it, so density and correctness win. Some work is human-to-model: a person has to look at the result and act on it, so legibility and "theory of mind for the user" win.

Picking the wrong one is a real cost. Route a human-facing design job to the model that optimized for machine handoff and you get something technically complete that nobody wants to use. Route an internal pipeline to the model that spends tokens on white space and you're paying for polish no agent will read.

What to take from a 3-minute opinion clip

Keep the honesty gate up: this is one team's subjective taste, on their own prompts, with no blind protocol and no external data. They say so — "it doesn't necessarily mean it's everybody's." Don't turn it into a ranking. But the operating lesson holds regardless of whose eye you trust:

The models are developing taste. Your job is to know what taste you're buying, and for whom.

Related Articles

Sources