Theo spent six figures in tokens testing OpenAI’s 5.6 Sol model to see whether if its better than Fable and what OpenAI have done to make it even better than 5.5. They breakdown why they both moved their agents to Linux boxes, how to actually burn $65k on a single loop, and the Codex vs Claude Code subagent gap that’s now bigger than the model gap itself.
Anthropic Fable is better than GPT 5.6 Sol. GPT 5.6 is better than GPT 5.5 and better than Opus 4.8. GPT 5.6 is the best of last generation.

Theo from t3.gg/T3 Code and Ben) on OpenAI’s GPT-5.6 Sol (the flagship of the GPT-5.6 family, released in limited preview late June 2026).
Overall Assessment
Very high quality for what it is. Not a benchmark-heavy review — it’s experiential, workflow-focused analysis from people who actually burned through massive usage (six figures in tokens, with one run alone costing ~$65k). A credible deeo view of how models behave in long, messy, real-world coding/agent loops, which benchmarks often miss.
GPT-5.6 Sol is a massive step up from GPT-5.5 (especially on long-running tasks and determination).
It is excellent at execution (“rottweiler that grabs the problem and doesn’t let go”).
It is not quite as smart/discerning as Anthropic’s Fable 5 (wise owl that thinks wider and asks better questions).
It shines in agentic workflows but still has rough edges (especially frontend).
One of the strongest parts is first impressions. The regression when switching back to 5.5 is real and widely reported — 5.5 suddenly feels broken once you’ve used Sol.
Long-running capability and not stopping prematurely is a genuine, consistent upgrade.
Subagent orchestration is meaningfully better.
Frontend is still weak/terrible (generic, over-vomiting elements, callout spam) — this complaint is extremely common.
Good at spatial/3D reasoning and CLI tools (Blender, etc.) — plausible and aligns with agentic strengths.
Code-Quality Gap (16:38–26:38)
Fair and nuanced. Better than 5.5 overall, but still overcomplicates APIs/SDKs and loves tests excessively.
Lacks Fable’s discernment (better at sniffing intent, asking good questions, avoiding bad early assumptions).
They correctly position Sol as “bottom of top-tier” / peak of the current generation, while Fable feels like the start of the next one. The analogy is PS3 vs early PS4.
Use Sol when you want something extremely determined that will just grind through long, complex tasks and subagent orchestration without stopping or lecturing you.
Use Fable when you want better discernment, taste, strategic thinking, and fewer rabbit holes.
The real power move is using both in hybrid workflows.
They identify the harness/tooling difference (Claude Code’s workflows vs Codex) can matter as much as or more than the underlying model. The ability to have models orchestrate across providers (5.6 calling Fable or vice versa) is the future.
Codex subagent UX vs Claude Code is a bigger gap than the raw models in many workflows. Claude Code’s dynamic workflow primitives (JS files with stages) give more flexibility than Codex’s more rigid tool-call-based system.
Fable corrects course better, Sol just powers through.
Sol has insane token burn via heavy subagent loops.

Brian Wang is a Futurist Thought Leader and a popular Science blogger with 1 million readers per month. His blog Nextbigfuture.com is ranked #1 Science News Blog. It covers many disruptive technology and trends including Space, Robotics, Artificial Intelligence, Medicine, Anti-aging Biotechnology, and Nanotechnology.
Known for identifying cutting edge technologies, he is currently a Co-Founder of a startup and fundraiser for high potential early-stage companies. He is the Head of Research for Allocations for deep technology investments and an Angel Investor at Space Angels.
A frequent speaker at corporations, he has been a TEDx speaker, a Singularity University speaker and guest at numerous interviews for radio and podcasts. He is open to public speaking and advising engagements.