Brett Chereskin
← Back to writing
AI & WorkMay 12, 2026 · 5 min read

Opus 4.8 Is My Daily Driver. Then Fable Cracked a Problem It Couldn't.

Share

Most weeks I don't write about a model release. The cadence is fast enough now that if I reacted to every one, this blog would be a changelog. But two things happened close together that are worth sitting with — not because of benchmarks, but because of what they told me about the shape of the work I do.

Anthropic shipped Claude Opus 4.8, a clear step up from 4.7. And for a brief window, Fable — a long-horizon model — showed up in the Claude family. I got to use both on a real, ugly piece of work. The contrast taught me something I haven't stopped thinking about.

First, Opus 4.8 as a daily driver

Opus 4.8 is excellent, and the jump from 4.7 shows up where it matters — in ordinary work, not demos. Fewer dropped threads in normal back-and-forth. Tighter reasoning when I push it on a decision. Less hand-holding to keep it on the rails. When I'm drafting, debugging something I directed it to build, or pressure-testing an idea before a meeting, 4.8 is what I reach for without thinking. That's the highest compliment I can give a tool — it disappears, and I just get to work.

For 95% of what an operator does in a day, that's the whole story. If you've been on 4.7, take the upgrade; you'll feel it inside an afternoon.

But this post is about the other 5% — where I hit a wall, and what happened next.

The task that broke the pattern

Part of my job as COO is figuring out what's actually working in how we reach people — which touches matter, in what order, and how they add up across a long, winding path before anything happens. Marketing attribution: multi-touch, many channels, many stages, each step depending on the assumptions and math of the step before it.

I'll keep the specifics generic, because the data isn't the point. The shape is the point. This wasn't one hard question. It was a long chain of medium-hard questions, each building on the last, where losing the thread anywhere in the middle quietly corrupts everything downstream. You don't get a clean error. You get an answer that looks plausible and is wrong because something twenty steps back got dropped.

Deep into that chain, Opus 4.8 — the model I just spent three paragraphs praising — started to strain. Not failing loudly. Slipping. Re-deriving a number it had already settled, or quietly contradicting an assumption we'd locked in early. Every individual step was fine. Holding all of them coherent at once, across the full length of the problem, is where it wobbled.

The failure wasn't intelligence. It was endurance. The problem wasn't too hard — it was too long. And those are completely different things.

Enter Fable

"Long-horizon" sounds like jargon, so here's the plain version: Fable is built to sustain coherence across very long, multi-step tasks without losing the thread. Not smarter in a single leap — built to hold the whole chain in view from start to finish.

I gave it the same sprawling attribution problem. And it just... held it. Assumptions locked early stayed locked. Numbers derived in step three were still intact in step thirty. It didn't re-litigate settled ground or wander off the spine of the work.

That was the wow moment. Not a flashier answer — a finished one, on a problem where "finished and still coherent" was the entire challenge.

Claude Opus 4.8

Outstanding general flagship and my daily driver — sharper and steadier than 4.7 for drafting, decisions, building, and the everyday back-and-forth. Where it strained: holding one giant, multi-stage chain coherent end to end.

Fable (long-horizon)

Built to sustain reasoning across very long, multi-step tasks. Where it shone: carrying the entire multi-touch attribution build start to finish — assumptions and derived numbers intact across the full length of the problem.

What this means for operators

Here's the trap to avoid: thinking about models as a ranked list. Better, worse. Newer, older. Pick the one at the top.

The right question isn't "which model is best?" It's "what shape is my task?"

Most of my work is a series of self-contained moves — answer this, draft that, debug this, decide that. For that shape, a strong general model like 4.8 is exactly right. Reach for it by default. But some work is one enormous connected chain, where the whole value is in holding it together across length: detailed modeling, sprawling analysis where step thirty depends on step three, long planning where an early assumption has to survive to the end. That shape rewards endurance — and that's what a long-horizon model is for.

Stop asking which model is smartest. Start asking what shape your task is. Short and self-contained, or long and connected? The answer tells you which tool to grab.

The honest take

I won't pretend a single problem is a benchmark — this is one data point from one operator on one ugly task. But it was a real glimpse of where things are heading, and once you've seen a long-horizon model carry something your daily driver couldn't, you can't unsee it.

So: Opus 4.8 is the upgrade you should take, and it's where I'll spend most of my hours. Long-horizon models are the thing I'm now actively watching, because they crack a specific kind of problem — the long, connected, easy-to-corrupt kind — that even an excellent general model wrestles with. Different tools for different task shapes. That's the whole lesson, and it's more useful than any leaderboard.

If you've hit the same wall — a model that's brilliant in bursts but loses the plot on something long and sprawling — I'd like to hear how you handled it. Drop it in the comments, or reach out through the contact page.

Comments

Loading comments…

Leave a comment

0/2000

— Read next
AI & WorkJune 12, 2026 · 5 min read

AI SEO Is the New Visibility Game. Here's How I Picked a Tool For It.

People are asking AI instead of Googling. That quietly rewrites the rules of being found. Here's how I evaluated Profound versus Athena for work — and why we went with Athena.

AI & WorkJune 5, 2026 · 5 min read

How I Used Claude to Fight a $600 Insurance Denial — and Actually Filed a Regulator Complaint

A routine visit to a specialist turned into a billing mess and a denied claim. Most people give up at that point. AI is the reason I didn't — and why I filed a formal complaint with the state.

AI & WorkMay 20, 2026 · 5 min read

I Tried the Big AI Note-Takers. I Keep Coming Back to Granola.

I ran the major AI meeting-notes tools through real work — including a head-to-head with the Gemini note-taker built into Google Meet. One quietly won. Here is why Granola earned the spot.