A Faster Opus Might Be Enough To Achieve Liftoff

Who needs Fable?

If you’ve been living under a rock, you might not have heard that access to Anthropic’s latest frontier model Fable has been placed under export control by the orange man and his fellow cronies, so basically nobody can use it. This is interesting for very many reasons, and raises questions such as:

  • Can anyone trust any US-based company ever?
  • Does this lead to AI trillionaires buying their own small island nations and capturing large parts of the world economy with unrestricted AI research?
  • Does the orange man know what a fable even is?

and more serious questions such as

  • Is Opus/Fable going to be the open-access capability maximum?
  • Will smarter models be available only to Blessed Individuals and Organisations?
  • Will economies other than the US and China manage to put together a competing intelligence provider before one of the current players releases something that takes all the air out of the room or the floor falls out from the market?
  • Given the huge amount of expense involved in training frontier models, if there can be no faith in the ROI (eg, models may be capriciously removed from public access) will anyone ever train frontier models again?

In the brief window that I had access to Fable, I had the perfect problem for it to solve. I’d come across a small, explainable, unsolved computational geometry / sequence problem in the OEIS with only the first 5 values proven and only the first 8 conjectured at. A few weeks prior, I’d let Opus 4.8 have a run at it for a few hours to see if it could solve the sequence and prove its results. Opus pretty quickly did a neat job of coming up with what appeared to be the correct sequence, but it couldn’t prove it in a reliable and particularly believable way given about 20 hours of runtime and several dozen subagents and approaches and so forth. I filed the project under “for smarter models.” My feeling was that Opus’s result was Good Enough For Government Work but not really Good Enough For Science. I was pretty confident in its sequence and working out, but it just couldn’t connect a few dots for me, and I don’t think anyone would have published a paper based on this work.

When Fable released a few weeks later, I fed it the same problem, and let it look at Opus’s work as well. About 30 minutes later, Fable had beautifully decomposed the primary logical hole that Opus had been stuck on into a set theory problem and (to my amateur eye) fully and believably proved the sequence. For bonus points I got it to verify the proof in Lean, which admittedly took quite some extra time due to the lack of computational geometry lemmas in Lean.

I still haven’t published Fable’s work anywhere, and I probably won’t. I don’t want to contribute to the massive AI slop tsunami actual researchers are drowning in. Also, the problem at hand has utterly no utility to anybody, so nothing is lost by not publishing, and an actual mathematician can still have the joy of solving it by hand some day.

Anyway, for me, this was strong evidence that Fable absolutely was more intelligent than Opus. As frontier models gain intelligence their improvements become much, much harder to pin down and it can be quite hard to find things that a frontier model can help you with that a frontier-minus-1 model can’t help you with. (Sidenote: I think this is why some model releases talk about how much more ‘autonomous’ a model is - to me, autonomy and intelligence are related but not the same, but ‘autonomy’ is easier to measure (and easier to train into a model)).

This reminds me of Paul Graham’s essay about Lisp and programming languages, and how it is very hard to look ‘up’ the power continuum:

As long as our hypothetical Blub programmer is looking down the power continuum, he knows he’s looking down. Languages less powerful than Blub are obviously less powerful, because they’re missing some feature he’s used to. But when our hypothetical Blub programmer looks in the other direction, up the power continuum, he doesn’t realize he’s looking up. What he sees are merely weird languages. He probably considers them about equivalent in power to Blub, but with all this other hairy stuff thrown in as well. Blub is good enough for him, because he thinks in Blub.

I think this is quite relevant for measuring ‘intelligence’.

Something I find interesting in this is that Opus was Good Enough For Government Work. It was correct in its result, but it failed to believably prove it was. Even if Fable is never released from its export controls prison, Opus might still prove to be intelligent enough for nearly everything you wouldn’t need a Von Neumann or a Hawking or an Einstein for. Especially, and this is what I’m really trying to get at, if it is delivered vastly faster than it is today.

For models, thought moves in terms of tokens. If the token rate of a relatively intelligent model gets scaled up to a high degree, what is the horizon of problems they can solve? What can they NOT do?

I am beginning to think that inference speed may be or might become the true capability boundary, rather than ‘intelligence per token’, once a given ‘intelligence per token’ is reached. What could you do with Opus if Opus could think at 500 tokens per second, or 1000, or 10000 tokens per second? Even more, what could Opus do with itself - if dozens of subagents are able to be spawned, output thousands of tokens of thought and an output, within seconds, what are the boundaries?

Related: https://www.anthropic.com/institute/recursive-self-improvement