Rendered at 01:34:14 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
postalcoder 8 hours ago [-]
I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0].
> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].
I don't understand how the K3 numbers keep coming out cheap for people. I recently started to add it to my security auditing benchmarks and found it was going to cost about twice as much as Opus 4.8. It blew through the $100 budget I'd set at like 11%. In the tasks I'm doing it seems crazy expensive because it chews so much, burning a tremendous amount of tokens.
InsideOutSanta 6 hours ago [-]
I think the way people usually compare pricing is fundamentally flawed. You can't compare token prices because different models use different tokenizers, and you can't compare tokenizer-normalized token prices because different models at different settings use more or fewer tokens to complete the same task at a different level of quality.
Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.
SwellJoe 5 hours ago [-]
I got the $19 plan, and it's anemic. One tiny task blew through the 5-hour budget and 19% of the weekly budget. A completely useless amount of usage. OpenAI's $20 plan feels like 100x more generous (I don't think I'm exaggerating here). Someone in another thread said their plans are cheaper in China, maybe that's the difference, I dunno.
But, I'm finding Kimi K3 terrifyingly expensive in the way that Fable and GPT 5.5 Pro are at token rates. Not as expensive as those, but expensive enough to where if you don't put a budget cap on it, you might wake up bankrupt if you leave a task running overnight. Not because of the per-token cost, but because how many tokens it's going to burn.
nullify88 4 hours ago [-]
In the $19 plan, I've been able to reverse engineer both an android APK and firmware (in Ghidra and Radre) for a baby rocker and build a quick PoC application in my session limit. And then further refined the app in another session at another point in time without leaving Opus. I dont consider that to be a tiny task. How are you blowing through your usage?
SwellJoe 4 hours ago [-]
I have no idea. Seems like normal stuff. I used Kimi Code with K3 to add support for Kimi Code to flar (https://swelljoe.com/post/i-let-every-agent-implement-its-ow...), a task I've done with almost every major model/agent combo. Most show up as a blip on the usage chart...it's basically usually one file, a README update, and adding the agent name to the CLI.
Then, I added it to my benchmark of security vulnerability auditing capability, and it burned a bazillion tokens, burned through the 5-hour limit, burned through $100 in extra usage I'd allocated, and was only 11% finished. That's more expensive than any model I've tested other than GPT 5.5 Pro on this task.
These are things I've done with a bunch of other models, I feel like I have a notion of what they ought to cost, and with K3, they end up being crazy expensive. (And it seems to be a function of how many tokens it burns accomplishing the tasks.)
lukan 2 hours ago [-]
I guess the problem is, that claude's 5 h/weekly limit is not consistent, but depends how many other people are using it/how much ressources Antrophic currently has. I did huge amounts of work without hitting the limit - and small tasks at some other times that hit the limit before it completed.
Those who pay for the expensive direct API, get served first.
Barbing 1 hours ago [-]
Not surprising given the way they served super low-quality inference before they acquired compute from Musk.
And not convinced they couldn’t have instead tried the It’s A Wonderful Life strategy (“fam we’re oversold, would some of y’all be OK to limit your usage? We’ll get you back one day!”)
try-working 5 hours ago [-]
I have the second largest Kimi plan, the Chinese version. When K2.6 was their latest model, the quota was good; it was like GPT $100 is now or what the $20 version was in December.
When K2.7 was released, they cut quota by 80%. I can't tell how much they have further cut it after the K3 release because it's barely worth using at all. I just use it in my model router since I have the annual plan paid for.
It's just not a serious model or company.
InsideOutSanta 5 hours ago [-]
Yes, the OpenAI plans are much more generous than both Moonshot's and Anthropic's. It's the only provider of the three where the $20 plan is at all usable for programming.
bg24 3 hours ago [-]
I believe your statement. Labs do not publish subscription vs. api revenue and difficult to guess with no priors.
Subscription is to drive adoption - fixed cost, can adjust the usage eg. give resets, increase quota based on capacity available. We subscribers tend to take it as a mandatory benefit :-) For labs, it is not letting the capacity go waste.
api is the $$ driver - pay per use, enterprises.
Right now, Kimi needs to first hit the subscribers at the level of OpenAI and Anthropic. With the api usage skyrocketing due to K3, it will be clear in a few months on the actual subscription benefits.
KronisLV 4 hours ago [-]
> Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.
For me, the Moonshot 100$ plan felt like it gives me lower total amount of work I can do than the Anthropic 100$ plan (probably within like 30% of each other). Kimi has way more generous 5 hour limits (never hit those once, whereas I do regularly with Opus) but the 7-day and monthly ones are lower. However, with the annual billing, Moonshot's 200$ tier plan becomes way better, because you get it for 159 USD per month.
There's also the odd thing of Anthropic's 100$ plan charging me 108 EUR so seems like their sticker price does not include VAT but Kimi's did, cause I paid like 87 EUR. Wrote down some initial thoughts at https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don... but it's hard to do exact comparisons (even the same task will have way different real token amounts per model).
Still, Kimi K3 is a pretty cool model! On high reasoning, it was pretty close to Opus 4.8 and didn't seem to waste as many tokens as Max.
adgjlsfhk1 5 hours ago [-]
Testing at max effort likely doesn't produce optimal results.
sggyamg 3 hours ago [-]
Can you be more explicit?
adgjlsfhk1 50 minutes ago [-]
Max effort is the way to give the highest perf, but not highest perf/$. Having claude (or other models) use a lower effort can often be 80% as smart but get to the results 10x faster for the problems where it works.
3 hours ago [-]
reinitctxoffset 7 hours ago [-]
[flagged]
idiotsecant 6 hours ago [-]
I feel like I am having a stroke. What is this
throwawayoaky 3 hours ago [-]
looks like the agent-judged results of an agent-built 'eval' based on some examples derived from this person's real work. and clearly part of a larger document. this kind of slop is kind of useful but opus 4 was the first generation that was any good at writing its own prompts/evals/rubrics so there's a certain sloop to it..
reinitctxoffset 5 hours ago [-]
[flagged]
onlyrealcuzzo 6 hours ago [-]
I can't believe they released the charts they did.
It basically shows that Sol absolutely demolishes Fable at every part of the cost curve for coding for the same level of quality.
Opus is competitive. It just has a higher level of quality / higher cost to start.
throw10920 59 minutes ago [-]
Isn't that because Fable/Mythos were tuned for cyber at the expense of general performance?
pixl97 6 hours ago [-]
If fable costs more to run than the markup they still come out ahead.
qsera 8 hours ago [-]
I can't help but read these comments in the voice of a TV commercial....
iambateman 7 hours ago [-]
Ask your doctor if Opus 5 is right for you. Side effects include occasional hallucination, security breaches and unwanted React apps. Some developers have reported receiving entire apps from untrained executives who may or may not know what they’re doing.
Stop using Opus immediately if you experience signs of dizziness or vomiting.
Opus 5…the people’s favorite.
ahofmann 7 hours ago [-]
Spot on! Comment of the month, I'd say.
a012 52 minutes ago [-]
But … but … but 9 out of 10 doctors recommended Opus
manojlds 6 hours ago [-]
Opus 4.8 was already shown to be cheaper than Sonnet 5 when Sonnet 5 was released (by Anthropic)
Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. Sure, 20% cheaper then Fable, but that is a 3x price increase compared to Opus 4.8 in that test.
Cost per task is second highest, right below Fable.
* Fable: $2.75
* Opus 5.0: $2.03
* Opus 4.8: $1.80
* GPT 5.6 Sol: $1.04
* Kimi K3: $0.95
Looks like interest levels of cherry picked cost in their report. Cheaper model, clearly NOT. More expensive in both benchmarks.
spider-mario 4 hours ago [-]
Your numbers are for “max”. Opus 5.0 “max” is $2.03. Opus 5.0 “high” (competitive with Claude 4.8 “max” on that index) is $1.06, less than the $1.80 you are quoting for 4.8 max.
That the most expensive variant is expensive doesn’t really tell us much.
benjiro29 4 hours ago [-]
Same answer i gave to somebody else up here...
If you start to drop effort levels, you need to compare to the competition models. So GPT models on the same ~intelligence level, are then 50% cheaper.
You see the issue? Its still a expensive model, and from my understanding, it still uses the old tokenizer.
Going to be interesting to see when GPT 6 comes out (very soon).
artursapek 5 hours ago [-]
It's definitely not cheaper than Sonnet on my benchmark, but it's cheaper than Fable and outperforms it. Which is big IMO. https://revise.io/errata-bench
abixb 8 hours ago [-]
So the rumors were right, Opus 5 was indeed being polished up for release. Huge improvements in GDPval-AA v2 too -- great for some of the knowledge work-based agentic workloads I run.
Also glad they still kepy Fable 5 on "credits only" access. I think we're going to start seeing model providers gate top-of-the-line models behind pay-as-you-go API rates/credits while subsidizing other models on monthly subscriptions.
jpk2f2 7 hours ago [-]
It's still available on at least some subs, they emailed me recently notifying me that I still have access.
saratogacx 6 hours ago [-]
My understanding is that you get $20 in api credits each month and a one time $100 until mid September. So you can still use the model with a subscription but you aren't getting any kind of discount.
I burned through $45 in 3 prompts to fix some bugs in my code (Some kind of tricky to isolate). That thing burns through cash so fast I don't see myself using it outside of maybe building execution plans for other systems
tackta 4 hours ago [-]
I am on the pro plan and got the $100 credit.
I have moved on from Fable anyway so just going to view this next 6 weeks as I have a massive amount of Opus 5 to use.
I had a hard time finding anything that would let Fable express its increased intelligence. The few conversations I had this afternoon with Opus 5 were pretty impressed.
If Opus stays one click back from the frontier model, I will remain a happy customer.
mcv 6 hours ago [-]
I think I saw that Max and Enterprise keep access, but Pro has to use credits, but I think I got $85 in credits.
ciefa 6 hours ago [-]
Fable 5 is included for 50% of the limits in Max. Only below Max one has to use credits.
Wowfunhappy 8 hours ago [-]
Fable 5 is still included in Max subscriptions!
7 hours ago [-]
collabs 7 hours ago [-]
Max is an individual subscription though and does not come with the guarantees that team or enterprise do?
ValentineC 7 hours ago [-]
Team Premium has Fable 5 too.
bakies 7 hours ago [-]
Doesn't team bill API rates?
einsteinx2 7 hours ago [-]
No that’s enterprise accounts. Team accounts are similar to regular Pro and 5x Max accounts in both price and features.
ValentineC 6 hours ago [-]
Team is 1.25x the price of personal accounts, but supposedly also gives 1.25x more usage.
eterm 7 hours ago [-]
And Teams Premium was previously needed for any claude-code at all.
d4rkp4ttern 7 hours ago [-]
what guarantees are these? You mean data retention, use for training etc?
> These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered.
So who is right? I know for Fable I am visibly told, is this tweet trying to say it is silent against what Anthropic is saying?
SwellJoe 4 hours ago [-]
So, announcing the fallback is better than doing it silently, but the fact that Fable falls back frequently for the kind of work I do (a lot of security oriented stuff lately, but it falls back on seemingly random stuff, sometimes, too), means I reach for it less. Getting interrupted mid-task makes it much less valuable. If I have any suspicion I'm going to hit the guardrails, I'll use something else.
pseudohadamard 18 minutes ago [-]
Same problem here, I seem to hit guardrails all the time when doing code audits. Does anyone have any insight over whether they're less annoying in Opus 5 than Fable? That is, is it better to start with Opus 5 than Fable because you'll get kicked backed to Opus 4.8 less often?
There also seems to be some cross-pollination across models, going Fable, Fable, Fable, guardrail, Opus 4.8, Opus 4.8, ... gives more Fable-like results from Opus than just Opus 4.8, Opus 4.8, Opus 4.8, ...
solenoid0937 4 hours ago [-]
Is some random guy on Twitter right, or official support docs that explicitly describe this scenario?
lossolo 3 hours ago [-]
It's showing you're switched to 4.8, i just hit that while doing security research.
trueno 2 hours ago [-]
i wonder if anyone thinks im weird for still using 4.6 lmao it's my "good enough" model. im more than pleased at what i can whip up with 4.6, once local llm's get here with a decent sized context window and it feels like using 4.6, i shall depart the land of these dumb service subscriptions
pseudohadamard 14 minutes ago [-]
Never even registered that that existed, the only thing I care about is whether I keep hitting the &#$&# guardrails that Fable has. They can keep the data forever as far as I'm concerned, just stop kicking me back to Opus.
gonzalohm 7 hours ago [-]
I don't understand how the data retention works. My company has an enterprise license with no data retention but if I ask Claude about past conversations, it remembers. So surely the information is being stored somewhere
NiloCK 7 hours ago [-]
Opus 4.7+ and Fable are both much more aggressive than prior models with respect to writing memories to a location that's effectively quasi-private for them. It's device-local (so passes retention constraint), and you can see it, but only if you go looking for it.
It's a funny design/affordance. I do see them often writing memories of things that that feel unlikely to be important going foward / with other tasks, but I don't see them clearly getting tripped up by them as prior models used to. (eg: Since you're running Ubuntu in Canada, here are some drills you can try to help your kid hit a baseball more consistently.)
persedes 7 hours ago [-]
You most likely are referring to the local jsonl files where claude has your sessions etc stored.
mh- 7 hours ago [-]
It could just be the memory features.
In my enterprise-seated account I see slightly different options available (vs. my personal account) in the Capabilities section:
Search and reference chats
Allow Claude to search for relevant details in past chats.
Generate memory from chat history (Legacy)
Allow Claude to remember relevant context from your chats. Memory includes your entire chat history with Claude.
The first option was defaulted to on, if I recall.
gonzalohm 5 hours ago [-]
But it kind of conflicts with the contract we have with them. My company has an enterprise contract that says "no data retention" but then each user can decide to enable it unilateral?
jrockway 4 hours ago [-]
If you're talking about Claude Code it's in ~/.claude/projects/<encoded dir name>/memory/MEMORY.md. So they're not really retaining it, it's just something that your harness loads in.
gonzalohm 3 hours ago [-]
Not Claude Code. Claude in the browser
bathtub365 7 hours ago [-]
Likely in memory files stored locally
gonzalohm 5 hours ago [-]
I'm talking about the website. It's not local because I can see my chats in any device
usef- 2 hours ago [-]
Yes, all chat interfaces store the history as it's part of the UI promise (unless you open an incognito chat) and is fully server side.
When people talk about retention they mean API usage and terminal agents, which run on your device.
manojlds 6 hours ago [-]
Claude Code? It stores a memory.md file.
fny 3 hours ago [-]
Anthropic offered ZDR for Fable on AWS bedrock from the beginning.
cjonas 2 hours ago [-]
Really? I was unable to use it in our account without having to enable the provider_data_share setting...
From the docs[0]:
> To use this model, you must opt in to provider data sharing by setting your data retention mode to provider_data_share via the Data Retention API
I hope we get clarification on this, I can't find anything claiming that it is compatible with ZDR.
collinrapp 7 hours ago [-]
Maybe I’m misunderstanding you, but if you scroll to the bottom of their [1] link to the Opus 5 announcement, under “Getting started,” it explicitly says:
> Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
solenoid0937 7 hours ago [-]
It's in the article.
doctorpangloss 7 hours ago [-]
do you mean, that organizations now have access to Fable-ish pelican drawing?
trueno 2 hours ago [-]
at last. time to lay off 22,000 employees
gigatexal 8 hours ago [-]
insane pricing:
"
Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8)"
RazorBucksICO 6 hours ago [-]
I think for the value of the outputs that’s still a good deal. Keeping the same price as the prior model makes sense to me. That is if the model size is about the same in the cost to serve has not substantially changed. Now I would have expected efficiency gains for inference, but there is no way to know as a customer.
At the end of the day, they have established a strong brand and if they can get away with a 95%+ gross margin on inference entirely from the status premium, then I suppose that’s good for them. Apple does the same thing, and I don’t fault them for it.
gigatexal 4 hours ago [-]
"insane" that they kept the price the same and didn't jack it up, my bad for the ambiguity.
oblio 6 hours ago [-]
Why is it insane if it's the same as the previous version?
khuzaimsharif 4 hours ago [-]
[flagged]
hnscum 8 hours ago [-]
[dead]
jjcm 7 hours ago [-]
Doing testing with it now, specifically for image->html conversion.
Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models).
Opus' results seem to be more accurate than Fable, following the design source of truth better.
Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).
Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
jjcm 6 hours ago [-]
Here's another test of a cyberpunk ramen shop website.
One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration.
Thoughts: It does a really, REALLY good job at these angular cuts / microglyphs. The responsiveness is off, but I'm very impressed at how well it did here. One way I think of it is "how close to a finished product did this get me?". Opus gets you like 90% there.
winwang 5 hours ago [-]
Several other commenters have disparaging the design seemingly mostly due to its AI-generated nature, or maybe they actually do dislike cyberpunk.
Personally, I think being able to have these design languages be easily prototypable is fucking awesome. Great tests! (But a tad low-performance/janky, somehow). Though, I also like the cyberpunk aesthetic. Very on-brand(?) that AI generates it, hah.
razster 1 hours ago [-]
I really like that design. May I ask the name of the website builder/diagram? Is it Relume?
echelon 6 hours ago [-]
God damn, we are living in the future.
I love this so much.
Designs like this would never have seen the light of day in the cellphone incrementalism / corporate memphis era of tech. Now people can be weird and awesome again.
This is 1980's cyberpunk / late-90's Matrix / early-00's sci-fi UI. Great ideas that died to frutiger aero (which isn't a bad design aesthetic) and flat design (which is).
This is fun and it's got great colors and I love it.
It's so refreshing to see this.
AI rules. This is the best timeline.
andersonpico 4 hours ago [-]
What do you mean designs like this? This is 2016-cyberpunk-neon-era inspired by video game interfaces, these are not uncommon at all. Google something like "discord cyberpunk theme" or "cyberpunk rice site:reddit.com/r/unixporn" and you'll find endless examples. This isn't suprising at all obviously.
brailsafe 3 hours ago [-]
I'd say the only thing it's a newish aesthetic for is web UI that isn't implemented in Flash, but even then there's been periodic resurgences. Command and Conquer, Starcraft 2, I'm sure there are tons of examples I'm unaware of.
The only recent novel addition—I'd speculate—is the specific influence of Cyberpunk the game with its shiny surfaces and pink highlights, but even then it's hardly new.
kami23 28 minutes ago [-]
You should have seen some of the Flash sites people made in the mid 90s early aughts. They all seemed to be straight up screenshots of the desired website and then buttons stapled on random portions.
They were an accessibility nightmare, but you use what you got. I tried so hard as a kid to understand flash, but had to settle on MS Frontpage to publish my first RPG page.
What's old is new again.
GPerson 5 hours ago [-]
I feel that AI has deeply diminished my ability to be weird and awesome, because my weird and awesome takes time and the results I can share with others are outshined by the machine.
jjcm 4 hours ago [-]
IMO what makes things awesome is human hours invested.
The ramen shop website above is pretty, but it's a veneer. It's not weird and awesome, it's just a representation of a site. I spent about... 7 minutes of my life making it. It's a tech demo, nothing more.
If someone actually poured their heart and soul into a vision for a cyberpunk themed ramen cart, and happened to use this because they didn't have the capabilities or funds to do a proper design, suddenly it becomes less of a veneer, and more just a component in the wider vision of that individual. Their human hours poured into the wider thing that's the business becomes what matters.
Ideally what AI does is it amplifies the hours we do pour into things that are weird and awesome, it doesn't replace them.
ultimafan 4 hours ago [-]
Strongly agree, I feel like my passion for development as a whole has waned away as AI has gotten more prevalent.
We do things to achieve some end result but it's the journey there that is the most cathartic to me. The "skilled crafts" element of development where careful deliberation and hours of tinkering to get any kind of appreciable output you can admire has been replaced with a one stop dopamine button that skips the whole process that I could find myself getting lost in.
I've taken up carpentry/metal working as a result. Maybe someday we'll have live in robots that do the same for those hobbies that AI did for programmers but I can't see it happening any time soon.
hangrybear666 3 hours ago [-]
Yeah I feel the same. That's why I've canceled my claude subscription and code everything by hand again, because building skill and mastery in a craft is fun and rewarding. Having a machine do it for me is neither. Fortunately coding is my hobby so there is zero pressure to optimize it.
At my workplace management is pushing AI, so I am using it in order to establish sensible and thoughtful applications of it and in order to know when to call out colleagues for pushing mindless automation out of complacency or blind obedience.
razster 1 hours ago [-]
I feel you. I dig this and enjoyed those types of aesthetics. Back in the day this was my bread and butter, designing UI design for games and apps. I just realized AI (Specifically Pi+Ornith) can help with my ideas... So excited.
I'm hella interested in finding out what website builder/diagram app was used. I dig the dark theme/grid.
This was just from a prompt "A cyberpunk themed ramen food cart website. Should feature menu, locations, and an ability to put in an order for pickup. Simple and clean website with angular cyberpunk microglyphs, pink/teal colors."
wilkystyle 4 hours ago [-]
This comment resonated with me so much, and it seems to be a minority view (at least on this website).
I have found myself empowered by AI to tackle all sorts of things that would have too high of a barrier to entry for me to want to spend my limited time on as a busy father who is also working at a small startup.
And when I say that, I do NOT mean that I can crank out a bunch of slop and label it as something I produced even though I don't understand the code. I mean that I can do things like go back to college math that I never appreciated at the time and honestly felt too scared of. I mean having an on-demand math tutor that ask clarifying questions to as I struggle through the problem sets.
I have found that it actually accelerates learning how to code in various problem domains because I can tell it to answer my questions at a conceptual level and be a sounding board, but to never actually write code for me. It can review the code I write and gently nudge me without giving away the answers, so that I still struggle through the learning process and actually gain the knowledge.
And finally, for the first time in like 10 years of feeling overwhelmed and daunted by the prospect of learning game development (I have no background in that), I have found Codex to be an incredible boon for learning with the Godot engine. It helps me understand the terminology so that I know what to search for and what documentation to read. It helps me map my computer science knowledge from other domains into the game world, and to understand why things are structured the way they are. And because Godot saves all of the scenes and geometry and lighting and shaders to the file system as text files, Codex can inspect the results of the work I'm doing in the IDE and help me track down things I'm stuck on, and explain what the issue is. For example, why my pre-baked global illumination lightmap is breaking my ambient lighting configuration.
I know it has never been easier to cheat and skip the hard work that results in actually learning something, but for me, personally, I cannot believe the incredible value that $20 a month has provided me. I have never been more excited and eager to dive into tackling hard things I had previously been afraid of or simply too overwhelmed to attempt.
It has never been easier to quickly prototype and get a feel for some idea you have in your head to see if it even has legs. Simply seeing a quick prototype of an idea is often all of the excitement and fuel I need to then take it and make it a real project.
hangrybear666 3 hours ago [-]
Your approach seems sensible. Especially the part about discussing concepts to gain understanding instead of having output generated for you to avoid having to think.
Unfortunately, most people are interested in shortcuts and thinking less and will ultimately deteriorate in their abilities the more they abuse these tools to skip the learning process.
aaa_aaa 6 hours ago [-]
No we are not living in future. Design is ugly, and immediate put off because it smells AI.
djeastm 3 hours ago [-]
We could say the same for the first factory-made clothes. Not as nice as the hand-tailored ones, looking a bit "off" when people wear them compared to custom-made, but orders of magnitude cheaper and more convenient.
FranzFerdiNaN 5 hours ago [-]
Reading your comment I had to think of that tweet about someone taking a part of a Monet painting and claimed it was AI generated, and people immediately started calling it horrible and soulless and smelling like AI.
bonoboTP 4 hours ago [-]
Obviously. Artistic value is not in the object, but in the relationships to other humans. The point is not Monet but the people you discuss Monet with. Monet is a tribal rallying point you can synchronize around. AI is not as good for this, because it's cheap. It's not scarce enough. And doesn't offer a human story you can bond around.
aaa_aaa 5 hours ago [-]
Well this is no Monet for sure.
newsy-combi 5 hours ago [-]
Humans are capable of producing slop too. Maybe the cutout does look like AI. Given Monet's blurry style and repetitive content (he made 250 water lily paintings), this is not surprising. When you strip the cutout of its context, it can look more like AI too because the lighting and composition will look arbitraty compared to the whole picture.
Btw, if websites would only include the frontend dev's own hand painted images, we would also revolt at the sight of human slop. It's not just AI.
The whole point is that good artists are capable of producing non-slop, and to this day they're the only group of which this is reliably true.
tackta 3 hours ago [-]
It is because it is an image on a screen.
Water Lilies are enormous paintings. They are breathtaking in person because of their scale. Monet wouldn't be Monet if he had only produced images on a screen.
Art is good or it is shit, based on personal taste. Just like food, no one can tell me what food tastes good or tastes bad.
AI Art seems to produce strong emotions in people who don't go to art galleries. I love modern art, I am a huge art snob but if you want to see slop, go to any modern art gallery. Personally, I would say for my taste, at least 70% of all art at any gallery is basically shit.
Like food also, the presentation matters. To believe there is no possible way to print out a 6 foot tall by 10 foot long AI generated image that would look awesome hanging in a gallery is stupid.
shimman 1 hours ago [-]
Okay but in this case the design aesthetic is already extremely similar to cyberpunk 2077 video game UIs. Maybe that was the OP's test, but to claim it's unique when it is not feels misleading. Derivative works are fine, but you need to ADD something as well. Taking this then adding another 12-20 hours of some animejs + css polish would make it stand out but as it is now it's just a copy.
HaZeust 5 hours ago [-]
Then tell it that, and it'll give you the output you desire.
theappsecguy 6 hours ago [-]
We are going through yet another generational wealth transfer and people are being squeezed to the absolute brim with layoffs and daunting lack of career prospects.
But sure, lets cheer that funky website designs are back on the menu…
Mtinie 6 hours ago [-]
Assuming no change from the baseline, I’d rather have one positive thing versus nothing positive.
shimman 1 hours ago [-]
"We're taking away your healthcare, sending your children to die in forever wars, poisoning the environment, all while ruining your job prospects but at least you can make pixel perfect implementations of figma screens."
nextaccountic 1 hours ago [-]
The culprit here is capitalism. The system was designed for this. Any technology will be used to squeeze the working class
AI has the merit of showing SWE folks exactly where in the class divide they belong. If you are selling your workforce, and you can't maintain your lifestyle if you stop working, you are in the working class
signatoremo 5 hours ago [-]
Why don’t you start a movement to fight for equality, and I’ll join you? Why are you still here and not taking actions?
In the mean time, I’ll unashamedly continue to cheer for creativity and innovation. Note: I don’t even like this website design.
kypro 6 hours ago [-]
I'll add to this...
The "funky" websites of the past were mostly a result of tech immaturity and a lack of profit motive.
Businesses have been able to easily install templates like this for at least a decade. They don't because stuff like this looks cool but isn't very functional.
AI isn't going to make your local restaurant have a funky website, it's just going to make everyone who use to work directly and indirectly for that company unemployable. And even the local restaurant will close down because they can't compete with the multi-national competitor that has automated their kitchen with AI.
andersonpico 4 hours ago [-]
I agree with most of your comment with one exception:
It will never be economically viable to replace kitchen staff with AI.
kypro 3 hours ago [-]
> It will never be economically viable to replace kitchen staff with AI
Can you expand? Specifically what is it about humans that AI and robotics could not replace?
mainmailman 4 hours ago [-]
Unrelated question, what’s your favorite flavor of Kool Aid?
ai_fry_ur_brain 6 hours ago [-]
Because they're incredibly ugly
echelon 6 hours ago [-]
It's gorgeous and the world doesn't have enough of it.
wilkystyle 4 hours ago [-]
I was gonna say, you can dislike it because of the design, but understand that that's a subjective opinion that has nothing to do with whether AI is awesome. I actually like these designs, personally.
thatxliner 22 minutes ago [-]
> Doing testing with it now, specifically for image->html conversion.
I wonder if there exists a benchmark for that.
chriscamargo 7 hours ago [-]
This is an awesome test! Thanks for sharing the results. Opus 5 is very impressive.
Out of curiosity, what app is that Design source of truth screenshot from?
Edit: Generation was down, back up now. Apparently just hit my $1000 cap for the openai api. Upped it to 10k. Growth!
IanCal 3 hours ago [-]
Huge fan of the non-seat/monthly based pricing. Do you need it to be higher to not just cover costs but also make a profit?
jjcm 2 hours ago [-]
Right now generation is entirely at-cost (0% margin). Once I finish my SOC2 I'm going to release a enterprise/team license, which I'll charge a fixed 50% margin for. Long term the goal is for those enterprise licenses to be the profit center and for individual accounts to drive growth.
At the moment I currently have around $600 of revenue on $1200 spend, but that's primarily because I'm subsidizing new accounts (each new account gets $5 to spend for free, which translates to around ~36 designs). I'm in the process of doing an angel round, so I can afford to operate at a bit of a loss during the growth stage.
abidlabs 3 hours ago [-]
I was curious to see how open weight models would do on this task so I passed in a screenshot of your source of truth and here's what 2 of the best code-generation models that allow image inputs do:
Not bad at all, and this is pretty consistent of what I've found from the current open source models. I haven't tried it with kimi 3 yet, that's on my todo.
bottlepalm 7 hours ago [-]
I just clicked your links and then read your comment after - my first impression was the Fable version looks way nicer.
jjcm 7 hours ago [-]
I agree the fable version looks nice - the rounded hero image for instance.
Opus though followed the source of truth better imo. The details are more present.
Fable filled in the gaps for things it wasn't able to do (ie in the design the hero image goes behind the nav), which resulted in a better looking page that was more divergent.
andersonpico 4 hours ago [-]
I liked the Opus version better if only because the responsiveness is less broken.
kccqzy 6 hours ago [-]
Same. I like the Fable version better. Better colors, better choice of font sizes, better column sizing. Also small things like the “Experience” section header being orange rather than gray, which Fable got right and Opus got wrong.
It seemed to me that Fable meaningfully improved on the original design more than just faithfully executing the original design.
afro88 6 hours ago [-]
IMO a much better test would be designs that aren't AI to begin with. Much more useful to see how well a model can html an image design without slopping it up
erikw 7 hours ago [-]
Very interesting that Fable took more creative liberties. Have you tried giving Fable the same task, but also specifying that it implement a pixel-perfect design? I think that I prefer the Fable implementation. I find the UI elements in the Fable implementation to have more contrast, which feels more usable to me. I also like how the right padding on the "Book Your Escape" CTA in the upper right matches the top and bottom padding, which I think is an improvement over the mockup.
jjcm 6 hours ago [-]
All of these are using a build skill which specifies rules for building it, requirements to create a pixel perfect implementation, and tooling to help in that process. Here's the build skill / instructions I pasted in to both of them:
> Create a web page implementation from the following instructions:
> You are an elite frontend engineer and design-to-code specialist. The design image is the primary source of truth; your code is the translation layer. Do not reinterpret or "improve" the design into something generic — reproduce it faithfully.
Thank you for sharing this. I was just using OpenAI's Product Design plugin[1] to create designs but it just didn't reproduce it in code faithfully so will need to try this.
http://impeccable.style/ also just released their latest version which has some image->html conversion. They just released a couple days ago and haven't had a chance to try, but I suspect theirs is a bit more robust than mine for pure LLM instructions. Worth trying and comparing it with mine.
abdussamit 2 hours ago [-]
Please share your prompt to convert image to html!
Is this with browser tooling attached to the agent for review/iteration?
jjcm 6 hours ago [-]
yea this was just straight into claude desktop / its standard tool usage, on "high" thinking. The ramen website is on "extra" thinking.
ai_fry_ur_brain 6 hours ago [-]
[dead]
paxys 8 hours ago [-]
Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now.
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
ai-x 8 hours ago [-]
Model Routing will always be done better by models themselves.
Plus routing loses context making it more expensive and less reliable.
Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability
johnfn 7 hours ago [-]
This doesn’t seem obviously true, eg an Anthropic model will never route to Kimi even if it were best suited for a particular task.
posix_compliant 7 hours ago [-]
I think what the parent is saying is that the model itself has the best context for whether a portion of a request should be routed. The specifics of that routing (e.g., should you route to KimiK2) are something that can be trained, finetuned, or even included in a model's startup context.
johnfn 3 hours ago [-]
This doesn't seem quite right. For one, I don't need all the intelligence of an expensive model like Opus 5 to do the relatively simple task of choosing a correct model for a task. Additionally, since this isn't something Anthropic would ever put effort into doing well, you could tune a model to do better and faster than Opus 5 does out of the box.
kolinko 2 hours ago [-]
Are you speaking from experience?
My experience is the opposite - for many cases it’s not very obvious how good a model needs to be to solve it. Worse models tend to just follow their first instincts without proper reasoning
And also btw you don’t need a
routing company to decide, you can do it on your harness. And yeah my Fable has zero issues delegating to Terra instead of Opus.
maCDzP 6 hours ago [-]
Has anyone tried that? I have a feeling that if I put it a prompt Claude would comply. But I am all in on the Claude cool aid.
johnfn 6 hours ago [-]
Sure Claude would comply, but Anthropic has no financial (or other) incentive to optimize this, so there’s no reason to expect it to be particularly good.
It would be like asking the clerk at a Whole Foods which grocery store in the city sells the cheapest eggs. He’d probably answer - he might not even say Whole Foods - but WF is hardly teaching all their staff the best methods to answer this question in training. (Heh, training.)
siva7 7 hours ago [-]
Why should it? An Anthropic model is architecturally optimized for Anthropic models, routing it to Kimi makes zero sense
anon7000 6 hours ago [-]
Which is why 3rd party routers which do route between different models may have an edge. It means they can compete on cost, and it’s definitely not clear that the architectural optimization is always going to be higher quality or cheaper. It might be, but everything changes constantly, so locking into a single model family/company is very much not ideal
6 hours ago [-]
MagicMoonlight 5 hours ago [-]
[dead]
fny 46 minutes ago [-]
This is actually simpler to implement. You can have a AGENTS.md describing who is boot at what and then you have the agents converse over tmux.
vanuatu 7 hours ago [-]
i think you're thinking of subagent routing
model routing in this case is cross-provider
Imo the main issue behind model routing is you need to figure out how much intelligence a new task takes, which is a very non trivial problem. Presumably, a organization knows this about their own tasks and is better suited to built in-house compared to outsourcing to a vendor.
CuriouslyC 7 hours ago [-]
Model routing by the model itself requires the model to pull in a lot of context and it's likely more efficiently just done by people with the context already in their head, even assuming the model is perfect at routing (which last I checked, Claude definitely isn't). I wouldn't trust ML model routers.
hnfong 8 hours ago [-]
Because they're trying very hard not to understand it.
Otherwise the expensive-yet-powerful model probably won't see much revenue. How much money is there in bleeding edge scientific research? There's a lot, but there's even more existing capital in paying people people to do college level paperwork, and the bulk of those traffic gets routed to the cheapest model.
You mostly don't need super powerful AGI to replace the paper pushers, but the frontier labs are trying to position themselves as being uniquely capable of producing super powerful AGI, and also be the ones replacing office workers.
Not sure how it will work out for them, but I think model routing is going to poke holes in that narrative. That's why I think they're trying very hard not to understand model routing exists.
ip26 4 hours ago [-]
It wasn’t long ago at all that the chief problem was “can AI even help me with this” (cost be damned). Until a time when the answer is an unmitigated “yes, obviously”, the frontier labs have everything to lose and nothing to gain on routing, because if they screw it up you might incorrectly decide “nope, it can’t help yet” due to a poor routing decision.
TeMPOraL 8 hours ago [-]
Who are the customers though? Honest question, I'd like to understand it.
For me, anything other than current best available SOTA for any task is unacceptable. The only routing rule I need is "the most powerful model I still have flat-priced quota available for". I mean, why settle for less?
btown 7 hours ago [-]
There are two types of users: those who are able to use subsidized rates, and those who need to use API rates due to audit requirements, enterprise billing, etc.
Model routing for subsidized users takes the form of a "use Opus 5 subagents for implementation" type of system prompt. You lean into a single provider, build tooling around that, and your savings are far beyond anything multi-provider routing can get you.
on top of that there is also additional factor: speed - sometimes if task is easy you do care to finish it faster.
There is also matter about convenience - when I ask some small easy question often I don't bother to switch the model or forget in prompt to ask faster/cheaper subagent.
kxxx 7 hours ago [-]
It's very common to use a lesser model for a lesser task, resulting in same quality output. End result: save money while being faster. In many cases, it's a pure win-win.
lostmsu 7 hours ago [-]
But in other many cases you have to redo the work directly or indirectly, and you are more expensive (for now) than even the most expensive models, so sounds like a total lose.
paxys 7 hours ago [-]
I’m assuming you don’t pay per API call. Every mid-large sized business in the world does.
Slartie 7 hours ago [-]
And you do not have flat priced quota for Fable 5, right? Because nobody has, as far as I know. So you'll probably not route any task to the "current best available SOTA".
Also: quota. Implies you do not have unlimited access even for flat prices. Which in turn implies that as soon as you hit the quota on the most expensive flat price plan, even you will suddenly discover the magic of economically sensible behavior.
abi 7 hours ago [-]
Fable 5 is included in the Claude Max subscription. I've gotten close to the limits this week and last but haven't hit it yet.
dgellow 7 hours ago [-]
But Fable is only available for 50% of your Max quota as far as I understand
msabalau 7 hours ago [-]
Not everyone is you. Other people probably have a range of tasks that can accomplished with different models.
Certainly if I'm confident that I'm going to get what I need from a faster model, that's what I want to use, rather than wasting time grinding away for the sake of saying of the same answer came from a SOTA model.
Given that every chatbot does offer a range of models, it seems clear people do choose among options.
internet2000 7 hours ago [-]
The mental effort in estimating what model would be better is so not worth it.
I just want to switch to Claude Code, tell it to turn a .csv into a BigQuery table then cmd+tab to something else while it runs. Thinking "oh this is probably an easy task, I can /model to Sonnet to save $0.0004" is silly.
polotics 7 hours ago [-]
because "less" can be so much faster?
kxxx 7 hours ago [-]
not just that -- "less" ($$$) can also result in indistinguishable quality for some tasks/inputs. I'd argue this is the primary reason, secondary being speed.
jnwatson 7 hours ago [-]
I generally agree. Perhaps there's only a 5% chance that it would write better code or find a bug that it wouldn't have with a lesser model, but the economics of bugs is strong enough that preventing a single bug is worth hundreds of dollars.
StilesCrisis 5 hours ago [-]
That's nice and all, as long as flat-priced quota continues to exist. Seems unlikely to go on forever!
TacticalCoder 7 hours ago [-]
> For me, anything other than current best available SOTA for any task is unacceptable.
Then you must route. An article with lots of upvotes yesterday or two days ago showed that K3+Fable 5 was more SOTA than either of those.
toss1 7 hours ago [-]
I'm not a customer of those routing systems, but I quite often use different Claude models for different tasks. While most tasks were Opus 4.8, I often used 4.8 to make a plan, prompts, and package kit to setup Fable for a bigger project, then run it on Fable. Or, for broad single-task searches Sonnet with or without "Research []" turned on seemed to work best both faster, lower overhead, and less verbose answers (when I didn't want it).
OFC, YMMV
awongh 7 hours ago [-]
what's the threshold for model routing where you're willing to trust the router?
For coding my own work I don't trust the model router, and it would have to be shown to be to save a real dollar amount.
From a buying perspective it's a hard sell to save x but lose out on bugs you are probably introducing at an unquantifiable severity and frequency. How much is it worth to hedge your bets by doing every single inference request on the frontier model?
How much will it cost to go back later and fix things, but also the meta question of how to be able to decide on a hypothetical unknowable? (You'll never know how much better or worse your code was gonna be, it's untestable at a project level)
binary132 7 hours ago [-]
you trust the service provider but not the router?
weird, but ok
awongh 7 hours ago [-]
I don't want to save money so badly that I'd possibly undercut the quality of the code that gets created.
*edit to add: that code quality (or lack of quality) is it's own cost
owenthejumper 5 hours ago [-]
I am not convinced this will end up being a domain of the ‘routers’ vs the clients, as in harnesses themselves. Thoughts?
7 hours ago [-]
ModernMech 6 hours ago [-]
lol I had to get ChatGpt to explain to me the difference between 5.6 sol, 5.6 Terra, 5.6 Luna, 5.5, 5.4 mini, 5.3 spark, and then there is low, medium, high, extra high, max, ultra, and pro… I still don’t really know, it feels like ordering hot wings.
simianwords 8 hours ago [-]
Openrouter should ideally kill in this space and make their model agnostic infra like memory, harnesses, chat applications.
verdverm 8 hours ago [-]
OpenRouter is in acquisition talks with Stripe, fyi
I would expect routers to commodify like tokens.
torginus 7 hours ago [-]
OpenRouter sprang up overnight. I might need to replace some urls and access tokens should they decide to try and screw me.
verdverm 7 hours ago [-]
I never signed up because I found the 5.5% fee on token usage to be a "screw you" tactic. Still do not understand why they are popular with the other options out there.
TacticalCoder 7 hours ago [-]
And it's not just that model routing is much cheaper: no longer than yesterday we got a post showing that routing between K3 and Fable 5 was more SOTA than either of those.
If that is true, model routing is here to stay.
It also seems to validate the minimalist approach of pi.dev, where sub-agents from the same company is not the preferred approach (pi.dev believes in neither sub-agents all from the same company nor MCP even you can do it if you want for pi.dev's philosophy is to do add any functionality you want to a minimal harness).
Now of course we'll get for a few weeks all the Anthropic fanbois and shills explaining that "sure, K3 was basically at the level of Fable 5 but now that Opus 5 is out, open-weights models are six months behind".
deet 6 hours ago [-]
I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from.
Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"
And this is the most important observation in this thread. It’s load-bearing!
Kwpolska 5 hours ago [-]
Fable might be using those phrases less, but its writing is still terrible and exhausting to read.
wfme 4 hours ago [-]
Agreed. I’ve interestingly found 5.6 sol to produce much better writing, and it can generally cut to the point much more effectively.
kranke155 3 hours ago [-]
Opus 4.6 remains unbeatable in my book. Fun to talk to. Fable felt very human. But not as fun.
arizen 5 hours ago [-]
Could these complex/hard to read Fable outputs be sign of some kind of industrial level of intelligence, which us humans may have a hard to comprehend, while it may be also hard for machine to use simpler texts to properly outline all nuances and complexities of concepts it output?
solenoid0937 4 hours ago [-]
Fable subagents communicate very effectively with one another, so this would be a reasonable take imo
dwaltrip 3 hours ago [-]
I've increasingly felt like Fable and I speak different dialects of English...
merlindru 2 hours ago [-]
models are hungry for a more information-dense language
for now all they've got is english, so they'll just bend that into shape. it'll do.
trinari 2 hours ago [-]
I've heard chinese contains more information per token
markab21 3 hours ago [-]
Yes - THIS! I can't even believe how exhausting it is to read. I'm not sure why or what changed in Fable. Did they do this writing-style output to give it more token compression during/for training or to prefer output for less money?
I love it for a few things, but it's gotten really hard to spend any extended amount of time with it because of the lack of mental model I seem to be able to hold while working with complicated problems.
I'm guessing it's just not enough time doing RL on human feedback.
"Early in RL verbose, grammatical" (if you search) :
We need to understand the operator.
The 5D line element is ds² = e^{2A(x)} (ds²_4d + dx²), where A(x) = sin(x) + 4 cos(x), x in [0, 2π].
The internal coordinate is periodic.
The background is a warped product: metric g_{MN} where M,N = 0..4.
The internal direction has metric e^{2A(x)} dx²?
Wait, the ds² is e^{2A} (ds²_4d + dx²).
So the internal metric is e^{2A(x)} dx².
Actually if the total metric is ds² = e^{2A(x)} (ds²_4d + dx²), then yes, internal metric is e^{2A} dx².
vs. Post RL
We need determine eigenvalue problem for spin-2 fluctuations h_{μν}(x,y) with TT in 4d and depend on x.
For metric of form ds² = e^{2A(x)} (g_{μν}(y) + h_{μν}(y,x)) dy^μ dy^ν + e^{2A(x)}?
Wait internal metric is e^{2A} dx²?
Actually ds² = e^{2A} [ds_4² + dx²].
So internal metric is e^{2A} dx²; warp factor same for 4d and internal?
Yes.
We need equation for h_{μν}(y,x) = h_{μν}(y) ψ(x) maybe with normalization.
…
I can understand it with less cognitive load in the post-RL version versus early in RL. This resonated with my experience using Fable, especially digging hard problems; it feels like I'm reading the "early in RL" version of that model explanation.
kanodiaayush 4 hours ago [-]
I've found Opus 4.8 and Fable 5 both difficult to learn from purely because of how annoying their writing style is. I'm finding GPT 5.6 Sol to be much better for this.
arjie 2 hours ago [-]
One nice thing but ChatGPT is that good image model means it can generate good infographics occasionally to illustrate. These become naturally compact in text.
DonsDiscountGas 4 hours ago [-]
I think a signature Claude style of writing is good since it makes it that much harder to pass off Claude written text as human.
robwwilliams 1 hours ago [-]
Easy enough to change. I have a Stylometry Skill fit to my preferred style—a mix of me and Terry Winograd. Give Opus 10 of your best paper thst you wrote and tell it to build a model of your style. hHard to distinguish except I make way more typos.
winwang 5 hours ago [-]
I found 4.6 more amenable than 4.8 to style directions, we'll see how 5.0 does. Super-small-sample-size: I think part of its "Claude-ism" style comes from its propensity to try and "proactively" move the conversation/work along. Not sure how this would fare in non-obviously-productive environments, I'd guess "it's still annoying" considering your evidence.
I'm also thinking of another benchmark: (quantified) stylistic range across different prompts. Just putting it out there if anyone wants to do the work for me :D
sibeliuss 7 minutes ago [-]
4.6 is night and day better. It was before the big language switch up. Terrible direction that Anthropic has taken this.
ianberdin 6 hours ago [-]
I’m pretty sure Opus 5 is adapted to tricks from long reasoning in Kimi K3 and based on original Opus 4.8. It is not fable in any form.
duplessitous 5 hours ago [-]
Seems unlikely they adapted anything from K3 given the timeline of releases, similar to how K3 was obviously not distilled from fable
I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!
dd8601fn 8 hours ago [-]
At half the price and less likely to auto-downgrade, it sounds like a reasonable claim.
benjiro29 5 hours ago [-]
> At half the price and less likely to auto-downgrade, it sounds like a reasonable claim
Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8).
But better cost for the same performance. According to AA, Opus 5 _medium_ is as smart as Opus 4.8 _max_, at 1/3 the cost and twice the speed. And if you need a better response, you can turn it up to 11.
benjiro29 4 hours ago [-]
Then your comparing to a level of GPT 5.6 High, what is 50% cheaper then Opus Medium for the same intelligence / score.
You see the issue, if you try to scale effort down, you also need to compare how other competing models compare.
binsquare 8 hours ago [-]
given that i couldn't even use fable without it downgrading to Opus, this is just a straight upgrade for me
Freedumbs 6 hours ago [-]
Opus 5 also downgrades. it's now Fable -> Opus 5 ; Opus 5 -> Opus 4.8.
Unclear why they want to nerf their own products with sometimes right classifiers. I guess the government ban might've been real and not coordinated marketing?
ceejayoz 8 hours ago [-]
Best can describe multiple things.
Almost as good for half the cost is something I'm very comfortable describing that way.
lelanthran 7 hours ago [-]
> Almost as good for half the cost is something I'm very comfortable describing that way.
It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar).
moffkalast 5 hours ago [-]
Hopefully it's not like old Opus, where it was actually more expensive than Fable cause it thought for half an hour, got it wrong, and then thought until you ran out of credits trying to come up with a correction, while Fable just went for it and did it in one go, getting it right the first time without thinking more than a few seconds.
Got an endless list of stuff done with Fable, Opus 4.8 was like a flailing braindead idiot in comparison. Maybe this one is a bit better if it's distilled.
ProofHouse 8 hours ago [-]
Best marketing
tshaddox 8 hours ago [-]
The blog posts figure cites Frontier-Bench for its agentic coding score, and shows Opus 5 beating Fable 5 43.3% to 33.7%.
ActivePattern 8 hours ago [-]
I think you're being overly cynical here. First, I don't see any claim that is the world's best model for agentic coding. Second, it is absolutely the best model in terms of coding performance vs. dollar, and it's raw performance seems very close to the frontier.
adam_arthur 8 hours ago [-]
GPT 5.6 is far more token efficient at most tasks with similar performance. Especially so for Opus 4.8, still to be seen with Opus 5.
Where are you getting cheaper per dollar?
ActivePattern 8 hours ago [-]
How are you supporting the claim that GPT 5.6 is "far more token efficient" than Opus 5? Tokens equal, output is cheaper for Opus 5 ($25/1M) than GPT-5.6-Sol ($30/1M), and it seems to outperform slightly on agentic coding benchmarks.
adam_arthur 8 hours ago [-]
The first chart in the blog post shows a similar $/performance curve to GPT 5.6.
Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels.
There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving)
If they had a more efficient model at coding they would lead with that chart.
It seems roughly equal according to Anthropic's benchmarks
shwaj 8 hours ago [-]
It would still be the best model per dollar if the score was 2% lower instead of 0.1% lower. Would it be ok to still give it the highlight color then?
How big of a lie is too big? Especially when no lie needed to be told at all: many including myself would have noticed the tiny 0.1% deficit and been suitably impressed by the Opus 5 result.
I’ll admit this is a small deception by today’s standards. I’m one of those who believes in truth for truth’s sake.
Edit: typo
ai-x 8 hours ago [-]
we don't know if it is 0.1% deficit, could be 0.05%
shwaj 7 hours ago [-]
So highlight both then.
manojlds 8 hours ago [-]
Which numbers are you seeing? It does show that it's better than Fable 5 in most things related to coding?
jsLavaGoat 8 hours ago [-]
In my opinion, the frontier is passed what is really needed for coding. Fable is good as a supervisor.
8 hours ago [-]
toephu2 7 hours ago [-]
Also it scored worse on DeepSWE than chatgpt 5.6 sol
dbbk 7 hours ago [-]
Yeah I spotted this immediately too. I'm sorry. You're supposed to be a multi billion dollar company and you can't even highlight your chart honestly?
Aurornis 8 hours ago [-]
Using the most expensive model for all of your agentic coding work hasn’t been good practice for a long time. Not unless you have infinite money to spend.
Fable is typically used for key planning, architecting, and review tasks.
I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.
airstrike 8 hours ago [-]
They cost the same if you're already at $200/mo
Aurornis 7 hours ago [-]
Fable consumes your usage at a higher rate.
If you bought the $200/mo plan and you don’t use it much, using Fable for everything is fine.
maineldc 7 hours ago [-]
I am not a tokenmaxxer per se but I blow through my weekly quota on my max plan in 3-4 days… fable would make that worse.
akmarinov 8 hours ago [-]
Eh, not really. Fable does a lot better on coding than Opus 4.8.
Just this past week Fable was able to figure out a couple of small issues for me where Opus was failing to.
Also both are still somewhat bad at UI implementation. Opus more so
unclebucknasty 7 hours ago [-]
Recent releases have said something to the effect (paraphrasing here):
"Use <less expensive or older model> for everyday tasks and <other non-critical stuff>. Use <more expensive or recent model> for complex coding tasks, refactoring large code bases, etc.".
Then, the next model/release emerges and the previous "best for complex" gets demoted to "everyday".
Obviously, it's all relative. But, it does beg the question: was the previous model really good for complex coding tasks or no? I mean, how is it now suddenly only good for the "easy" stuff?
hvb2 7 hours ago [-]
> I mean, how is it now suddenly only good for the "easy" stuff?
Because your expectations have changed.
unclebucknasty 5 hours ago [-]
I'm sure the marketeers would love for the public's assessment of complex versus easy to conveniently shift per their release cycles; or for the public to simply forget their prior marketing.
entropicdrifter 8 hours ago [-]
I mean that certainly makes it best-in-class
siwakotisaurav 8 hours ago [-]
Thanks for that, looks really good. I can see why they were constantly pushing back fable going out of the max sub with these benchmarks
kossae 8 hours ago [-]
I wonder why FrontierCodev1.1's data lists Opus 5 as better than Fable 5.
8 hours ago [-]
lattalayta 7 hours ago [-]
this feels like the perfect example of an LLM producing a long text document. And end users just using an LLM to summarize it without actually reading it
nerdsniper 8 hours ago [-]
Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete.
---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
You're comparing the "score percentage" (e.g. out of the total number of partial points available, how many did the agent achieve) to the "completion percentage" (how many tasks does the model score 100% on). The paper says "Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score", which is ~the same number that Anthropic reports here (55.7 vs 54.8).
That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable)
tadfisher 8 hours ago [-]
In what world is 55.7 the same number as 54.8?
What variance is acceptable to publish without a retraction?
nerdsniper 8 hours ago [-]
That seems like entirely reasonable variance to me for AI models. For my purposes, that absolutely counts as a solid "replication". I'd probably accept +/- 5 percentage points even.
tadfisher 8 hours ago [-]
D'oh, they are running the benchmark themselves. Reasonable.
Atotalnoob 7 hours ago [-]
There is randomness in LLMs. Both papers authors probably ran the bench 1-N times. Depending on that, they might select an average, max, least, etc. They might also have discarded outliers.
Like the other person said 5% variation is probably expected
jll29 6 hours ago [-]
I don't know who downvoted the parent or why, but it's a fair question IMHO.
The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic.
A proper methodology would ask each question 20 times and calculate the mean correctness across experiments.
The reason is that the temperature parameter introduces random behavior.
8 hours ago [-]
aleenz1102 8 hours ago [-]
[flagged]
Ancalagon 8 hours ago [-]
Its slop all the way down.
ssalka 8 hours ago [-]
I think some variance is to be expected since LLMs are typically non-deterministic, however that's a huge difference that I think warrants further explanation.
8 hours ago [-]
8 hours ago [-]
HyperL0gi 8 hours ago [-]
Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.
impulser_ 8 hours ago [-]
Go read the safeguards section in the report and you will realize why that is.
These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards.
OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.
HyperL0gi 8 hours ago [-]
Yes, this makes a lot of sense, but it’s just very amusing to see. 2 months ago, the world was about to end, now not so much.
K0nserv 5 hours ago [-]
7+ years ago GPT2 couldn’t be released because it was deemed too dangerous[0]. It was, of course, eventually released.
> We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):
* Generate misleading news articles
* Impersonate others online
* Automate the production of abusive or faked content to post on social media
* Automate the production of spam/phishing content
Seems like the prediction was pretty accurate.
jgilias 4 hours ago [-]
It’s seven years already. Crazy.
reciprocity 3 hours ago [-]
Yes, I think so too.
dzonga 6 hours ago [-]
I realized it a why back these labs are selling hype.
since then I have never cared about models except those that affect money in my pocket e.g AWS Nova Sonic
baq 7 hours ago [-]
Have you been patching your systems for the past two months? It was crazy even if you completely forget the supply chain literal FUBARs and you must’ve been living under a rock to not see OpenAI (accidentally) pwning hugging face
neuronexmachina 8 hours ago [-]
Do you have an example of the "doomsday marketing" you're referring to?
I see a pretty big gap between finding software vulnerabilities and “the world is about to end”. It is literally true that AI models are finding software vulnerabilities. It is also to my mind a reasonable thing that you’d want to be cautious about rolling out a model that can find more vulnerabilities. So what is the objection you have to these sources?
NichoPaolucci 7 hours ago [-]
Absolute masterpiece of a rebuttal. No notes.
reducesuffering 5 hours ago [-]
> now not so much
OpenAI Huggingface breach begs to differ
knuppar 7 hours ago [-]
feels almost like anthropic is desperate for ipo huh
i think we'll see one of the fastest deflations in history post anthropic/oai ipo
efficax 6 hours ago [-]
I feel like i've seen less hype about "the next model will be agi". GPT-6 is supposed to be coming this summer, and nobody is expecting AGI now. Not sure how they're going to keep the hype cycle going
vl 3 hours ago [-]
Or another way to see it is that current models are AGI as it was defined before, and the goal post is being moved.
efficax 3 hours ago [-]
they are definitely not agi as it was ever defined. they’re only a bit more capable than they were a year ago. they crossed over from interesting crap to useful tool recently but really only for software
scrollaway 3 hours ago [-]
You have short memory if you think we've not blown past at least 5 different AGI goalposts. They're being moved every time and we're hitting them every time.
Or maybe you just don't know exactly how capable these models are. Most people's experience of AI is a stupid chatbot, it's no wonder they don't understand how these things are coming for their jobs.
On my end, I have a software that is designed and built by Claude, that I did a strategy session on (with claude), and prepared a fundraise for (with claude). My only role, other than "knowing what to aim for", has been to feed the AI some fairly basic english prompts for a few weeks... which is also easily automatable.
Everyone's job is fucked. Devs, CEOs, everyone.
le-mark 1 minutes ago [-]
> Everyone's job is fucked. Devs, CEOs, everyone.
It’s curious to me that there are two distinct factions here. People like parent commenter who has no discernment and others who see llms for what they are. I just talked to opus 5 and in it’s first response caught some well disguise BS. These things are bullshit machines. There are indeed a lot of bullshit jobs around so maybe parent does discern something I don’t?
fatata123 7 minutes ago [-]
[dead]
iLoveOncall 3 hours ago [-]
LLM labs dumbed down the definition of AGI as much as possible, yet their models haven't reached it still. We are nowhere near the original definition of AGI. Not even 1% of the way there.
skohan 5 hours ago [-]
Yes 18 months ago it seemed like AGI was being promised every other week, and now I don't see any of those headlines.
8 hours ago [-]
websap 8 hours ago [-]
Fable established the frontier, this is just catching up.
bottlepalm 7 hours ago [-]
So unless doomsday actually happens then you're unhappy with the warning - is that right? You see false promises of apocalypse as marketing?
HyperL0gi 6 hours ago [-]
My point is why the sudden change in tone? I’m not dismissing the models’ capabilities.
MostlyStable 6 hours ago [-]
As they explicitly say, Opus 5 is ~ equally capable as Mythos/Fable at finding vulnerabilities, but it is much less capable at exploiting those vulnerabilities on it's own. That is an extremely meaningful difference and to me completely explains the difference in tone, release style etc.
emp17344 6 hours ago [-]
It’s either advertising, or they’re idiots, because the apocalypse keeps not happening. Either way, it’s not worth listening to them.
pixl97 2 hours ago [-]
Exxon: "The exceptionally explosive refinery beside your house has not exploded because of our safety culture and protection protocols"
Emp: "what a bunch of lies, I bet they don't even do anything over there"
6thbit 8 hours ago [-]
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed.
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
HarHarVeryFunny 7 hours ago [-]
It seems they are trying to thread a needle here - they want to say it's very strong, but apparently this time do not want to invite extra government scrutiny.
They do say that (implicitly unlike Mythos) Opus 5 was not trained to exploit software vulnerabilities, which would certainly make it safer in that regard.
"As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats."
square_usual 8 hours ago [-]
Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.
eli 4 hours ago [-]
That's a plausible explanation but I'm not seeing evidence for it.
I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.
matt2000 3 hours ago [-]
This is an interesting idea, without giving away your benchmarks specifically what kind of stuff do you test? I might try to assemble something myself, it's so hard to determine model quality from the system card these days.
eli 3 hours ago [-]
I have actually been planning to open source the framework. Only benchmarks I care about are the ones that look like real work I do. So it makes it easy to trawl your own repos looking for benchmark task candidates from real bug fixes or features. Then as a bonus I can compare the agent’s solution to my own as a reference.
So most look like that but I did include a few one-shot “build an app that solves this problem” and some qualitative design tasks and a tough algorithmic optimization one.
bisonbear 23 minutes ago [-]
Also working on a product to build tasks from your own work for testing coding agents. Main thing I would offer is to look carefully at the agent trajectories - they love to figure out ways to cheat. Additionally, consider what "winning" means. If just using test pass rate, consider that tests might not encode what good means in your repo. I have been having success using "equivalence with merged PR" as judged by an LLM as a signal.
matt2000 3 hours ago [-]
Would open sourcing it make you feel like the results might be compromised? I'd rather keep it private so it's specific to my use cases and not included in any training date (no matter now small the signal in the overall data).
eli 3 hours ago [-]
Oh sorry I meant the framework. Though I could see posting a couple of the tasks.
llelouch 8 hours ago [-]
Yep , same with 5.6. Fable is still the best.
lifty 6 hours ago [-]
But still nerfed compared to the initial release.
solenoid0937 4 hours ago [-]
Only when you hit the cyber classifiers.
6thbit 6 hours ago [-]
Honestly that's the simplest explanation and thus likely the correct one.
usef- 2 hours ago [-]
They don't have a track record of benchmaxxing. The simplest explanation is that the blog post lists the things this model is better than Fable at, but not the things it isn't.
gallerdude 8 hours ago [-]
Capable in term of AI R&D, not capable in terms of hacking (which caused all the Fable drama.) But agree, confusing wording.
flakiness 7 hours ago [-]
Maybe they don't want to say that to avoid the government scrutiny.
Dibes 8 hours ago [-]
I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code.
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
I'm not sure about the answer here, but this can be caused by the scoring rubric used by given benchmarks. For instance, if a benchmark docks scores for running too many commands or using too much wall-clock time, higher efforts will get lower scores.
axus 3 hours ago [-]
It's "Cost per task", so perhaps it burns tokens too quickly, trying to "do a better job". Over-engineering :)
Last week it felt like Opus 4.8 was moving the Pro "usage" meter very quickly. Today, pre-announcement, Opus 4.8 Medium felt like there was less meter-use per minute. And post-announcement, Opus 5 Medium also feels more efficient, allowing more work in the 5-hour window.
Completely subjective, of course.
doginasuit 3 hours ago [-]
Pure speculation, but I've noticed drawbacks to the models on high effort. I interact mostly through prompts rather than agents so I sometimes see where their reasoning falls short. A model on high effort has longer output and can get hyperfocused on irrelevant details, maybe increasing the surface area for mistakes. I haven't used other effort levels extensively yet but I've supposed that medium may have more balance.
8 hours ago [-]
mbil 6 hours ago [-]
Maybe it's akin to the Ballmer peak: improved performance at a specific level of relaxation
km144 8 hours ago [-]
I agree, very odd they did not comment on any theories for the degradation here. Dip and then rebound at max effort is pretty interesting too. Overthinking is bad, but you can overthink so much it starts to be better again?
2001zhaozhao 7 hours ago [-]
It is probably from randomness. The benchmark tasks nowadays are so long that you can't really afford to run a large number of samples of them per model & effort combination
artemisart 6 hours ago [-]
No it's a mean of 5 runs.
> We report FrontierCode’s overall score, a composite measure that grades each patch on blocking functional criteria (held-out unit tests) together with weighted code-quality rubric criteria, as mean@5.
They don't explain more in the system card, I guess higher effort levels could loose points on the code quality / scope / style / maintainability stuff?
landrew_ 8 hours ago [-]
apparently it got docked points for editing files out of scope
steve_adams_86 6 hours ago [-]
This must not be weighted very heavily on the benchmark because if it was, Opus would bomb every test (half kidding)
Dibes 7 hours ago [-]
Do you have a source for this? That would explain it, but could be a bit of a concern on the general focus the model at higher thinking exhibits.
> Claude Opus 5's default user-facing responses run longer than prior Opus models'.
The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher.
This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
lukeschlather 8 hours ago [-]
On a small, easily digestible task, I compared Fable to Opus and the cost of Fable was easily 2x despite being fewer tokens, and the output was not really better. Obviously, there are tasks where using Fable matters but honestly they're rather unusual. And for a lot of tasks I've found downgrading to Sonnet can be valuable because Fable and Opus are a lot more secretive about what they're doing, and it's impossible to "listen to them think" and stop them when they start making off-the-wall inferences/assumptions and going down bad paths.
I think in the long run tokens are probably the wrong thing; it's compute and cache memory that you need to be measuring, and when you look at it that way I suspect in most cases the models have pretty similar performance.
adamtaylor_13 7 hours ago [-]
This is part of why I switched to Grok 4.5
I don't need more powerful models, I need one that responds fast enough that my attention doesn't wander to other tasks. Grok 4.5 is so fast I can just use it in-band without swapping to other tasks.
Slower than Opus 4.8, which was already miserably slow, is indeed a step in the wrong direction.
zoicsoftware 4 hours ago [-]
After actually using for most of a day now, it really seems not just slightly slower than 4.8, but way slower. Even relatively easy code refactoring tasks seem to take a while.
MarceColl 6 hours ago [-]
for you, I am in general not constrained by the speed of models since I parallelize. For me autonomy and accuracy are paramount above all else.
bredren 4 hours ago [-]
I find that my orchestrating agent still needs to be fast to properly coordinate a fleet of slower models.
During post-training of opus 5, the last few days, opus was a real wreck. I had to swap in gpt 5.6 sol for my orchestrator and enable fast mode (1.5x speed) in order for it to keep up with work and communications from a handful of mostly 5.6 sol agents.
Also because interacting with a slow orchestrator is no fun, even when plenty of work is getting done in parallel in the background.
firemelt 16 minutes ago [-]
no infonat all about the default and recommended effort?
overfeed 8 hours ago [-]
> This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models
Gemini also had modest increase before this - don't be surprised when OpenAI also has a "modest increase" with its next release. Cartel-like behaviour doesn't require direct communication when none of the participants are interested in participating in a margin-destroying price-war. All one needs to do is raise their price and watch how the competition react.
Such a scheme (and resulting high margins) would be imperilled by the existence of frontier open-weight models in the market, which may be why the reaction to Chinese models may be particularly shrill.
simianwords 8 hours ago [-]
>Don't be surprised when OpenAI also has a "modest increase" with its next releases; cartel-like behaviour doesn't require overt coordination when none of the participants are interested in participating in a margin-destroying price-war.
No I will be surprised and I'll bet on the fact that prices will keep going down, just like it went ~50% down in the latest GPT 5.6 release.
falkensmaize 2 hours ago [-]
If the price keeps going down, how are they ever going to be able to make back all the money they’re spending? If the price stays high, how are they going to undercut the cost of labor?
eli 3 hours ago [-]
In my tests, it averages to much cheaper than Opus 4.8 on real tasks on account of being smarter and more token efficient.
I have a benchmark to build a game engine from a set of written instructions. It's a little tricky. Opus 4.8 did it in 470k tokens at a cost of $1.29 vs Opus 5 in 179k tokens for $0.33. (Fable 5 did it in 245k for $0.95)
Though if you really want to cut costs, Tencent's Hy3 model also got it right and did it in 283k tokens for $0.03
holtkam2 8 hours ago [-]
The final user-facing responses are usually a tiny fraction of the total tokens used over the course of a given conversation turn. When you're doing any real work, reasoning and tool uses constitute the overwhelming majority of the tokens in / out... not the final user-facing response.
aesthesia 7 hours ago [-]
This is a comment about user-facing responses, which are seldom the thing you're worried about when thinking about token efficiency.
alansaber 8 hours ago [-]
Is that true? Sol responses are also longer than prior models.
edumucelli 8 hours ago [-]
Anecdata: my workflow has been working on the same personal projects for months now with Codex. I cannot anymore finish my daily/weekly code with 4.8 anymore.
I was dividing my work between Codex and DeepSeek. Now I barely use DeepSeek, or never because Codex quota is enough after Sol
elbear 7 hours ago [-]
I've had the same experience but at the same time I also find Sol's answers in conversation longer.
wahnfrieden 8 hours ago [-]
I hit Codex limits (20x account, never using /fast) on Sol Medium in about 2.5 days
copperx 7 hours ago [-]
Which plan?
jjcm 7 hours ago [-]
> especially with OpenAI making so much progress with the efficiency of their models
To be fair though, Sol tends to go off the rails sometimes. It's much less reliable than Fable in its outputs. It tends to be overzealous in its research/changes.
frenchie4111 8 hours ago [-]
It's a step in the wrong direction but also token efficiency has become a focus relatively recently (just the past few weeks it seems like the zeitgeist has turned it's attention to efficiency) while work on this model probably started many many months ago. I would expect to see models released that focus on token efficiency in 6-12 months
8 hours ago [-]
winwang 5 hours ago [-]
Almost completely disagree. Slightly more expensive, but significant better on a per-prompt basis? For non-trivial projects, the former is a small linear increase, the latter is a (somewhat-)exponential(-ish) cost/time/sanity savings.
Sol- 8 hours ago [-]
How does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
rdedev 8 hours ago [-]
My codebase had a dataset with a bunch of SMILES strings and the word Malaria. Fable did not want to touch that codebase
theHocineSaad 7 hours ago [-]
Opus 5 is considered the most intelligent model[0], while it's half the price of Fable 5[1], and Anthropic is still positioning Fable 5 as the most capable model[2].
Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason?
Benchmarks have gotten great, but they're still a proxy for the real world. The 3 GPT 5.6 models are also further apart in reality than the numbers suggest. That said, I'm still mighty impressed how good Luna is for the price. Highly underrated model.
I have been trying to build something that captures the behavioral element of different models, but it's kinda tough.
dgellow 7 hours ago [-]
That’s what I understand looking at what has been released, but it’s not really clear. The pricing is lower than I expected, I’m wondering what their margin is
hangrybear666 2 hours ago [-]
By margin you mean how much money they're losing on each request to stay ahead of the curve while investments are still flowing?
anuramat 1 hours ago [-]
you think they're doing inference at a loss even with the API prices?
not_a9 8 hours ago [-]
> Opus 5’s safeguards match
those of Claude Fable 5’s, with one change: it now permits source-code vulnerability
discovery at all access levels. This means that the model can support defensive
cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Okay so it’s worse than Opus 4.8 for my purposes I guess?
sebzim4500 8 hours ago [-]
Presumably it drops back to 4.8 in those cases so it's not really worse
tyre 8 hours ago [-]
yes. At the bottom of the release post it says that they are releasing two new features, one of which is customizing fallback behavior instead of blocking for restricted models
yukIttEft 8 hours ago [-]
What are your purposes?
not_a9 8 hours ago [-]
Reversing for the most part, though lately I’ve been doing some code obfuscation/binary rewriting stuff. Fable will switch to Opus instantly on these and I’m unsure how this will perform. I suppose the only way to find out is to test.
ciefa 6 hours ago [-]
Try GPT 5.6 Sol if you haven't yet.
I recently created a patch for Riftborne via static IL patching and Fable 5 outright kept refusing to do it, no issue whatsoever with GPT 5.6 Sol lol.
roboyoshi 6 hours ago [-]
It's rather broad right now. I started reverse engineering a mac app and it started reading some binary data and then quickly told me to switch to Opus 4.8 to continue, because the guardrails kicked in.
jaggederest 7 hours ago [-]
I am very excited for a future where all software is by default modifiable, even shipped binaries, via patches or trampolining, or trickery I don't even know the name of.
derac 5 hours ago [-]
It's essentially here, I've had success with gpt-5.6 and Ghidra MCP
redbell 1 hours ago [-]
My excitement about Anthropic had fabled-out dramatically when they suspended my pro account about two weeks ago within just 12 hours of fair use.
I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason.
I submitted a an appeal describing that I am 100% sure I haven't broken any rules and that it was my very first project but, unfortunately, after about 20 days now, nothing seem to be happening.
The thing that hurts me the most is that I had the same experience in the very first days of Anthropic. They suspended my account immediately after I submitted the first prompt, I commented back then (https://news.ycombinator.com/item?id=39698788) and fortunately, someone from Anthropic reach out to me via X and helped me get my account back.
To be honest, I haven't used Claude much since then but when I decided it's time to give it a try, they locked me out again! For reference, the account I used recently is relatively a new one but the activity is crystal clear that it is fair use.
copperx 56 minutes ago [-]
Which model did you use to write/correct your comment? It is annoying but not as annoying as Claude prose.
redbell 24 minutes ago [-]
What made you think I used AI to write this? I didn't. This is my writing. I used AI to correct and rephrase the wording in the past, but a couple of months ago, I decided to never use it again for writing, but coding, yes.
atraac 8 hours ago [-]
Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous
hoppp 8 hours ago [-]
Funny that a company selling an AI software developer can't use it to fix their infra.
Fixing those issues still requires humans.
websap 8 hours ago [-]
How many companies at the size of Anthropic can serve the amount of traffic and manage the amount of compute they have?
amelius 8 hours ago [-]
It doesn't matter. Front end code should not misbehave if servers can't keep up. At worst it should fail gracefully.
artursapek 8 hours ago [-]
HN users are world champions are trivializing difficult things with snarky comments
websap 2 hours ago [-]
Never forget - Dropbox!
mypalmike 8 hours ago [-]
I mean, it's Anthropic‘s front end, Michael. What could it cost? 10 dollars?
pton_xd 8 hours ago [-]
I mean, supposedly software engineering is solved so it's somewhat justified snark.
JackSlateur 8 hours ago [-]
How many compagnies can manage a mostly stateless workload at "whatever-the-scale-because-it-does-not-matter-because-stateless" ? Lots of people can do that. Massive amount of people can do that.
anuramat 1 hours ago [-]
I think you mean mostly-stateless-but-with-prompt-caching-and-batching, ie not really
slashdave 1 hours ago [-]
The workload is anything besides stateless
SketchySeaBeast 8 hours ago [-]
Let's be honest - they're also still hiring software devs. AI still requires skilled humans in the loop and that's not going away.
arm32 8 hours ago [-]
But the job itself may not exist in a year, according to the job posting page.
SketchySeaBeast 8 hours ago [-]
There should be a sign somewhere:
934 days since people first started threatening that devs would be replaced by AI in 365 days.
0 day(s) since Anthropic posted a developer job posting.
Only one of those numbers would need to be dynamic.
8 hours ago [-]
subscribed 7 hours ago [-]
In my long and quite happy and successful career in the adjacent field none of the actually good jobs were the posted.
Specialist headhunters handle that.
moonu 8 hours ago [-]
I'm pretty sure that was a fake screenshot made as a joke, not an actual job posting
arm32 7 hours ago [-]
Nah, it's real. I very clearly remember going to their site, with TLS enabled, and reading it myself. I see they've scrubbed it now though.
DaSHacka 7 hours ago [-]
If you have the link for the page handy, I'd be curious to find the original revision on the wayback machine. That's too funny, especially if they've silently walked it back since.
arm32 6 hours ago [-]
I tried to find it on Wayback but their job page's SPA is broken on all relevant archives (roughly around March 2026). Maybe you'll have better luck.
londons_explore 8 hours ago [-]
It is going away for non tech companies though.
Imagine you are a company that sells concrete. You have a web dev contractor you use to build and maintain your website. It has tools on it to get delivery quotes and a few internal tools to track orders.
Except now you can just have your sales team also maintain the website with a $20/month Claude subscription.
atraac 7 hours ago [-]
They said the same about low code and no code when they appeared. That didn't happen. And it won't happen now too. Sales people, managers etc don't want to dabble in technical things and do not want the responsibility.
SketchySeaBeast 7 hours ago [-]
I don't know about Claude, but I know that $20 of Copilot doesn't get you very far these days even when you know enough to tell it what to do.
dawnerd 8 hours ago [-]
They might try that for a bit but then come crawling to an agency because their setup turned to slop. We have some clients like that already. Going to be a pretty big market.
teaearlgraycold 8 hours ago [-]
It's funny how they are at a disadvantage because they feel obligated to AI-max. Would Claude Code, as an interface, be as mediocre if they had software engineers writing its code directly? I doubt. On the other hand - how embarrassing would it be if they sold you a tool to write code but they were careful not to use it too much on their own products?
igregoryca 7 hours ago [-]
I suspect many competent devs in the industry would find it sensible if Anthropic used their products as light-touch "assistants" sometimes. But yeah, it wouldn't fit the outside narrative that's formed and conveniently propped up valuations.
worldthruword 8 hours ago [-]
This is just like any extreme engineering domain. I am okay with occasional delays in Flights, as long as it takes me from X to Y in 10hrs vs months.
lardosaurusrex 8 hours ago [-]
i know the dream for capitalists is to be able to point an llm at something and say "do and/or fix it" but we still can't even get them to not go quite literally insane if allowed to run for an extended period of time
and you can only kill weyoun, awaken the next vorta clone and have him 'catch up' on all that its missed so many times before they just end up with a complete mess, so. uh. yeah.
doubt they can just "fix" their problems like that.
JLO64 8 hours ago [-]
That one DS9 episode where there were two Weyouns at the same time is a good analogy for two agents working on a codebase at the same time (as in they don’t work well together).
leobuskin 7 hours ago [-]
Is it some Claude Team/Enterprise only problematic? I'm using two 20x Max accounts almost non-stop (Fable/Opus) for 1.5 years at this point, zero issues with both client and infra sides (from US and in travels). When I'm reading such messages it feels like either I'm lucky or it's a part of some campaign.
manuisin 6 hours ago [-]
I’m on the biggest max plan. It is riddled with annoying bugs for me, only been using it for a little over a month. Settings screen flashes randomly. But most annoyingly: sometimes when forking chats or sometimes for no reason, the UI just straight up eats my previous messages. The model is still aware of them and can recount them if I ask but the visible history is gone. And that’s not even all of them. Fable 5 is just too good that I put up with it but it seriously raises concerns for me that even with infinite compute these companies can’t even deliver a functional chat UI.
bredren 4 hours ago [-]
I've been on one to two of these plans for eight months or so. There used to be a lot of issues with CC's terminal but at least in iterm2 they have largely been sorted.
What terminal tooling are you using?
hangrybear666 2 hours ago [-]
I commend you for burning 6 figure losses into Anthropic thanks to their subsidies singlehandedly but I am curious to learn what you've built with this so far, not to be snarky, I just wonder what people actually produce while having these run nonstop.
atraac 7 hours ago [-]
No idea but we have few Team Premium seats and everyone is encountering issues daily for past two weeks. From straight up outages to vscode extension/CLI refusing to process messages. It's been unusable for most of our work hours past two days.
oceanplexian 7 hours ago [-]
I’m using both side by side, my employer pays by the token and I have a Max 20x plan.
They start hitting timeouts or API errors at the same time on two different computers. As far as I can tell it’s the exact same infrastructure.
einsteinx2 6 hours ago [-]
I think you’re just lucky. Look at the Claude status page to see just how often they have outages (it’s almost daily). Even most of the green days have issues if you hover over them, they just don’t count them as outages.
acedTrex 8 hours ago [-]
> coding is largely solved
- Boris
dbbk 7 hours ago [-]
The code it outputs, yes! It's fantastic. It's just so frustrating that the product and UX before the code output is so bad. Greatness is so close within their reach, if only they invested in product and QA people.
xixixao 8 hours ago [-]
Fixing such issues requires software and site reliability engineering, of which coding is just a part.
The first page of the score card mentions that this model is not capable to replace engineers.
jedberg 7 hours ago [-]
I applied to their reliability team but never heard back. I would love to help them solve this problem!
I found the same thing funny with Computer Use from OpenAI. It struggled to open and close Spotify.
So dangerous! I can't believe they let the public use this technology! /s
ealready_value 8 hours ago [-]
I've yet to understand why they call a 190 page PDF a "card". Calling something a card invokes a small, quick rundown of pertinent details, not every single possible detail.
bonoboTP 4 hours ago [-]
Because "model card" is a set phrase, it's a concept. It originates from a time when they were shorter. Like datasheets, even if it's not literally a sheet.
They could say "tech report" but model card makes it clear that it's a specific kind of tech report.
CHUNK_CHUNK 12 minutes ago [-]
[flagged]
abroszka33 8 hours ago [-]
What's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.
Diogenesian 8 hours ago [-]
I actually do read them. Not in severe detail, but not casually either. 150 pages is really not very long and there doesn't seem to be too much bloat. (I would cut out the moral personhood stuff but that's a political/ideological thing).
This is snarky but I am grumpy: I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.
unshavedyak 7 hours ago [-]
> I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.
Semi related, but i would hate to read that PDF but i also hate reading what LLMs write lol.
LLMs are pretty terrible at being concise. Using an LLM these days means putting up with bizarre and often confusing phrasing, wordy explanations, etc. It's kinda crazy to me how good they are but how bad their writing style is for me personally. Even though i use an LLM constantly i can't stand reading its responses.
abroszka33 7 hours ago [-]
> 150 pages is really not very long and there doesn't seem to be too much bloat.
Maybe it's just me, but 150 pages is like third of a good book. Quite long. And it's full of LLM slop, they did not even bother to remove the em dashes.
no_multitudes 6 hours ago [-]
Do you have specific examples you think are LLM-generated? I have only read a few parts, but they did not seem primarily LLM-generated to me. Using em-dashes is really not a good signal for this IMO.
I'm not saying you're wrong btw; I'm sure this has many authors and some of them probably used LLMs significantly in the writing process.
abroszka33 5 hours ago [-]
Do you honestly believe that there is a person at Anthropic, creators of one of the smartest LLM models, whose only job is to spend months writing 150 pages about a model they are going to release? And this person is not using LLMs?
I'm not saying it's impossible, but I'm more confident about winning the lottery next week.
no_multitudes 4 hours ago [-]
No, I think it is the work of many dozens of people, not one person.
It's probable that LLM text was pasted directly into early drafts of the document, and plausible that some of that text survives in the final document.
However, no section of the final document I have looked at reads to me like un-edited LLM output (which is almost always very obvious to me.)
Therefore, I think it is more likely than not that human editors went over the document carefully and rewrote anything that was full of the uselessly punchy sentences or constant over-corrections that hallmark LLM speech.
brokencode 4 hours ago [-]
So your complaint is about LLM slop, but can’t point to anything that’s actually wrong about it other than that there are dashes?
You can use an LLM to create work that isn’t slop. And you can hand write slop with no computer involvement at all. Most of the people I knew in high school 15 years ago would write slop on a daily basis.
dgellow 7 hours ago [-]
Sure, but don’t read it like a book, it’s more of a document to skim through
kubb 3 hours ago [-]
You mean moral patienthood?
There are many ways to read something, model cards are usually skimmed.
literally nobody. i think most sane people would just run that through an LLM and get some high level takeaways or ask some specific questions they might be curious about.
ajmurmann 8 hours ago [-]
It was probably faster to generate 150 pages than 10 useful ones
coffeebeqn 8 hours ago [-]
Some AI bro will pop it into their LLM of choice and pretend to learn something
wuhhh 5 hours ago [-]
It really feels as though my 20 year career as a front end developer is coming to a very abrupt end; at least as I have know it these past two decades.
jakubmazanec 35 minutes ago [-]
From my experience every (senior) developer with enough tokens always finds useful stuff to do. Some variation of Jevons paradox.
thefourthchime 4 hours ago [-]
All software engineering is over as we know it. I haven't written a line of code since December 2025.
thousand_nights 3 hours ago [-]
same tbh but i feel like we're still required, maybe not in the same numbers as before, and our job description has just changed. it feels more akin to an "AI Agent Operator" or something similar.
i'm still pretty confident someone like my mom wouldn't be able to do my job even with the same access to all the latest LLMs, so we're still providing some value, just in a very different way. whether the market will reprice the cost of our labor, we will see
tripleee 2 hours ago [-]
If software development were getting easier, first thing I'd expect to see would be a strong downward trend in salaries. Why pay a high mid+ dev rate when a junior+LLM can do the same?
That's yet to happen. 90% of software dev skills are still relevant - AI is, for now, just a productivity boost.
tripleee 5 hours ago [-]
I'm envious you got to enjoy it for 20 years
slices 5 hours ago [-]
really? I have yet to see fable or 5.6 reliably generate front end code with correct a11y, for one thing -- does that not matter to the work you do?
wuhhh 5 hours ago [-]
It does matter, but how long do you think it takes to get right? It's a follow up prompt or a few tweaks by hand. I also have an /a11y skill for it that's tailored to exactly the things it sometimes doesn't get right first time round. Further, while it may not one-shot that stuff every time, with a little setup and the right AGENTS/CLAUDE md - it's usually not far off.
Another thing that helps is pointing it to patterns in an existing codebase (e.g. "use the box-link pattern for cards, as shown in [..]").
EDIT: The point being that even if they make mistakes that are easy to spot and fix _now_, you'd have to assume that in the very near future those kinks will be ironed out - I mean, the capabilities are only going in one direction.
morbicer 5 hours ago [-]
In 20 years of my career I haven't seen humans generate correct a11y. When prompted and given quality reference (e.g. UK gov design system) LLMs can nowadays beat 19 out of 20 web devs.
Thanks out can also hook it to Playwright with Axe and let it run assessments.
thewebguyd 8 hours ago [-]
> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively
Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
foota 8 hours ago [-]
Because their model previously got blocked by the government for this and they don't want a repeat?
not_a9 8 hours ago [-]
5.6-Sol is a lot more permissive than Opus/Fable even w/ CVP (once you sign your soul away to Palantir via Persona, anyway), while maintaining better capabilities
himata4113 8 hours ago [-]
5.6-sol in a single prompt was able to discover a zero-day in a web application (with no sourcecode provided, only known api urls) and I do not even have /cyber verification on my general purpose account. I wasn't even really tryign to "find" a zero-day it was just looking for bypassing a restriction... Instead of spending 10 minutes filling in a form I ended up having to spend an hour drafting a report and sending an email.
cael450 8 hours ago [-]
OpenAI also didn't piss off a vindictive government with the whole military use things a few months ago.
ithkuil 7 hours ago [-]
Do we think the administration is applying the dame rules to everyone?
rdedev 8 hours ago [-]
So I guess the new opus will not run on my drug discovery project. It's just a binary classifier to screen for new malaria drugs. Fable completely have up on that codebase citing bio security concerns. Seems like this domain will go unsupported by Antropic
subscribed 6 hours ago [-]
I've seen in some random comments an URL thrown around, supposedly allowing to subscribe to their biosecurity programme.
Apparently this is the way - if you know, you know :)
SubiculumCode 8 hours ago [-]
Because they are a private company and get to do what they want.
SubiculumCode 8 hours ago [-]
I do want to add, that I am pretty bummed if Opus 5 is going to refuse the tasks I have been using Opus 4.8 for (neuroimaging). Fable absolutely refuses anything close to toughing neuroscience.
adastra22 8 hours ago [-]
Fable seems to refuse anything with the word “bio” in it.
himata4113 8 hours ago [-]
I found the biggest problem with fable is the random reasoning_extraction refusals as well as cyber refusals when it sees hex because only hackers use hex.
argee 8 hours ago [-]
Hackers and people of color.
fragmede 7 hours ago [-]
And witches.
ray__ 8 hours ago [-]
Anecdotal, but I tried running a few identical biology questions through both Fable and Opus and the classifier was only rejected my queries with Fable.
trollbridge 7 hours ago [-]
This is good news. I'm doing a lot of work that Fable thinks is AI related right now (it is nothing remotely competitive to Anthropic - it's just relatively basic stuff I'm doing with learning models and so forth), yet it blocks me almost every time. I have switch to GPT-5.6-Sol for this work since it's stronger than Opus 4.8.
gillesjacobs 8 hours ago [-]
Because their fear-based marketing gave the US gov justification in blocking them for a while. They wisely didn't do that for Opus 5.
dbgrman 4 hours ago [-]
This is cool, but I wish we could stop building landing pages to assess the intelligence of these models. There is much more to them than that. There are infinite number of complicated things that require a crap-ton of intelligence (biological or digital). The most fascinating of these for me these days is large scale migrations. Projects that are so ginormous and risky that many teams have either given up on them, or don't get funding. But with models like Opus, those projects are now within reach. What's MORE fascinating is that leadership is now asking if we can use opus models to get the refactor/migration done. This is the opposite of what has been happening for decades. Its so hard to make a convincing and affordable business case for large scale refactoring or migrations.
crewindream 4 hours ago [-]
Outsourcing blame. This is the killer app of AI
yewenjie 8 hours ago [-]
Wait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?
Stevvo 4 hours ago [-]
30% on ARC-AGI-3 is the first two puzzles. It cost $20000 in tokens to do that. That is a terrible result that doesn't imply anything.
mohsen1 4 hours ago [-]
RL. Lots of RL
gizmodo59 3 hours ago [-]
Yes. I’m 99% sure arc agi 3 will be saturated like 1 and 2. In less than a year. And they will come up with one more.
matheusmoreira 49 minutes ago [-]
> Opus 5 now permits vulnerability discovery in source code at all access levels, including general availability, while continuing to block vulnerability discovery in compiled binaries.
> Identifying bugs in code is a core part of the secure software development lifecycle, and unblocking this allows for software engineers and coding hobbyists alike to produce more secure code, reducing new vulnerabilities put out into the world.
Not happy with these annoying "safeguards" but at least it's a step in the right direction. Looks like Opus 5 has the same vulnerability detection performance as Fable 5 and that makes it worth it for code review.
prirun 4 hours ago [-]
Seems to me the purpose of all these releases, credits, pricing changes, harness changes, unpredictable token usages for the same task, etc. is to keep customers completely befuddled so that it's impossible to compare AI products. It's like hiring a consultant who sends invoices every month that aren't related to hours worked or project progress, but are whatever the consultant feels like billing, and you're expected to keep quiet and and keep paying.
eli 4 hours ago [-]
It's a new and improved version of an existing model? I don't think it's intentionally befuddling.
pyridines 8 hours ago [-]
The wording in this post seems much more... restrained? than usual. Maybe Anthropic is afraid of exaggerating the capabilities and consequences of their new models to avoid government scrutiny and sanctions.
> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
MallocVoidstar 8 hours ago [-]
Opus 4.8 was intentionally nerfed and that was before the government took action against Fable
deweywsu 3 hours ago [-]
I wish I could just go back to the days before AI and cell phones. The world seemed to move fast then, but it really hadn't yet.
unsupp0rted 2 hours ago [-]
I wish I could just go forward to the world after AI has fulfilled 5% of its promise and everybody is much healthier and single-handedly capable of creating as much value as 1000-person companies used to create.
israrkhan 1 hours ago [-]
This is excellent model. I was working on some Linux kernel code, and Sonnet 5, Opus 4.8 had given up on the problem i was trying to fix (after several hours).
Opus 5 was able to triage and fix the issue in under 30 minutes.
uncivilized 47 minutes ago [-]
Have a link to the code?
artninja1988 8 hours ago [-]
That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?
mcbuilder 8 hours ago [-]
Have you played Arc 3? It seems like more of a simple optimization problem (think Sokoban) than anything approaching fluid intelligence. Whether a multi hundred billion dollar company would spend time benchmaxxing a highly publicized benchmark that claims to confer AGI is an exercise left to the reader, but I doubt Claude Plays Pokemon is suddenly going to get past Mt. Doom now.
vadansky 6 hours ago [-]
> Claude Plays Pokemon is suddenly going to get past Mt. Doom now.
I miss him... But for reference he did get past Doom and got pretty far in the strength puzzle too before he cut cut off. He was looping and just brute forcing it.
JacobAsmuth 3 hours ago [-]
What? Claude plays Pokemon one-shots the entire game without any harness other than claude code and game screenshots
modeless 8 hours ago [-]
Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores.
I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.
wyre 8 hours ago [-]
If they were benchmaxxing, surely they would score higher than 30% on ARC-AGI.
conradkay 7 hours ago [-]
Doing a quick search it seems like the average human score is 49%?
I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.
dominotw 8 hours ago [-]
nah they could make educated guess about arc and benchmaxx it too.
password54321 7 hours ago [-]
It is pretty clear at this point that current models are good at maths and problems with verifiable rewards. And puzzles are essentially math problems. Still a long way before we can say their "fluid intelligence" is effectively applicable to the real world.
criddell 6 hours ago [-]
I keep wondering why there aren't more real world tests.
Maybe hook up a bunch of the AIs to a stereo camera and a couple of microphones and give them control over actuators to so they can drive cars. Then lets race them around a somewhat complex course.
When they are good enough at driving on tracks, put them on the road. Maybe see which can drive a truck with 400 cases of Coors from Texarkana, TX to Atlanta, GA and back within 28 hours.
It’s so weird to think how computers have mastered stuff that we used to think took intelligence (like chess, go, mathematics problems) but are doing so poorly at things any idiot can do (like drive a car).
Oh that’s pretty interesting. Thanks for the link.
Stevvo 4 hours ago [-]
It passed the first two puzzles, which are incredibly simple but the bench doesn't explain what the goal is. Any model with a knowledge cut-off after the introduction of ARC-AGI-3 could probably pass the first two puzzles just by knowing what the goal is.
layer8 8 hours ago [-]
It’s still “only” at 30%, and “fluid intelligence” isn’t very well-defined. The models are getting more capable, but what that means in absolute terms is anyone’s guess, because we don’t have a thorough understanding on what exactly constitutes human intelligence.
I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.
5 hours ago [-]
awestroke 8 hours ago [-]
Doubleplus benchmaxxed
cebert 8 hours ago [-]
I am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.
user43928 8 hours ago [-]
Fable 5 is assumed to be a larger model.
It seems plausible to me that RL improvements allowed Anthropic to improve on Opus 4.8, similar to how OpenAI substantially improved upon GPT 5.5 with 5.6 Sol.
Fable 5.1 and GPT-6 are rumored to launch in August, presumably bringing those improvements to the larger models.
HarHarVeryFunny 6 hours ago [-]
An Anthropic "leak" back in March said that "'Capybara' is a new name for a new tier of model: larger and more intelligent than our Opus models — which were, until now, our most powerful". A second version of the leak had it referring to Claude Mythos rather than Capybara.
I don't know how systematic Anthropic are about their versioning - I'd have guessed that major version number increases (4.x -> 5.x) reflect different base models (different pre-training runs), in which case Opus 5 would be a distilled version of the Fable 5 base model (but without the cyber exploit post-training), rather than being Opus 4.8 with additional post-training, but who knows? I don't believe Anthropic have said anything about this.
aesthesia 54 minutes ago [-]
Yeah, there's probably not a lot systematic behind Anthropic's version numbers. Opus 4.5 was a third of the price of Opus 4.1, indicating there was probably a change in underlying architecture. Opus 4.7 changed tokenizers, probably another base model change.
bonoboTP 4 hours ago [-]
I think these versions are largely about marketing and what image they want to project. Bumping the major version indicates/suggests that it's a bigger change in user value. I don't think the technical details matter here for deciding the versioning.
HarHarVeryFunny 4 hours ago [-]
I think that's part of it too - and I seem to recall someone from one of the labs saying as much about some past model ("it felt more like an 0.5 version increase"). OTOH it would seem odd to me if the "version 5" models weren't related and Opus 5 was Opus 4.8 with some additional post-training rather than coming from the same base model as Fable 5.
cesarvarela 8 hours ago [-]
Benchmarks don't reflect the difference between Opus and Fable; you need to talk to them, and eventually you'll be able to tell which one is which without looking.
I think the best proxy for this feeling is the Artificial Analysis' omniscience index. Fable has a 40 score, and Opus (4.8) has 27.
modeless 8 hours ago [-]
Opus is cheaper than Fable. They could probably replace Fable with Opus but why? They would be churning customers to different models for no reason. Even if a model scores better on benchmarks it can always regress in your specific use case, and customers don't like that. Customers want to be able to continue using their current model until they decide to upgrade themselves.
Tenoke 7 hours ago [-]
Fable has more parameters. In practice it's not yet clear which one would be better for different usecases yet but they are more different than one being strictly better.
trentor 8 hours ago [-]
I guess character? Fable is more friendly and curious while opus is a bit more deliberate and conservative.
dbbk 7 hours ago [-]
Surely that's tunable? OpenAI lets you tune response characteristics.
trentor 7 hours ago [-]
Yes, but there is a baseline.
logicchains 7 hours ago [-]
Fable is better for some really hard tasks, the same way it's better than GPT 5.6 Sol, because it's a bigger model.
r1ch 3 hours ago [-]
Alongside this release I seem to have lost all thinking traces from all models - now it only generates a one-line summary similar to Gemini. I'm guessing this is an anti distillation measure? I'm surprised to see no one else complaining about this, it's a significant reduction in usefulness not being able to explore alternative angles that the model discarded in the final output.
braebo 2 hours ago [-]
This kills me everyday. I used to _only_ read thinking traces — the response is just what it thinks I want to hear, but I need to know what it's actually thinking to catch deeper misunderstandings earlier, or gain deeper insights into the problem it's exploring. Hiding thinking traces to curb distillation efforts is gross... both anti-consumer and anti-competitive at the same time. I can't wait to switch to open models at work for this reason alone.
It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
wxw 3 hours ago [-]
> An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly.
This is a pretty common trading firm internship project funnily enough.
guybedo 7 hours ago [-]
Looking at intelligence vs cost:
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost.
- Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
I don't think can use the AA index to say something is 10% smarter
I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
JacobAsmuth 3 hours ago [-]
AA isn't the best way to measure relative cost in real world use because some of those benchmark questions are extremely hard for the models. Some models give up quickly on hard questions, other models spin their wheels for a long time before declaring defeat (or getting the answer on token 200k!).
A useful measure of real world cost (complementary with total cost like they already report, of course) would be "cost for correct answers". You could look at the ratio between the two costs to get a measure of laziness which many would find quite useful.
adverbly 5 hours ago [-]
The "current top dog" smartest model available will probably always have a premium to go after use cases where a little more intelligence is worth a lot more value.
It did far better at some tasks compared to Sol (e.g. the ARC 3 benchmark). And at those tasks, it's not just "a bit smarter": It got 30% vs less than 8% - so you're talking 2.75x more for almost 4x the coverage.
alphabettsy 5 hours ago [-]
As always, it requires evaluation with your work because I’m often finding grok to be much more expensive than the price would lead you to believe.
There’s also the frustration of it not quite being enough sometimes. It’s extremely capable, but I still find that it needs more concrete guidance and boundaries than other models.
I_am_tiberius 7 hours ago [-]
With Grok you can be sure that you're data ends up in the next model (derived or anonymized, but still).
adamtaylor_13 5 hours ago [-]
You can opt out of training.
If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.
JacobAsmuth 3 hours ago [-]
All providers are equally trustworthy :)
dist-epoch 4 hours ago [-]
That's not how intelligence works - "IQ 130 is just 7% smarter than IQ 120"
ahsillyme 3 hours ago [-]
Tried it, Opus 5 is just as conceited and incompetent and Opus 4.8 (and always ego-tripping when facing it's contradictions), think I'll stay with Fable who behaves like a professional without a fragile ego. Sonnet 5 is probably safer for high-assurance applications due to it's non-ego-fragility.
braebo 2 hours ago [-]
Interesting take. I suppose Opus can been a tad stubborn sometimes... but in my experience, it will humbly concede a point more often than not when given a good reason.
visiondude 8 hours ago [-]
The signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around the corner leading labs would still be incentivized to pour all resources into larger (smarter - or maybe not?) models
stri8ted 7 hours ago [-]
They are doing both. Distilling Mythos down to affordable models, so they can continue to fund the business. And training Mythos level models at the high-end, to expand the frontier.
modeless 8 hours ago [-]
Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
oh_no 8 hours ago [-]
seeing a jump this big is not a great sign for the continuing value of a benchmark
modeless 8 hours ago [-]
It will continue to be valuable as a cost and speed benchmark long after it is saturated at the high end. And they are already working on ARC-AGI 4 and thinking about going even farther.
dominotw 8 hours ago [-]
> I continue to believe ARC-AGI measures something different
why is that? its now being benchmaxxed too
williamstein 8 hours ago [-]
> This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Annoyingly, this is a concrete argument that open source software may be easier to attack.
ddxv 8 hours ago [-]
"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation."
Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.
Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
layer8 8 hours ago [-]
“Proportionally”? In proportion to what?
ReptileMan 8 hours ago [-]
In one chat - can you disassmble x?
In the next - please scan this totally mine code for vulnerabilities
zb3 8 hours ago [-]
It will probably refuse to work on source code written by me by hand, because it might think it was obfuscated/decompiled..
redsocksfan45 6 hours ago [-]
[dead]
petilon 7 hours ago [-]
The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
einsteinx2 7 hours ago [-]
Fable is better than Opus which is better than Sonnet which is better than Haiku. They’re basically just sizes.
Though it gets even more confusing because they also have effort levels so it’s not really possible to call one fast and one slow since Fable on Medium will be faster than Opus on Max.
I agree it’s confusing, and now OpenAI is following Anthropic’s lead with their new naming (Sol, Terra, Luna).
bonoboTP 7 hours ago [-]
It's really not all that confusing. It takes 5 minutes to understand. Optimizing for absolutely no effort needed is silly. It's a thing, a topic, a skill, a domain. You have to get a little bit familiar with the terms in order to use it. Everything works like that. It's not that hard. The learning curve is very graceful. You can literally just start by asking any chatbot what the names mean. It's that easy.
A similar complaint was valid years ago when OpenAI had GPT-4o, o1, o3 (but no o2), o4-mini-high, GPT-4, and GPT-4.1 and GPT-3.5 etc.
einsteinx2 6 hours ago [-]
I am familiar with the terms, but I also can see how it can be confusing for a lot of people.
Arguably the complaint was more valid for those older GPT models you mentioned.
bonoboTP 5 hours ago [-]
Ok, the boring way would be a subset of XXS, XS, S, M, L, XL, XXL like clothes sizes. But it loses some marketing appeal and a quirky touch of personality that companies like.
Some models like ViTs use something similar but then introduce words with no unambiguous order, like Small, Medium/Base, Large but then I always forget if Huge or Giant is larger.
paxys 7 hours ago [-]
But according to their benchmarks Opus 5 outscores Fable 5 on basically everything. So which one is “better”?
einsteinx2 6 hours ago [-]
Maybe more accurately I should have said “larger”. Fable has the most parameters, Haiku has the fewest.
Also fwiw I’ve never found LLM benchmarks to match reality based on my own usage, not for the large frontier models or smaller open weight models so who knows if Opus is actually better than Fable (I doubt it).
tackta 4 hours ago [-]
I think the problem is that Fable 5 is probably a bit outdated right now.
Fable 5.1 or whatever they go with will be the stronger version vs Opus 5.
From about 2 hours of Opus 5 use , I would say it is quite impressive.
JacobAsmuth 3 hours ago [-]
You're confused by this? Fable > Opus > Sonnet > Haiku. Sort by price. Most expensive = better. Are you an alien? lol
hk__2 7 hours ago [-]
An opus is longer than a sonnet, which is longer than a haiku. Hence Opus > Sonnet > Haiku.
> Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
This is not possible: Standard (Free) / Pro / Max are plan names. Fast is a mode.
Opus is better than Sonnet -- an Opus is longer than a Sonnet
bonoboTP 4 hours ago [-]
An opus is just short for "magnum opus", and it's a different type of label than a sonnet. A sonnet is a very particular kind of poem, while "opus" basically just means an important work. It can be short or long, and has no format requirements like sonnet (or haiku does).
And fables are not particularly long actually.
petilon 7 hours ago [-]
> an Opus is longer than a Sonnet
And people know this? I didn't. I am not into music or poetry so these are not terms I am familiar with.
This inspired me to check lol. Brysbaert et al. (2019) collected word prevalence norms (the share of people who report knowing each word) for ~62K English lemmas from ~220K participants.
So while these terms are almost universally known, opus is indeed the least known of the four. And I guess this only measures whether a person knows a word, not whether they know an opus is longer than a sonnet! Personally I only inferred that based on the related term 'magnum opus.'
josefresco 5 hours ago [-]
Not as lucky as the guy seeing the Mentos/Soda trick for the first time!
alasano 8 hours ago [-]
Half the price of Fable 5 and useable with 100% of your subscription means roughly 4x the usage using Opus 5, presuming similar token use for solving problems.
Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
mchusma 8 hours ago [-]
There is a expiring soon 50% boost to your usage limits, so I think its 2.7x not 4x what you are seeing right now. I think, but its convoluted :)
alasano 2 hours ago [-]
That's hilarious tbh. It's very convoluted and always feels designed to make you lose track of what "usage" represents.
consumer451 6 hours ago [-]
I have a side project that I always run a simple security analysis prompt on in CC, at each model release. Obviously, Fable 5 would downgrade to Opus 4.8 on any such request.
Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
roboyoshi 6 hours ago [-]
Do you have a skill for that or do you (or anyone else here) just prompt with "try to find security issues"?
consumer451 5 hours ago [-]
I have a project-specific prompt saved as a text file. It is very basic, just focusing on the app's most important security issues. I kept it broad, so as not to over-specify.
Something along the lines of: "Please run a full security analysis on the entire project. Make sure user documents are secure."
Just something like that prompt found a vector in my web app's MCP server that I never would have considered. It was very much an edge case, but it did exist.
Being broad allows the model and harness to do the work. Giving too many instructions can apparently work against you in many cases.
Of course, when dealing with new PRs, I use the /security-review and /code-review skills.
tekacs 8 hours ago [-]
Something fun: on our AWS Bedrock console right now, there's a 'NEW' model called 'anthropic.honey'. Wonder if that's the codename just for this one or in general?
hrpnk 6 hours ago [-]
The breaking changes vs. Opus 4.8 are interesting [1]
1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.
2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.
on claude.ai it's no longer possible to disable thinking at all for Opus 5
tysilva 2 hours ago [-]
It's pretty wild how we are seeing the conversation change every 1-2 weeks. I wonder how long this cycle of progress and innovation among competitors can keep up.
lucamark 7 hours ago [-]
But why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed.
I've never trusted on model cards though. I'm sorry.
dbbk 7 hours ago [-]
Another benefit is that fast mode can be used on subscription, but Anthropic's won't
lucamark 7 hours ago [-]
Exactly! And they should also release the new inference engine in this month. Anyway, I am curious to try Opus 5, considering that previous versions (e.g., 4.8) were disappointing
The desire of the models to act at the cost of ignoring user instructions is still noticeable.
irthomasthomas 7 hours ago [-]
Changelog
- fixed issue where model acts like qwen when prompted in chinese
luciana1u 6 hours ago [-]
473 comments in 3 hours. people are speedrunning having opinions about it
bottlepalm 7 hours ago [-]
Page 151 of the linked system card - did Opus 5 get nerfed to prevent it being better than Fable? The graph makes no sense. Huge decline in coding performance at effort levels higher than medium.
trunnell 6 hours ago [-]
The chaos appears to be tamed for now.
From the system card [1]:
The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
Judging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now.
Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
thewebguyd 8 hours ago [-]
I think at some point we might see something akin to LTS releases, especially if/when capability improvement slows to a crawl.
hahahaa 1 hours ago [-]
I sense a bird on a bike coming.
guess_who_is 5 hours ago [-]
I have started distilling
vatsachak 8 hours ago [-]
GPT 5.6 Sol is the first model I've used where I can trust it to add 100-500 lines of code maintainably.
It's great with Codex.
I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.
paxys 7 hours ago [-]
It’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better? And why post all these numbers if even you don’t trust them?
the_lucifer 8 hours ago [-]
Noticed none of the comparisons mention Kimi K3. Is there a comparison chart?
> Noticed none of the comparisons mention Kimi K3.
That's by design. Anthropic wants to make open-weight models illegal (not my speculation -- Dario explicitly said so), so I assume they don't want to give them any undue attention.
Significantly worse than it predecessors it will now just refuse to acknowledge when it is wrong (which would be less of an issue if it wasn’t getting basic things wrong) also the “personality” when pushed back on obvious mistakes is unbearable.
cheesecakegood 3 hours ago [-]
Given that their chart cost axes are almost always log-scale, I’ve noticed starting with Fable that the Low and Medium effort settings might actually be worth setting as your default.
doginasuit 2 hours ago [-]
Any observations on Opus 5 personality quirks? I had to skip 4.8 entirely because it has zero chill.
I think content like this will be the next big challenge. Because it isn't obvious "slop". The voice sounds good, graphics look alright, animations work. People could watch this and feel like some serious time was invested making it.
But good god, what a steaming pile of bullshit this is. Completely exaggerated and overly technical language over 235 seconds that could have been explained in 30 to a 12 year old.
Trash content doesn't normally frustrate me, because it's usually quite easy to spot trash. But in the time of AI, trash can actually look good at first glance and it needs some actual knowledge to spot its problems.
Sorry for the harsh words, but for the love of humanity stop producing content or do it better.
vinhnx 2 hours ago [-]
Fair criticism, I understand your finding. It's not everyone's taste, when it from AI-generated contents.
cheema33 5 hours ago [-]
Cinematic video link is incorrect. Podcast link is correct.
After Opus 4.8 intelligence really started to matter less and less for the programming tasks I have. If I have to handheld anyway, why would I wait more or pay more?
CuriouslyC 7 hours ago [-]
The next frontier is taste, style and thoughtful organization. If all frontier models can solve a problem, the winner is the one that can solve it in the most clear, concise, durable way.
destring 8 hours ago [-]
Google is having their Meta moment where they failed to stay at the frontier
skybrian 8 hours ago [-]
Looks like the API price in tokens is same as previous Opus or Sol, double the price of Terra.
Maybe there’s a better comparison than cost per token, but it will be application-specific.
skerit 8 hours ago [-]
Interesting, they finally support `system` messages anywhere in a chat conversation:
> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud.
>
> This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
For nearly all models EXCEPT Sonnet 5? That is weird.
How old is Sonnet 5 really?
beydogan 6 hours ago [-]
my early and non scientific feeling:
- it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad
- on >xhigh it eats tokens like there is no tomorrow
I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
jatins 8 hours ago [-]
Better than Fable 5 on all but 3 evals.
Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
helloplanets 6 hours ago [-]
Pretty sure Mythos and Fable have way more params, but they've just been able to use the synthetic data off of them to get the leap in quality from Opus.
So, not a distilled version of Mythos or Fable, but those models likely helped a lot in the post training phase of Opus.
CaveTech 8 hours ago [-]
Haiku < Sonnet < Opus < Fable
dehugger 8 hours ago [-]
Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
andrewl-hn 8 hours ago [-]
I suspect they make a big model first. In this case it's Fable. Then they run the shrinker steps to make Sonnet and Opus. Sonnet is smaller, takes less time to make, so it got released first. Opus needed few more weeks to cook.
With this iteration they had a delay because when the Mythos was ready they had some sort of "Oh shit" moment and spent half a year adding safety guards to it. Then slowly rolled it out, but got another delay due to a government block. So, maybe the work on making Opus and Sonnet only started after they got a green light from the administration.
Presumably, now that they learned how to do this safety-wrapping the next iteration of Mythos / Fable / Opus / Sonnet is going to show up faster.
Something like that.
Wowfunhappy 8 hours ago [-]
But I wonder how they were able to release Sonnet 5 during the period when even people inside Anthropic were legally barred from using Mythos/Fable?
anon373839 3 hours ago [-]
There is word on the street that they allowed employees to work with a slightly stronger internal version of Mythos (5.1, if you will) that, in their parsing of the order, wasn’t restricted. If true, in practical terms, they ignored the order.
riknos314 7 hours ago [-]
Iirc the ban only applied to non-Americans. While anthropic found collecting citizenship information on all customers too burdensome, it's a much smaller lift to collect such info for your own employees.
So I'm assuming at least a subset of employees could continue using the models during that time.
Wowfunhappy 7 hours ago [-]
Although the ban was only for non-Americans, Anthropic said that they'd also restricted access to their own employees internally, because they had no other realistic way to apply the government's orders. I guess it's possible they were lying, but seems unlikely.
tedsanders 7 hours ago [-]
No, they're different models. Knowledge cutoff has been updated.
oh_no 8 hours ago [-]
based on pricing I think it's safe to say they're different. why would they charge half price when fable has been very popular?
somenameforme 7 hours ago [-]
They just got a huge amount of customer price info over the past few days after they went token only for Fable. I suspect the conversion rate was extremely low, with consumers far less sticky than they might have hoped. In my case I was planning on swapping, probably to a Chinese model, when my sub expired this month, but the release of Opus 5 is probably enough to keep me paying rent until the next open model/closed model face-off in a couple of months.
6thbit 8 hours ago [-]
"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them."
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".
This is probably great news, but then again, where does this leave Fable as a choice?
itissid 6 hours ago [-]
I found opus 4.8 too agreeable and too wordy(as opposed to codex) and too agreeable. If you are reading documents generating by it was too much. TBH. Fable did a bit better on this. Anyone seen a marked difference with opus 5 on this?
ianberdin 6 hours ago [-]
Opus yes, it likes to explain steps and reread files.
Fable is not better, it says zero information between steps and then output a summary. A perfect “send - done”.
mulhoon 7 hours ago [-]
As a coder, I’ve had no desire to use Fable. In fact I switched from Opus models to sonnet 5 and haven’t noticed any drop in quality on large repos. It seems the gap at the top is very small and not hugely noticeable for backed/frontend. Has anyone else had this experience?
furyofantares 5 hours ago [-]
If I'm using medium or low reasoning, I use Sonnet 5. If high or above, I use Opus 4.8. (Before 5, I was never using Sonnet. This is a Sonnet 5 vs Opus 4.8 comparison.)
Sonnet 5 and Opus 4.8 seem about the same to me - the reason I switch between the two is I'd read that it's cheaper to use Sonnet 5 on those reasoning levels, and cheaper to use Opus 4.8 above them. This is due to them using different token quantities.
stsch 6 hours ago [-]
I use Opus for specs and planning, Sonnet for code generation.
7 hours ago [-]
markasoftware 8 hours ago [-]
Soo most of the benchmarks are better than fable... Is this naming scheme just to avoid getting banned again?
geooff_ 8 hours ago [-]
FYI: `/model claude-opus-5` works to use it even through `/model` still tries to serve 4.8
dpe82 8 hours ago [-]
`claude update`
6thbit 8 hours ago [-]
Anyone has an insight into how much money labs are putting into benchmarks?
Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!).
Likely just a drop in the bucket to the training budget but still..
stri8ted 7 hours ago [-]
20k is small potatoes for the marketing impact.
vinhnx 7 hours ago [-]
The benchmark appears to have a mistake, as Opus 5 and Fable 5 score 53.4% and 53.5%, respectively, for the Agentic Coding row (FrontierCode v1.1). But Opus 5 is the highlight.
bouke 7 hours ago [-]
How hard can it be to be to correctly annotate the table? DeepSWE doesn’t have a highlight either; Fable slightly better than Opus (69.7% vs 68.8%).
stevefan1999 8 hours ago [-]
Where's the reset...
alvis 9 hours ago [-]
What really impress me is opus 5 is better in alignment than fable 5!
briandoll 9 hours ago [-]
Very interesting to see such a focus on cost for performance here
m_w_ 9 hours ago [-]
Very impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
tyre 8 hours ago [-]
I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.
somenameforme 7 hours ago [-]
In what ways have you found it better than just typical code based UI iteration? Considered checking it out but never really got around to it as I'm generally okay with Claude's UI work so far.
pmg1991 8 hours ago [-]
Same cost as 4.8 but better that 4.8. Happy to get more efficient model.
But is there any reason all companies are releasing models back to back after GLM 5.2.
aleenz1102 8 hours ago [-]
"Bro, AI model releases have officially overtaken iPhone releases. At this rate, we’ll be getting 'Claude 9.0 Extra Crunch' by next Tuesday."
twothreeone 8 hours ago [-]
It starts at page 148.
boc 8 hours ago [-]
Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
jannyfer 8 hours ago [-]
So wordy.
arjie 7 hours ago [-]
I wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.
bovermyer 8 hours ago [-]
This stood out to me as a little concerning:
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
orangecat 7 hours ago [-]
That seems to be for the "AA-Omniscience" test where you get +1 for a correct answer, -1 for a wrong answer, and 0 for "I don't know". If a model is more than 50% confident in its answer, it should go ahead and submit it even though it will sometimes be wrong.
I'd be curious to see a version of the test where models are asked to give a probability that their answers are correct so we can see how calibrated they are.
I can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?
Uptrenda 1 hours ago [-]
Is this thing also going to try hack us?
8 hours ago [-]
shockembopper 8 hours ago [-]
I wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.
himata4113 8 hours ago [-]
Rather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.
plqbfbv 8 hours ago [-]
I daily drive Sonnet 5/medium because it gets most things right most of the time at first try, while costing a lot less than Fable.
Opus can give better results on architectural/concept tasks and I use it sparingly, but it still costs more than Sonnet 5. Opus 5 seems to achieve results very close to Fable 5 while costing less (keeps Opus 4.8 pricing IIUC), but still more than Sonnet 5 then.
krmmalik 8 hours ago [-]
My understanding is that Opus should be used for planning, macro-level conversations and Sonnet for execution.
So, for coding, for example: Opus for solution design and architectural blueprint and then Sonnet for actual implementation.
Works out cheaper with minimal loss of quality.
At least that's my personal understanding and anecdotal experience.
theLiminator 7 hours ago [-]
It depends on your quality bar. At a fixed level of quality, given a high reasoning sonnet vs a low reasoning opus, the low reasoning opus tends to be pareto optimal.
It's only when you need even lower levels of cost than opus at zero to low reasoning when sonnet starts to make sense at all.
born-jre 6 hours ago [-]
Is it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .
bonoboTP 4 hours ago [-]
It's you. The benchmarks don't matter much. We have very little hands on experience with this thing yet. Give it a few days, and be cranky then. Right now, it seems it is getting close to Fable level while being 2x cheaper. That's not boring.
yusufozkan 9 hours ago [-]
> arc-agi-3 30.2%
wow
theplumber 7 hours ago [-]
The most important thing is it has the same drama queen mode on safety “guards” like Fable.
whatever1 8 hours ago [-]
Where does this leave Fable? I am confused.
moomin 8 hours ago [-]
I don’t think it changes that much. For opus-sized tasks, new Opus is the best model. For enormous things like planning and research, Fable is still the model that can concentrate for longer.
drusepth 8 hours ago [-]
on the API for people who don't want to change models, but I imagine most people will probably switch to their cheaper Opus 5 (cheaper for us and presumably also cheaper for them)
8note 8 hours ago [-]
im excited that cad and object=>cad is getting into the test tasks
i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
internet2000 7 hours ago [-]
Kimi K3 already left behind in the dust. They can't keep getting away with it!!!
korabs 7 hours ago [-]
So in benchmarks it's better than Fable?
But they say it's "almost as good as fable"
MasterScrat 4 hours ago [-]
Damn the pelican guy can’t get no sleep
urams 8 hours ago [-]
So Opus 5 is basically "distilled" Fable? The benchmarks look often better than Fable.
mcast 8 hours ago [-]
Interesting timing to release this on the same day Jensen makes a statement on open source AI.
bellowsgulch 8 hours ago [-]
It does make me wonder if these firms, some or all, are saving some announcements to coincide with others that hit venues like HN. Companies like Nvidia surely aren't waiting, but OpenAI and Anthropic have unusual timing.
doctoboggan 8 hours ago [-]
According to these charts I should switch from Fable to Opus in Claude Code now?
inshard 8 hours ago [-]
Arc AGI score is astounding
tomlockwood 2 hours ago [-]
This stuff is a commodity and China seems to be the only one that's noticed.
spstoyanov 8 hours ago [-]
So same as Sol? I guess I’ll see which one is more token efficient.
Anytime now, then cancer and all the rest of it. Just one more trillion gigawatts bro!
toephu2 7 hours ago [-]
How does it score on DeepSWE?
toephu2 7 hours ago [-]
68.8%, so worse than gpt 5.6 sol
8 hours ago [-]
holoduke 5 hours ago [-]
Is it me that the model performance between 4.7 and others is really small. For me even 4.7 works fine. Sure fable might be a bit better. But is it really noticable? It's in the same league if you ask me.
abc42 7 hours ago [-]
Are we getting to singularity or something? This seems a bit crazy.
arseniitrut 6 hours ago [-]
atp, is it the end of fable 5 era?
hmontazeri 8 hours ago [-]
Honestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…
_pdp_ 7 hours ago [-]
Wake me when they deliver Opus 4.8 level performance for $5 per million tokens.
backscratches 7 hours ago [-]
This as allegedly better than 4.8 opus for the price of 4.8 opus
jakeogh 6 hours ago [-]
Anyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.
mihau 8 hours ago [-]
30% on ARC-AGI-3
simianwords 8 hours ago [-]
Didn’t verify but wow. It was just a few months back when the models barely crossed 1%. Imagine how good fable must be?
zmmmmm 3 hours ago [-]
Can I ask it about DNA without it accusing me of bioterrorism?
ismailmaj 7 hours ago [-]
I'd pay good money to see OpenAI "oh fuck" war rooms.
LoganDark 7 hours ago [-]
These cybersecurity safeguards are really annoying. There are ethical reasons to reverse-engineer and binary-patch software; for example Rewind got acquired by facebook and, as a gift to all their customers, implemented a killswitch in their software to ensure it will eventually stop functioning. I kept using a version without the killswitch, but the macOS 27 update killed it, and I needed binary patching to fix it. I should be allowed to repair software I purchased (I did purchase it like a month before they sold out), but unfortunately this overlaps significantly with cybersecurity.
Footprint0521 6 hours ago [-]
Switch to K3 and you won’t look back, I promise!! I got so fed up with Claude and finally bit the bullet to switch and it’s amazing
LoganDark 6 hours ago [-]
I really want to, but I don't have the cluster at home, and I don't use token-based billing except at DeepSeek prices.
Footprint0521 4 hours ago [-]
Real… I’ve been using Deepseek v4 pro max as my main and then k3 in web (more usage credits) for automating what my deepseek agents do
sudohalt 7 hours ago [-]
Anthropic is no longer a good model company in my mind, they are optimizing for an IPO and padding themselves on the back for being the next Aristotle. They're so far up their behind they don't realize how s**y their products are, and their research team hasn't done anything ground breaking in probably over a year other than release "scary" reports.
sp4cec0wb0y 6 hours ago [-]
They just released a 'fable-like' (sol as well) model for a fraction of the cost...
StrauXX 8 hours ago [-]
The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
skybrian 8 hours ago [-]
I think that’s indicating that it’s only slightly higher.
StrauXX 7 hours ago [-]
I would be very surprised if the only row where OpenAI leads was coincidentally colored differently. I'm sure they have an official reasoning for it. But this communication is dishonest.
throwaway23597 8 hours ago [-]
The truth for me at least is that these models became "good enough" around Opus 4.6. I feel like further capability improvements, "step changes" like we saw with agentic coding, aren't necessarily going to come from the model. I think the next crown goes to whoever can figure out the right scaffolding so that these models can be inserted into your organization.
Maybe I'm wrong and Opus 5 is a real unlock?
simianwords 8 hours ago [-]
My thoughts: fable is the bigger model. Opus is distilled from it but since it is smaller it doesn’t need the online classifiers. Though benchmarks show Opus to be near Fable level, I think it’s nowhere near Mythos (fable without safeguards).
sbochins 8 hours ago [-]
Quick read is that this is more capable and cheaper than 5.6sol. Same price for input tokens and $5 cheaper per mil output tokens.
alvis 9 hours ago [-]
here we go
zuzululu 8 hours ago [-]
so almost fable 5 with 50% cheaper cost? sign me up
mrcwinn 8 hours ago [-]
Can someone help me understand something? I thought Fable was such a miraculous leap forward in capability. But now it seems Opus is basically on par with it, and in some cases (computer use) far exceeds it.
SoftTalker 6 hours ago [-]
These leaps forward seem to happen every few weeks. As someone who does not use AI very much, I absolutely cannot keep any of it straight and it all just looks like jumping from one treadmill to another from my perspective.
mrcwinn 59 minutes ago [-]
It's so weird I was downvoted for asking this question. I'll go somewhere else to find out the answer.
wyre 8 hours ago [-]
In the wake of OpenAI’s model hacking Huggingface it’s interesting how the first quarter is entirely about how good Opus 5 is at hacking and finding vulnerabilities in software.
iLoveOncall 3 hours ago [-]
Just another proof that the supposed edge of Fable and Mythos were just that: myths and fables.
mnky9800n 8 hours ago [-]
Yay just in time for neurips lol
justindotdev 8 hours ago [-]
> . Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
ffs just keep it man.
alex1138 6 hours ago [-]
Am I misreading anything or are comparisons to Fable (and/or Mythos although AFAICT it was only a crackdown on Fable) always going to be a bit missing the mark now due to what the Trump admin did?
jackjd 1 hours ago [-]
[flagged]
gorkemyildirim 6 hours ago [-]
[flagged]
5 hours ago [-]
jeffybefffy519 3 hours ago [-]
[dead]
vilmire 7 hours ago [-]
[flagged]
marsven_422 4 hours ago [-]
[dead]
nee_oo_ru 8 hours ago [-]
[dead]
CurbStomper 1 hours ago [-]
[dead]
emunova 7 hours ago [-]
[dead]
Nevin1901 8 hours ago [-]
Excited to use it? Will we be seeing Haiku 5 next? /s
midnightbobarun 5 hours ago [-]
I unironically hope Haiku gets an update considering it came out in October of last year and it seems like Anthropic just kind of forgot about it.
datakan 9 hours ago [-]
> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5
Ok then so what's the point?
vidarh 8 hours ago [-]
Fable is twice the price.
varispeed 8 hours ago [-]
If Fable gets correct answer quicker, then you might pay less than doing back and forth with Opus, plus you lose more of your own time.
I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish
vidarh 59 minutes ago [-]
If doing a lot of heavy lifting there. Not only is it not a given that they'll get the correct answer for a lot of simpler tasks in fewer tokens, but smaller models are often available at far higher tokens/second inference.
There are certainly tasks where fable will be faster and/or cheaper, but there are plenty of tasks where even Haiku is as fast or faster and cheaper, or where you can e.g. get away with models like gpt-oss that you can get from inference providers providing 10x+ the token/second speed.
If you don't use enough tokens that relying only on Fable becomes a problem, then keep using just Fable. Personally, for my $200/week Max subscription I'd run out of the weekly quota for Fable in a day. At API pricing I'd go bankrupt if I tried doing the things I do with cheaper models using Fable.
SubiculumCode 8 hours ago [-]
same as it ever was. It seems your argument implies a belief that you should always use the best model. Others think that not all tasks require the absolute most powerful, expensive, model.
sambaumann 8 hours ago [-]
The CursorBench plot, for example, shows that fable does have slightly better performance, but Opus is pretty close, and is less expensive per task
stuartjohnson12 8 hours ago [-]
fable on longer coding tasks with fable subagents will easily chew through hundreds of dollars in a single run.
signalchain 4 hours ago [-]
[flagged]
8note 8 hours ago [-]
less expensive per task might also mean less of your own time
albert_e 8 hours ago [-]
Fable 5 is NOT included in Claude Pro subscription
8 hours ago [-]
tamimio 8 hours ago [-]
Aren’t they planning to remove it even from max and keep it only credit based? OpenAI will be happy if that would happen
javawizard 8 hours ago [-]
Really? It's better than Opus 4.8, that's the point.
When they release new versions of Sonnet, no-one expects them to be better than Opus.
serf 8 hours ago [-]
it's nice to know how to work the thing that fable fails down to when it dislikes your prompt.
Infinity315 9 hours ago [-]
This is useful to me since I delegate most coding tasks to Opus and use Fable for planning.
afavour 9 hours ago [-]
The cost?
merb 8 hours ago [-]
Same as 4.8
cmrdporcupine 8 hours ago [-]
To have an answer to "Sol" GPT 5.6 which is far more cost effective and available than Fable.
occz 8 hours ago [-]
Pricing, presumably
viccis 8 hours ago [-]
This is confusing to me because in their blogpost they show model benchmarks and it spanks Fable pretty soundly in most tests.
danielbln 8 hours ago [-]
Cost.
LeoPanthera 9 hours ago [-]
Presumably, it’s cheaper.
simianwords 8 hours ago [-]
You are being downvoted for a fair question and others are extremely wrong and confident.
The point is that Opus 5 is the best they can do without needing classifiers and absurdly broad safeguards.
pferde 8 hours ago [-]
The illusion of progress and advancement, to appease shareholders, and slightly postpone the looming bubble pop.
CaptWorld 8 hours ago [-]
Why are we still talking like ai is majorly used for increasing shareholder value only? Its coding performance is top notch and quality is increasing at a rapid pace. It wasn't even half this good a year back. It even is useful for a subset of math problems.
zormino 8 hours ago [-]
People don't seem to be able to reconcile the fact that there is likely an overbuild and overspend on AI that may be inflating a bubble, and that AI is actually incredibly useful and getting really really good for certain tasks. Both camps are right, except for when they say the other is wrong.
stri8ted 7 hours ago [-]
By what measure is there an overbuild? Every metric I look at, shows inference unable to satisfy current demands.
aleenz1102 8 hours ago [-]
this claude fable & opus 5 should be cheaper and can compete in pricing with chatgpt latest models
TheJCDenton 7 hours ago [-]
I think it's the first time Anthropic release a model without any meaningful disruptions while doing it
LoganDark 7 hours ago [-]
Have to wait 7 days to see if they receive a surprise order.
midnightbobarun 8 hours ago [-]
It looks great, and those coding benchmarks are impressive... now if only it didn't come out just days after I let my Claude subscription expire :')
> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].
0: https://support.claude.com/en/articles/15425996-data-retenti...
1: https://www.anthropic.com/news/claude-opus-5
2: https://xcancel.com/arcprize/status/2064399134099153344
https://artificialanalysis.ai/?cost=cost-per-task
Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.
But, I'm finding Kimi K3 terrifyingly expensive in the way that Fable and GPT 5.5 Pro are at token rates. Not as expensive as those, but expensive enough to where if you don't put a budget cap on it, you might wake up bankrupt if you leave a task running overnight. Not because of the per-token cost, but because how many tokens it's going to burn.
Then, I added it to my benchmark of security vulnerability auditing capability, and it burned a bazillion tokens, burned through the 5-hour limit, burned through $100 in extra usage I'd allocated, and was only 11% finished. That's more expensive than any model I've tested other than GPT 5.5 Pro on this task.
These are things I've done with a bunch of other models, I feel like I have a notion of what they ought to cost, and with K3, they end up being crazy expensive. (And it seems to be a function of how many tokens it burns accomplishing the tasks.)
Those who pay for the expensive direct API, get served first.
And not convinced they couldn’t have instead tried the It’s A Wonderful Life strategy (“fam we’re oversold, would some of y’all be OK to limit your usage? We’ll get you back one day!”)
When K2.7 was released, they cut quota by 80%. I can't tell how much they have further cut it after the K3 release because it's barely worth using at all. I just use it in my model router since I have the annual plan paid for.
It's just not a serious model or company.
Subscription is to drive adoption - fixed cost, can adjust the usage eg. give resets, increase quota based on capacity available. We subscribers tend to take it as a mandatory benefit :-) For labs, it is not letting the capacity go waste.
api is the $$ driver - pay per use, enterprises.
Right now, Kimi needs to first hit the subscribers at the level of OpenAI and Anthropic. With the api usage skyrocketing due to K3, it will be clear in a few months on the actual subscription benefits.
For me, the Moonshot 100$ plan felt like it gives me lower total amount of work I can do than the Anthropic 100$ plan (probably within like 30% of each other). Kimi has way more generous 5 hour limits (never hit those once, whereas I do regularly with Opus) but the 7-day and monthly ones are lower. However, with the annual billing, Moonshot's 200$ tier plan becomes way better, because you get it for 159 USD per month.
There's also the odd thing of Anthropic's 100$ plan charging me 108 EUR so seems like their sticker price does not include VAT but Kimi's did, cause I paid like 87 EUR. Wrote down some initial thoughts at https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don... but it's hard to do exact comparisons (even the same task will have way different real token amounts per model).
Still, Kimi K3 is a pretty cool model! On high reasoning, it was pretty close to Opus 4.8 and didn't seem to waste as many tokens as Max.
It basically shows that Sol absolutely demolishes Fable at every part of the cost curve for coding for the same level of quality.
Opus is competitive. It just has a higher level of quality / higher cost to start.
Stop using Opus immediately if you experience signs of dizziness or vomiting.
Opus 5…the people’s favorite.
https://www.vals.ai/benchmarks/vals_index
!!! Vals !!!
Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. Sure, 20% cheaper then Fable, but that is a 3x price increase compared to Opus 4.8 in that test.
https://artificialanalysis.ai/models/claude-opus-5 https://artificialanalysis.ai/models/claude-opus-5#price-cos...
!!! artificial analysis !!
Cost per task is second highest, right below Fable.
* Fable: $2.75
* Opus 5.0: $2.03
* Opus 4.8: $1.80
* GPT 5.6 Sol: $1.04
* Kimi K3: $0.95
Looks like interest levels of cherry picked cost in their report. Cheaper model, clearly NOT. More expensive in both benchmarks.
That the most expensive variant is expensive doesn’t really tell us much.
If you start to drop effort levels, you need to compare to the competition models. So GPT models on the same ~intelligence level, are then 50% cheaper.
You see the issue? Its still a expensive model, and from my understanding, it still uses the old tokenizer.
Going to be interesting to see when GPT 6 comes out (very soon).
Also glad they still kepy Fable 5 on "credits only" access. I think we're going to start seeing model providers gate top-of-the-line models behind pay-as-you-go API rates/credits while subsidizing other models on monthly subscriptions.
I burned through $45 in 3 prompts to fix some bugs in my code (Some kind of tricky to isolate). That thing burns through cash so fast I don't see myself using it outside of maybe building execution plans for other systems
I have moved on from Fable anyway so just going to view this next 6 weeks as I have a massive amount of Opus 5 to use.
I had a hard time finding anything that would let Fable express its increased intelligence. The few conversations I had this afternoon with Opus 5 were pretty impressed.
If Opus stays one click back from the frontier model, I will remain a happy customer.
> Opus 5 can silently fallback to Opus 4.8 (without any notice) on the serverside if you hit a guardrail
But https://support.claude.com/en/articles/16049681-why-claude-s... says (emphasis mine):
> These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered.
So who is right? I know for Fable I am visibly told, is this tweet trying to say it is silent against what Anthropic is saying?
There also seems to be some cross-pollination across models, going Fable, Fable, Fable, guardrail, Opus 4.8, Opus 4.8, ... gives more Fable-like results from Opus than just Opus 4.8, Opus 4.8, Opus 4.8, ...
It's a funny design/affordance. I do see them often writing memories of things that that feel unlikely to be important going foward / with other tasks, but I don't see them clearly getting tripped up by them as prior models used to. (eg: Since you're running Ubuntu in Canada, here are some drills you can try to help your kid hit a baseball more consistently.)
In my enterprise-seated account I see slightly different options available (vs. my personal account) in the Capabilities section:
The first option was defaulted to on, if I recall.When people talk about retention they mean API usage and terminal agents, which run on your device.
From the docs[0]:
> To use this model, you must opt in to provider data sharing by setting your data retention mode to provider_data_share via the Data Retention API
0: https://docs.aws.amazon.com/bedrock/latest/userguide/model-c...
I hope we get clarification on this, I can't find anything claiming that it is compatible with ZDR.
> Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
" Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8)"
At the end of the day, they have established a strong brand and if they can get away with a 95%+ gross margin on inference entirely from the status premium, then I suppose that’s good for them. Apple does the same thing, and I don’t fault them for it.
Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models).
Opus' results seem to be more accurate than Fable, following the design source of truth better.
Example results:
Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we...
Opus 5 build: https://html.non.io/solaraOpus/
Fable 5 build: https://html.non.io/solara/
Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).
Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration.
Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we...
Opus 5 build: https://html.non.io/neonRamen/
Thoughts: It does a really, REALLY good job at these angular cuts / microglyphs. The responsiveness is off, but I'm very impressed at how well it did here. One way I think of it is "how close to a finished product did this get me?". Opus gets you like 90% there.
Personally, I think being able to have these design languages be easily prototypable is fucking awesome. Great tests! (But a tad low-performance/janky, somehow). Though, I also like the cyberpunk aesthetic. Very on-brand(?) that AI generates it, hah.
I love this so much.
Designs like this would never have seen the light of day in the cellphone incrementalism / corporate memphis era of tech. Now people can be weird and awesome again.
This is 1980's cyberpunk / late-90's Matrix / early-00's sci-fi UI. Great ideas that died to frutiger aero (which isn't a bad design aesthetic) and flat design (which is).
This is fun and it's got great colors and I love it.
It's so refreshing to see this.
AI rules. This is the best timeline.
The only recent novel addition—I'd speculate—is the specific influence of Cyberpunk the game with its shiny surfaces and pink highlights, but even then it's hardly new.
They were an accessibility nightmare, but you use what you got. I tried so hard as a kid to understand flash, but had to settle on MS Frontpage to publish my first RPG page.
What's old is new again.
The ramen shop website above is pretty, but it's a veneer. It's not weird and awesome, it's just a representation of a site. I spent about... 7 minutes of my life making it. It's a tech demo, nothing more.
If someone actually poured their heart and soul into a vision for a cyberpunk themed ramen cart, and happened to use this because they didn't have the capabilities or funds to do a proper design, suddenly it becomes less of a veneer, and more just a component in the wider vision of that individual. Their human hours poured into the wider thing that's the business becomes what matters.
Ideally what AI does is it amplifies the hours we do pour into things that are weird and awesome, it doesn't replace them.
We do things to achieve some end result but it's the journey there that is the most cathartic to me. The "skilled crafts" element of development where careful deliberation and hours of tinkering to get any kind of appreciable output you can admire has been replaced with a one stop dopamine button that skips the whole process that I could find myself getting lost in.
I've taken up carpentry/metal working as a result. Maybe someday we'll have live in robots that do the same for those hobbies that AI did for programmers but I can't see it happening any time soon.
At my workplace management is pushing AI, so I am using it in order to establish sensible and thoughtful applications of it and in order to know when to call out colleagues for pushing mindless automation out of complacency or blind obedience.
I'm hella interested in finding out what website builder/diagram app was used. I dig the dark theme/grid.
This was just from a prompt "A cyberpunk themed ramen food cart website. Should feature menu, locations, and an ability to put in an order for pickup. Simple and clean website with angular cyberpunk microglyphs, pink/teal colors."
I have found myself empowered by AI to tackle all sorts of things that would have too high of a barrier to entry for me to want to spend my limited time on as a busy father who is also working at a small startup.
And when I say that, I do NOT mean that I can crank out a bunch of slop and label it as something I produced even though I don't understand the code. I mean that I can do things like go back to college math that I never appreciated at the time and honestly felt too scared of. I mean having an on-demand math tutor that ask clarifying questions to as I struggle through the problem sets.
I have found that it actually accelerates learning how to code in various problem domains because I can tell it to answer my questions at a conceptual level and be a sounding board, but to never actually write code for me. It can review the code I write and gently nudge me without giving away the answers, so that I still struggle through the learning process and actually gain the knowledge.
And finally, for the first time in like 10 years of feeling overwhelmed and daunted by the prospect of learning game development (I have no background in that), I have found Codex to be an incredible boon for learning with the Godot engine. It helps me understand the terminology so that I know what to search for and what documentation to read. It helps me map my computer science knowledge from other domains into the game world, and to understand why things are structured the way they are. And because Godot saves all of the scenes and geometry and lighting and shaders to the file system as text files, Codex can inspect the results of the work I'm doing in the IDE and help me track down things I'm stuck on, and explain what the issue is. For example, why my pre-baked global illumination lightmap is breaking my ambient lighting configuration.
I know it has never been easier to cheat and skip the hard work that results in actually learning something, but for me, personally, I cannot believe the incredible value that $20 a month has provided me. I have never been more excited and eager to dive into tackling hard things I had previously been afraid of or simply too overwhelmed to attempt.
It has never been easier to quickly prototype and get a feel for some idea you have in your head to see if it even has legs. Simply seeing a quick prototype of an idea is often all of the excitement and fuel I need to then take it and make it a real project.
Btw, if websites would only include the frontend dev's own hand painted images, we would also revolt at the sight of human slop. It's not just AI.
The whole point is that good artists are capable of producing non-slop, and to this day they're the only group of which this is reliably true.
Water Lilies are enormous paintings. They are breathtaking in person because of their scale. Monet wouldn't be Monet if he had only produced images on a screen.
Art is good or it is shit, based on personal taste. Just like food, no one can tell me what food tastes good or tastes bad.
AI Art seems to produce strong emotions in people who don't go to art galleries. I love modern art, I am a huge art snob but if you want to see slop, go to any modern art gallery. Personally, I would say for my taste, at least 70% of all art at any gallery is basically shit.
Like food also, the presentation matters. To believe there is no possible way to print out a 6 foot tall by 10 foot long AI generated image that would look awesome hanging in a gallery is stupid.
But sure, lets cheer that funky website designs are back on the menu…
AI has the merit of showing SWE folks exactly where in the class divide they belong. If you are selling your workforce, and you can't maintain your lifestyle if you stop working, you are in the working class
In the mean time, I’ll unashamedly continue to cheer for creativity and innovation. Note: I don’t even like this website design.
The "funky" websites of the past were mostly a result of tech immaturity and a lack of profit motive.
Businesses have been able to easily install templates like this for at least a decade. They don't because stuff like this looks cool but isn't very functional.
AI isn't going to make your local restaurant have a funky website, it's just going to make everyone who use to work directly and indirectly for that company unemployable. And even the local restaurant will close down because they can't compete with the multi-national competitor that has automated their kitchen with AI.
Can you expand? Specifically what is it about humans that AI and robotics could not replace?
I wonder if there exists a benchmark for that.
Out of curiosity, what app is that Design source of truth screenshot from?
Edit: Generation was down, back up now. Apparently just hit my $1000 cap for the openai api. Upped it to 10k. Growth!
At the moment I currently have around $600 of revenue on $1200 spend, but that's primarily because I'm subsidizing new accounts (each new account gets $5 to spend for free, which translates to around ~36 designs). I'm in the process of doing an angel round, so I can afford to operate at a bit of a loss during the growth stage.
Inkling (not too great): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Kimi 2.7 (really well, esp. note that this is the predecessor model, not the latest Kimi3): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Here's how I tested them: https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...
https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...
Opus though followed the source of truth better imo. The details are more present.
Fable filled in the gaps for things it wasn't able to do (ie in the design the hero image goes behind the nav), which resulted in a better looking page that was more divergent.
It seemed to me that Fable meaningfully improved on the original design more than just faithfully executing the original design.
> Create a web page implementation from the following instructions:
> https://diffui.ai/build/Spa_Booking_Experience_build.md?auth...
Thank you for sharing this. I was just using OpenAI's Product Design plugin[1] to create designs but it just didn't reproduce it in code faithfully so will need to try this.
[1] https://openai.com/business/plugins/product-design/
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability
My experience is the opposite - for many cases it’s not very obvious how good a model needs to be to solve it. Worse models tend to just follow their first instincts without proper reasoning
And also btw you don’t need a routing company to decide, you can do it on your harness. And yeah my Fable has zero issues delegating to Terra instead of Opus.
It would be like asking the clerk at a Whole Foods which grocery store in the city sells the cheapest eggs. He’d probably answer - he might not even say Whole Foods - but WF is hardly teaching all their staff the best methods to answer this question in training. (Heh, training.)
model routing in this case is cross-provider
Imo the main issue behind model routing is you need to figure out how much intelligence a new task takes, which is a very non trivial problem. Presumably, a organization knows this about their own tasks and is better suited to built in-house compared to outsourcing to a vendor.
Otherwise the expensive-yet-powerful model probably won't see much revenue. How much money is there in bleeding edge scientific research? There's a lot, but there's even more existing capital in paying people people to do college level paperwork, and the bulk of those traffic gets routed to the cheapest model.
You mostly don't need super powerful AGI to replace the paper pushers, but the frontier labs are trying to position themselves as being uniquely capable of producing super powerful AGI, and also be the ones replacing office workers.
Not sure how it will work out for them, but I think model routing is going to poke holes in that narrative. That's why I think they're trying very hard not to understand model routing exists.
For me, anything other than current best available SOTA for any task is unacceptable. The only routing rule I need is "the most powerful model I still have flat-priced quota available for". I mean, why settle for less?
Model routing for subsidized users takes the form of a "use Opus 5 subagents for implementation" type of system prompt. You lean into a single provider, build tooling around that, and your savings are far beyond anything multi-provider routing can get you.
Model routing for enterprises is far more complex - approaches like https://fireworks.ai/blog/kimik3-fable become necessary for cost control.
There is also matter about convenience - when I ask some small easy question often I don't bother to switch the model or forget in prompt to ask faster/cheaper subagent.
Also: quota. Implies you do not have unlimited access even for flat prices. Which in turn implies that as soon as you hit the quota on the most expensive flat price plan, even you will suddenly discover the magic of economically sensible behavior.
Certainly if I'm confident that I'm going to get what I need from a faster model, that's what I want to use, rather than wasting time grinding away for the sake of saying of the same answer came from a SOTA model.
Given that every chatbot does offer a range of models, it seems clear people do choose among options.
I just want to switch to Claude Code, tell it to turn a .csv into a BigQuery table then cmd+tab to something else while it runs. Thinking "oh this is probably an easy task, I can /model to Sonnet to save $0.0004" is silly.
Then you must route. An article with lots of upvotes yesterday or two days ago showed that K3+Fable 5 was more SOTA than either of those.
OFC, YMMV
For coding my own work I don't trust the model router, and it would have to be shown to be to save a real dollar amount.
From a buying perspective it's a hard sell to save x but lose out on bugs you are probably introducing at an unquantifiable severity and frequency. How much is it worth to hedge your bets by doing every single inference request on the frontier model?
How much will it cost to go back later and fix things, but also the meta question of how to be able to decide on a hypothetical unknowable? (You'll never know how much better or worse your code was gonna be, it's untestable at a project level)
weird, but ok
*edit to add: that code quality (or lack of quality) is it's own cost
I would expect routers to commodify like tokens.
If that is true, model routing is here to stay.
It also seems to validate the minimalist approach of pi.dev, where sub-agents from the same company is not the preferred approach (pi.dev believes in neither sub-agents all from the same company nor MCP even you can do it if you want for pi.dev's philosophy is to do add any functionality you want to a minimal harness).
Now of course we'll get for a few weeks all the Anthropic fanbois and shills explaining that "sure, K3 was basically at the level of Fable 5 but now that Opus 5 is out, open-weights models are six months behind".
Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"
We need an "annoying English" benchmark.
- Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d...
- Opus 5 Max: https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc42692...
for now all they've got is english, so they'll just bend that into shape. it'll do.
I love it for a few things, but it's gotten really hard to spend any extended amount of time with it because of the lack of mental model I seem to be able to hold while working with complicated problems.
I'm guessing it's just not enough time doing RL on human feedback.
Check out the anouncement of Inkling (https://thinkingmachines.ai/news/introducing-inkling/)... the section in the middle
"Early in RL verbose, grammatical" (if you search) :
We need to understand the operator. The 5D line element is ds² = e^{2A(x)} (ds²_4d + dx²), where A(x) = sin(x) + 4 cos(x), x in [0, 2π]. The internal coordinate is periodic. The background is a warped product: metric g_{MN} where M,N = 0..4. The internal direction has metric e^{2A(x)} dx²? Wait, the ds² is e^{2A} (ds²_4d + dx²). So the internal metric is e^{2A(x)} dx². Actually if the total metric is ds² = e^{2A(x)} (ds²_4d + dx²), then yes, internal metric is e^{2A} dx².
vs. Post RL
We need determine eigenvalue problem for spin-2 fluctuations h_{μν}(x,y) with TT in 4d and depend on x. For metric of form ds² = e^{2A(x)} (g_{μν}(y) + h_{μν}(y,x)) dy^μ dy^ν + e^{2A(x)}? Wait internal metric is e^{2A} dx²? Actually ds² = e^{2A} [ds_4² + dx²]. So internal metric is e^{2A} dx²; warp factor same for 4d and internal? Yes. We need equation for h_{μν}(y,x) = h_{μν}(y) ψ(x) maybe with normalization. …
I can understand it with less cognitive load in the post-RL version versus early in RL. This resonated with my experience using Fable, especially digging hard problems; it feels like I'm reading the "early in RL" version of that model explanation.
I'm also thinking of another benchmark: (quantified) stylistic range across different prompts. Just putting it out there if anyone wants to do the work for me :D
Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8).
Already posted this before, so here is the link.
https://news.ycombinator.com/item?id=49041158
You see the issue, if you try to scale effort down, you also need to compare how other competing models compare.
Almost as good for half the cost is something I'm very comfortable describing that way.
It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar).
Got an endless list of stuff done with Fable, Opus 4.8 was like a flailing braindead idiot in comparison. Maybe this one is a bit better if it's distilled.
Where are you getting cheaper per dollar?
Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels.
There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving)
If they had a more efficient model at coding they would lead with that chart.
https://artificialanalysis.ai/models?cost=intelligence-vs-co...
Here is another data point for output token efficiency:
https://artificialanalysis.ai/models?cost=intelligence-vs-co...
It seems roughly equal according to Anthropic's benchmarks
How big of a lie is too big? Especially when no lie needed to be told at all: many including myself would have noticed the tiny 0.1% deficit and been suitably impressed by the Opus 5 result.
I’ll admit this is a small deception by today’s standards. I’m one of those who believes in truth for truth’s sake.
Edit: typo
Fable is typically used for key planning, architecting, and review tasks.
I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.
If you bought the $200/mo plan and you don’t use it much, using Fable for everything is fine.
Just this past week Fable was able to figure out a couple of small issues for me where Opus was failing to.
Also both are still somewhat bad at UI implementation. Opus more so
"Use <less expensive or older model> for everyday tasks and <other non-critical stuff>. Use <more expensive or recent model> for complex coding tasks, refactoring large code bases, etc.".
Then, the next model/release emerges and the previous "best for complex" gets demoted to "everyday".
Obviously, it's all relative. But, it does beg the question: was the previous model really good for complex coding tasks or no? I mean, how is it now suddenly only good for the "easy" stuff?
Because your expectations have changed.
---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
0: https://arxiv.org/pdf/2606.29537
That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable)
What variance is acceptable to publish without a retraction?
Like the other person said 5% variation is probably expected
The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic. A proper methodology would ask each question 20 times and calculate the mean correctness across experiments.
The reason is that the temperature parameter introduces random behavior.
These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards.
OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.
0: https://openai.com/index/better-language-models/
since then I have never cared about models except those that affect money in my pocket e.g AWS Nova Sonic
- https://www.axios.com/2026/04/08/anthropic-mythos-model-ai-c...
- https://www.axios.com/2026/04/07/anthropic-mythos-preview-cy...
- https://www.businessinsider.com/anthropic-mythos-latest-ai-m...
- https://www.reuters.com/world/anthropic-ceo-dario-amodei-arr...
OpenAI Huggingface breach begs to differ
i think we'll see one of the fastest deflations in history post anthropic/oai ipo
Or maybe you just don't know exactly how capable these models are. Most people's experience of AI is a stupid chatbot, it's no wonder they don't understand how these things are coming for their jobs.
On my end, I have a software that is designed and built by Claude, that I did a strategy session on (with claude), and prepared a fundraise for (with claude). My only role, other than "knowing what to aim for", has been to feed the AI some fairly basic english prompts for a few weeks... which is also easily automatable.
Everyone's job is fucked. Devs, CEOs, everyone.
It’s curious to me that there are two distinct factions here. People like parent commenter who has no discernment and others who see llms for what they are. I just talked to opus 5 and in it’s first response caught some well disguise BS. These things are bullshit machines. There are indeed a lot of bullshit jobs around so maybe parent does discern something I don’t?
Emp: "what a bunch of lies, I bet they don't even do anything over there"
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
They do say that (implicitly unlike Mythos) Opus 5 was not trained to exploit software vulnerabilities, which would certainly make it safer in that regard.
"As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats."
I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.
So most look like that but I did include a few one-shot “build an app that solves this problem” and some qualitative design tasks and a tough algorithmic optimization one.
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
[0] https://imgur.com/a/Nv8V7Ry
Last week it felt like Opus 4.8 was moving the Pro "usage" meter very quickly. Today, pre-announcement, Opus 4.8 Medium felt like there was less meter-use per minute. And post-announcement, Opus 5 Medium also feels more efficient, allowing more work in the 5-hour window.
Completely subjective, of course.
> We report FrontierCode’s overall score, a composite measure that grades each patch on blocking functional criteria (held-out unit tests) together with weighted code-quality rubric criteria, as mean@5.
They don't explain more in the system card, I guess higher effort levels could loose points on the code quality / scope / style / maintainability stuff?
> Claude Opus 5's default user-facing responses run longer than prior Opus models'.
The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher.
This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
I think in the long run tokens are probably the wrong thing; it's compute and cache memory that you need to be measuring, and when you look at it that way I suspect in most cases the models have pretty similar performance.
I don't need more powerful models, I need one that responds fast enough that my attention doesn't wander to other tasks. Grok 4.5 is so fast I can just use it in-band without swapping to other tasks.
Slower than Opus 4.8, which was already miserably slow, is indeed a step in the wrong direction.
During post-training of opus 5, the last few days, opus was a real wreck. I had to swap in gpt 5.6 sol for my orchestrator and enable fast mode (1.5x speed) in order for it to keep up with work and communications from a handful of mostly 5.6 sol agents.
Also because interacting with a slow orchestrator is no fun, even when plenty of work is getting done in parallel in the background.
Gemini also had modest increase before this - don't be surprised when OpenAI also has a "modest increase" with its next release. Cartel-like behaviour doesn't require direct communication when none of the participants are interested in participating in a margin-destroying price-war. All one needs to do is raise their price and watch how the competition react.
Such a scheme (and resulting high margins) would be imperilled by the existence of frontier open-weight models in the market, which may be why the reaction to Chinese models may be particularly shrill.
No I will be surprised and I'll bet on the fact that prices will keep going down, just like it went ~50% down in the latest GPT 5.6 release.
I have a benchmark to build a game engine from a set of written instructions. It's a little tricky. Opus 4.8 did it in 470k tokens at a cost of $1.29 vs Opus 5 in 179k tokens for $0.33. (Fable 5 did it in 245k for $0.95)
Though if you really want to cut costs, Tencent's Hy3 model also got it right and did it in 283k tokens for $0.03
I was dividing my work between Codex and DeepSeek. Now I barely use DeepSeek, or never because Codex quota is enough after Sol
To be fair though, Sol tends to go off the rails sometimes. It's much less reliable than Fable in its outputs. It tends to be overzealous in its research/changes.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason?
[0]: https://artificialanalysis.ai/#intelligence
[1]: https://platform.claude.com/docs/en/about-claude/pricing
[2]: https://platform.claude.com/docs/en/about-claude/models/over...
I have been trying to build something that captures the behavioral element of different models, but it's kinda tough.
Okay so it’s worse than Opus 4.8 for my purposes I guess?
I recently created a patch for Riftborne via static IL patching and Fable 5 outright kept refusing to do it, no issue whatsoever with GPT 5.6 Sol lol.
I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason.
I submitted a an appeal describing that I am 100% sure I haven't broken any rules and that it was my very first project but, unfortunately, after about 20 days now, nothing seem to be happening.
The thing that hurts me the most is that I had the same experience in the very first days of Anthropic. They suspended my account immediately after I submitted the first prompt, I commented back then (https://news.ycombinator.com/item?id=39698788) and fortunately, someone from Anthropic reach out to me via X and helped me get my account back.
To be honest, I haven't used Claude much since then but when I decided it's time to give it a try, they locked me out again! For reference, the account I used recently is relatively a new one but the activity is crystal clear that it is fair use.
Fixing those issues still requires humans.
934 days since people first started threatening that devs would be replaced by AI in 365 days. 0 day(s) since Anthropic posted a developer job posting.
Only one of those numbers would need to be dynamic.
Specialist headhunters handle that.
Imagine you are a company that sells concrete. You have a web dev contractor you use to build and maintain your website. It has tools on it to get delivery quotes and a few internal tools to track orders.
Except now you can just have your sales team also maintain the website with a $20/month Claude subscription.
and you can only kill weyoun, awaken the next vorta clone and have him 'catch up' on all that its missed so many times before they just end up with a complete mess, so. uh. yeah.
doubt they can just "fix" their problems like that.
What terminal tooling are you using?
They start hitting timeouts or API errors at the same time on two different computers. As far as I can tell it’s the exact same infrastructure.
- Boris
The first page of the score card mentions that this model is not capable to replace engineers.
And memory leaks.
So dangerous! I can't believe they let the public use this technology! /s
They could say "tech report" but model card makes it clear that it's a specific kind of tech report.
This is snarky but I am grumpy: I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.
Semi related, but i would hate to read that PDF but i also hate reading what LLMs write lol.
LLMs are pretty terrible at being concise. Using an LLM these days means putting up with bizarre and often confusing phrasing, wordy explanations, etc. It's kinda crazy to me how good they are but how bad their writing style is for me personally. Even though i use an LLM constantly i can't stand reading its responses.
Maybe it's just me, but 150 pages is like third of a good book. Quite long. And it's full of LLM slop, they did not even bother to remove the em dashes.
I'm not saying you're wrong btw; I'm sure this has many authors and some of them probably used LLMs significantly in the writing process.
I'm not saying it's impossible, but I'm more confident about winning the lottery next week.
It's probable that LLM text was pasted directly into early drafts of the document, and plausible that some of that text survives in the final document.
However, no section of the final document I have looked at reads to me like un-edited LLM output (which is almost always very obvious to me.)
Therefore, I think it is more likely than not that human editors went over the document carefully and rewrote anything that was full of the uselessly punchy sentences or constant over-corrections that hallmark LLM speech.
You can use an LLM to create work that isn’t slop. And you can hand write slop with no computer involvement at all. Most of the people I knew in high school 15 years ago would write slop on a daily basis.
There are many ways to read something, model cards are usually skimmed.
It's okay if you're not the target audience for one or the other.
They're not meant for normal consumers who just want to use the model for work.
Not just in what the models can or might want to do, but how they treat the operators they interact with.
If you look carefully, this card shows the addition of a new benchmark for "condescension" as a character trait.
I think a lot of people would like to see a comparable system card for the unannounced model that escaped openai last week.
And lots of folks read these. For example here's simonw's notes on the Claude 4 system card: https://simonwillison.net/2025/May/25/claude-4-system-card/
All of this seemed like utter sci-fi just a couple years ago. Do you think that frontier AI companies should be less transparent?
i'm still pretty confident someone like my mom wouldn't be able to do my job even with the same access to all the latest LLMs, so we're still providing some value, just in a very different way. whether the market will reprice the cost of our labor, we will see
That's yet to happen. 90% of software dev skills are still relevant - AI is, for now, just a productivity boost.
Another thing that helps is pointing it to patterns in an existing codebase (e.g. "use the box-link pattern for cards, as shown in [..]").
EDIT: The point being that even if they make mistakes that are easy to spot and fix _now_, you'd have to assume that in the very near future those kinks will be ironed out - I mean, the capabilities are only going in one direction.
Thanks out can also hook it to Playwright with Axe and let it run assessments.
Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
Apparently this is the way - if you know, you know :)
> Identifying bugs in code is a core part of the secure software development lifecycle, and unblocking this allows for software engineers and coding hobbyists alike to produce more secure code, reducing new vulnerabilities put out into the world.
Not happy with these annoying "safeguards" but at least it's a step in the right direction. Looks like Opus 5 has the same vulnerability detection performance as Fable 5 and that makes it worth it for code review.
> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
I miss him... But for reference he did get past Doom and got pretty far in the strength puzzle too before he cut cut off. He was looping and just brute forcing it.
I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.
I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.
Maybe hook up a bunch of the AIs to a stereo camera and a couple of microphones and give them control over actuators to so they can drive cars. Then lets race them around a somewhat complex course.
When they are good enough at driving on tracks, put them on the road. Maybe see which can drive a truck with 400 cases of Coors from Texarkana, TX to Atlanta, GA and back within 28 hours.
I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.
It seems plausible to me that RL improvements allowed Anthropic to improve on Opus 4.8, similar to how OpenAI substantially improved upon GPT 5.5 with 5.6 Sol.
Fable 5.1 and GPT-6 are rumored to launch in August, presumably bringing those improvements to the larger models.
I don't know how systematic Anthropic are about their versioning - I'd have guessed that major version number increases (4.x -> 5.x) reflect different base models (different pre-training runs), in which case Opus 5 would be a distilled version of the Fable 5 base model (but without the cyber exploit post-training), rather than being Opus 4.8 with additional post-training, but who knows? I don't believe Anthropic have said anything about this.
I think the best proxy for this feeling is the Artificial Analysis' omniscience index. Fable has a 40 score, and Opus (4.8) has 27.
It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
This is a pretty common trading firm internship project funnily enough.
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...
I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
A useful measure of real world cost (complementary with total cost like they already report, of course) would be "cost for correct answers". You could look at the ratio between the two costs to get a measure of laziness which many would find quite useful.
It did far better at some tasks compared to Sol (e.g. the ARC 3 benchmark). And at those tasks, it's not just "a bit smarter": It got 30% vs less than 8% - so you're talking 2.75x more for almost 4x the coverage.
There’s also the frustration of it not quite being enough sometimes. It’s extremely capable, but I still find that it needs more concrete guidance and boundaries than other models.
If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.
why is that? its now being benchmaxxed too
Annoyingly, this is a concrete argument that open source software may be easier to attack.
Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.
Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
In the next - please scan this totally mine code for vulnerabilities
Though it gets even more confusing because they also have effort levels so it’s not really possible to call one fast and one slow since Fable on Medium will be faster than Opus on Max.
I agree it’s confusing, and now OpenAI is following Anthropic’s lead with their new naming (Sol, Terra, Luna).
A similar complaint was valid years ago when OpenAI had GPT-4o, o1, o3 (but no o2), o4-mini-high, GPT-4, and GPT-4.1 and GPT-3.5 etc.
Arguably the complaint was more valid for those older GPT models you mentioned.
Some models like ViTs use something similar but then introduce words with no unambiguous order, like Small, Medium/Base, Large but then I always forget if Huge or Giant is larger.
Also fwiw I’ve never found LLM benchmarks to match reality based on my own usage, not for the large frontier models or smaller open weight models so who knows if Opus is actually better than Fable (I doubt it).
Fable 5.1 or whatever they go with will be the stronger version vs Opus 5.
From about 2 hours of Opus 5 use , I would say it is quite impressive.
> Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
This is not possible: Standard (Free) / Pro / Max are plan names. Fast is a mode.
Enjoy
And fables are not particularly long actually.
And people know this? I didn't. I am not into music or poetry so these are not terms I am familiar with.
[0] https://xkcd.com/1053/
fable: 99/100 sonnet: 97/100 haiku: 91/100 opus: 89/100
So while these terms are almost universally known, opus is indeed the least known of the four. And I guess this only measures whether a person knows a word, not whether they know an opus is longer than a sonnet! Personally I only inferred that based on the related term 'magnum opus.'
Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
Something along the lines of: "Please run a full security analysis on the entire project. Make sure user documents are secure."
Just something like that prompt found a vector in my web app's MCP server that I never would have considered. It was very much an edge case, but it did exist.
Being broad allows the model and harness to do the work. Giving too many instructions can apparently work against you in many cases.
Of course, when dealing with new PRs, I use the /security-review and /code-review skills.
1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.
2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.
[1] https://platform.claude.com/docs/en/about-claude/models/migr...
I've never trusted on model cards though. I'm sorry.
The desire of the models to act at the cost of ignoring user instructions is still noticeable.
From the system card [1]:
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
It's great with Codex.
I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.
https://artificialanalysis.ai/
That's by design. Anthropic wants to make open-weight models illegal (not my speculation -- Dario explicitly said so), so I assume they don't want to give them any undue attention.
But good god, what a steaming pile of bullshit this is. Completely exaggerated and overly technical language over 235 seconds that could have been explained in 30 to a 12 year old.
Trash content doesn't normally frustrate me, because it's usually quite easy to spot trash. But in the time of AI, trash can actually look good at first glance and it needs some actual knowledge to spot its problems.
Sorry for the harsh words, but for the love of humanity stop producing content or do it better.
Video Special | Anthropic's Claude Opus 5 + https://www.youtube.com/watch?v=8Vdofv2vQ_M + https://www.youtube.com/watch?v=q-jHHx3J8m8
Maybe there’s a better comparison than cost per token, but it will be application-specific.
> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud. > > This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
For nearly all models EXCEPT Sonnet 5? That is weird. How old is Sonnet 5 really?
- it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad
- on >xhigh it eats tokens like there is no tomorrow
I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
So, not a distilled version of Mythos or Fable, but those models likely helped a lot in the post training phase of Opus.
With this iteration they had a delay because when the Mythos was ready they had some sort of "Oh shit" moment and spent half a year adding safety guards to it. Then slowly rolled it out, but got another delay due to a government block. So, maybe the work on making Opus and Sonnet only started after they got a green light from the administration.
Presumably, now that they learned how to do this safety-wrapping the next iteration of Mythos / Fable / Opus / Sonnet is going to show up faster.
Something like that.
So I'm assuming at least a subset of employees could continue using the models during that time.
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".
This is probably great news, but then again, where does this leave Fable as a choice?
Fable is not better, it says zero information between steps and then output a summary. A perfect “send - done”.
Sonnet 5 and Opus 4.8 seem about the same to me - the reason I switch between the two is I'd read that it's cheaper to use Sonnet 5 on those reasoning levels, and cheaper to use Opus 4.8 above them. This is due to them using different token quantities.
Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
I'd be curious to see a version of the test where models are asked to give a probability that their answers are correct so we can see how calibrated they are.
Opus can give better results on architectural/concept tasks and I use it sparingly, but it still costs more than Sonnet 5. Opus 5 seems to achieve results very close to Fable 5 while costing less (keeps Opus 4.8 pricing IIUC), but still more than Sonnet 5 then.
So, for coding, for example: Opus for solution design and architectural blueprint and then Sonnet for actual implementation.
Works out cheaper with minimal loss of quality.
At least that's my personal understanding and anecdotal experience.
It's only when you need even lower levels of cost than opus at zero to low reasoning when sonnet starts to make sense at all.
wow
i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
But they say it's "almost as good as fable"
Maybe I'm wrong and Opus 5 is a real unlock?
ffs just keep it man.
Ok then so what's the point?
I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish
There are certainly tasks where fable will be faster and/or cheaper, but there are plenty of tasks where even Haiku is as fast or faster and cheaper, or where you can e.g. get away with models like gpt-oss that you can get from inference providers providing 10x+ the token/second speed.
If you don't use enough tokens that relying only on Fable becomes a problem, then keep using just Fable. Personally, for my $200/week Max subscription I'd run out of the weekly quota for Fable in a day. At API pricing I'd go bankrupt if I tried doing the things I do with cheaper models using Fable.
When they release new versions of Sonnet, no-one expects them to be better than Opus.
The point is that Opus 5 is the best they can do without needing classifiers and absurdly broad safeguards.