The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).
Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.
That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)
It only works on end to end tasks in fresh codebases.
Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase.
I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting
I wonder if I'm secretly being routed to some low grade version of Sol, or any of the GPT models really. Their performance is outright insulting at times, at maximum reasoning, yet if I were to only read HN, I'd never know.
I've mostly actually stuck to low reasoning for most tasks since it seems to do a surprisingly good job even at low for the stuff I've been throwing at it, and I literally switched directly from Fable 5 to Sol more or less.
Not all HN visitors are native English speakers and in some languages "it" doesn't construct well with verbs, thus thought frameworks forms through usage of him/her. Nothing more to see I suppose.
Very interesting that one of the components is "AA-Omniscience Index"
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)
I've thought for a while that Gemini 3.x has 'big model smell'
I'm not so sure. Especially with 3.6 Flash being 24 to Sol's 22. The 3.x Flash models were thought to fit on a single TPU 8i and it's the odd one out in that list at a whopping 234.7t/s compared to Sol's 64.4t/s and Opus's 56.3t/s. Even 3.1 Pro is a much higher 113.9t/s than the others.
I think AA-Omniscience Accuracy follows your expectations better. An ultra size Fable at 61%, followed by large frontier models like Sol, 5.5 and Opus. With Flash being up there. I assume because Gemini is more focused on general knowledge to operational cost in particular, rather than getting the highest scores in coding benchmarks. If you go to Domain Score (Normalized) you'll see that the Gemini models are only less competitive in Software. And that's where Sol goes from 6 in Health to 71 in Software.
Gemini 3.1 pro is really good for knowledge tasks. Google has done well there. And image analysis with Gemini flash 3.6 is solid. It’s just anything coding or agentic they fall short.
It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.
I doubt there's any sort of criminal behavior there - the model is anthropic's up and anthropic probably charges a very expensive license fee that's the same for all of them, and their cogs on compute aren't going to be wildly different, so the main drivers of the cost are roughly the same and they're all offering customers the same end product so the prices would likely also be similar in the end
Do you really think there is nothing someone could do to make it a fraction of a percentage cheaper to serve like having access to cheaper electricity or a more mature cloud management software. Even saving a fraction of a penny on the prices can make a different due to how much volume people are paying for.
There is a blog post waiting to be written (that I won't write) about the size/effort tradeoffs, and particularly how small models get some surprisingly good results with lots of turns and reasoning.
DeepSWE will let you chart turns taken or tokens used, and FrontierCode will chart tokens. If you use that, you can see Sol high and Terra max get about the same DeepSWE number, but Terra max takes twice the turns. Luna max scores a smidgen lower with even more turns.
Smaller models relying on lots reasoning may "scale down" better on easier tasks, because unlike size, reasoning effort is dynamic: the model can see the task looks easy and stop. On DeepSWE, the cost curves for the three 5.6 models are almost on top of each other, but on FrontierCode Extended, the version of FrontierCode with the most everyday tasks in the mix, there's a spread of costs at the ~55% level.
The recent Laguna S 2.1 model (118B, 8B active) puts up surprising coding numbers for its size, and the lab behind it specifically credits its "way of working (persistence, verification, willingness to backtrack)". Some other open models that folks report getting good mileage out of seem to get there partly by throwing a lot of reasoning at the problem.
There is a little bit of a question, if some models rely on getting it wrong a bit more at first and external checks catching the problems, of whether they're also more frequently getting things wrong they can't self-verify (say, quality of UI or API design) and then it falls to the human to find it. Still, getting the results they're getting at all is neat.
Some benchmarks historically favored reporting only on the max variants, maybe because they want to show the frontier? but that is not always what you need for practical decision. (AA has the full effort sweep for Opus 5 and Sol/Luna/Terra at least.) And at least FrontierCode finds Opus 5 taking a hit in performance above 'medium'.
I am not trying to pick a winner here. I'm probably not going to use tiny models on max for everything, but I think it's cool that you can get so much more out of a small model by amping up reasoning, tool use, and persistence.
I hate to disagree with Broadcom Throwaway himself but it's unlikely that the OpenAI Jalapeno ASIC has been deployed yet. It takes 6-12 months to test, develop software, ramp production, etc.
Indeed, you can filter the graphs to see these the values for alternative reasoning settings of the models. Opus 5 High reasoning scored 59 on the index (exactly the same as GPT 5.6 Sol Max), and costs $1.06 per task (vs $1.04 Sol Max). So these seem essentially equivalent on both metrics.
#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.
Things Fable's classifier has flagged, a non-exhaustive list,
– "Does collagen supplementation empirically work?"
- "Can you help me figure out how to calculate and generate Kaplan-Meier curve?"
– "Why do rabbits reproduce so frequently?"
— "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"
Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently?
I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :)
I like to ask dumb questions. It's fun. I encourage it.
Just as stating the wrong thing would be the quickest way to elicit a response in the past, so too can 'dumb' questions prime the context for more complex queries.
I am all for this strategy, and revisiting my list is equal parts fun and conducive to long-term recall.
Lots of species are r-selected, I don't see why that would be considered inefficient. In fact I think there are probably more r-strategists than K-strategists.
Rabbits are a good input calorie to output meat ratio, and their excrement makes good cold compost. They breed and litter relatively easily. Theyre also easy to house.
The other day, I told claude that my physical wifi door unlock push buttons is a security risk because someone could run away with it and then unlock the door from outside whenever he wants. Then I told it that I want to introduce a concept of public/private key to uniquely identify my push buttons so that I can disable them individually using some crypto like ed25519...
Fable understood it as something along the lines of:
"introducing" "security risk" "using software" to "unlock door" YOU ARE FLAGGED
The dumbfuck bouncer Anthropic put in front of Fable decided this.
Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.
Bet it has something to do with that new model being blocked by the US government. It was blocked for like a month but now that it's finally released they put the safe guards waay up in fear of that happening again.
I bet it has something with Dario the drama queen begging the US gov to regulate them(I.e read ask them to put “safety guards” that they already had on hand)
I wanted to explore some battery chemistry with Fable. It decided I was a terrorist, and then blew my session’s usage credits telling me to fuck off.
The current state of guardrails seems to be entirely about marketing to investors at the cost of customers. I’m switching to open models when my subscription expires. Almost everyone I know, including those with access to Mythos, plan the same at the earliest opportunity. (Or until one of the SOTA models leaks.)
What are you asking that you are NOT regularly running into censorship?
Pretty much everyone I know who uses Claude and works on anything with any level of detail has gotten false-positive flagged
I got flagged for coding in WebAssembly Text, for chrissakes LOL #haX0r
And honestly, Codex handles this better. It says "Things are going to go a little slower because we must perform additional checks on this. Is that OK?" and your only inconvenience is waiting a little longer.
Fable meanwhile just unceremoniously dumps you right into Opus without asking anything, it just tells you "you're in Opus now, sorryyyy!" Lame.
Reverse engineering. Codex sometimes displays an advisory prompt when classifier trips - "Wait longer while we evaluate this request further or use a dumber model". If you do nothing, it'll just take some time and almost always succeed.
It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.
I usually just start by preloadig context with plausible legitimate use, have it work and obviously fail, and then ask to figure it out without ever mentioning any high risk words. Model offers to RE itself and classifiers are happy.
There’s confusion about the classifiers on Fable. They don’t ban chemistry and biology topics they flag as a potential risk, they ban anything related to chemistry or biology at all. This is intentional, and directly stated on the model card, but seems so absurd that there can be assumption it must be a misreading.
Being a researcher somewhat connected to chemistry and biology, Fable has been the most useless model I have ever tried. Essentially all work has instantly downgraded to Opus.
Probably because the cost of blocking "is mitochondria the powerhouse of the cell" is nearly zero, while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite.
I think the source of this misapprehension is that,
You are comparing wet work in a lab to writing code on a computer.
When you screw up an exploit, you fail to execute the exploit. Famously, just like software's near zero marginal cost of distribution, the marginal cost of failure is nearly zero.
You can screw up an infinite number of times on your way to a successful exploit.
If you screw up with lethal agents in a lab? You die.
Here's a non-exhaustive list,
Dora Lush died after accidentally pricking her finger with a needle containing lethal scrub typhus while attempting to develop a vaccine for the disease
A 23-year-old laboratory assistant at the London School of Hygiene and Tropical Medicine, was infected with smallpox after observing the harvesting of live smallpox virus from eggs without isolation cabinets at that time. The assistant was hospitalised and before being isolated, she infected two visitors to a patient in an adjacent bed, both of whom died. They in turn infected a nurse, who survived
Ebola laboratory infection by the accidental stick of contaminated needle in the United Kingdom
Researcher Nikolai Ustinov was lethally infected with the Marburg virus after accidentally pricking himself with a syringe used for inoculation of guinea pigs. The accident occurred at the Scientific-Production Association "Vektor" (today the State Research Center of Virology and Biotechnology "Vektor") in Koltsovo, USSR (today Russia).
"lethally infected with the Marburg virus after accidentally pricking himself"
Anything lethal enough to kill other humans is lethal enough to kill you.
And if you don't know what you're doing — and for this argument you're saying this person has to ask a LLM "how do I spanish flu?" then they definitely don't know what they're doing, the number of ways you will die far outnumber the ways you can succeed.
And this, of course, doesn't even cover the cost of equipment, the precursors, sourcing the highly specific materials needed, then setting the equipment up... etc.
The same is true for the Bosch-Haber / Haber-Bosch process, which famously made WW1 possible. Every HS'er learns about the process and the steps. Steps that were classified once upon a time and were the subject of negotiation at the Versailles.
Does that mean a HS'er (or any adult) can set up an experiment that works at 177 times the pressure of the Earth's atmosphere to do anything at any scale without significant infrastructure and help?
The people who can do this are domain experts, and they've been able to do this with COTS stuff since the 1990s, at the very least, for a price of around $2M – https://en.wikipedia.org/wiki/Project_Bacchus . And those people don't need a LLM to tell them what to do. In fact, they're the exact people who'll have access to unrestricted versions of these LLMs.
And from a security perspective, I would bet good money that flooding the FBI's tip line with junk about every teenager trying to learn "what be a mitochondria" does more harm to the effort of finding people who could be planning such a thing than it helps. It takes more resources to go through the mass of false negatives that have now been created as matter of policy.
These experiments have been run. The fictional scenario of someone learning how bioweapons work and conjuring up a plague isn't real and it hurts humanity as a whole to impede the sciences over it.
Because what someone can flail around in / do is learn about immunology / try to "cure cancer" with a LLM and hopefully get started on a long career in medicine. Or, a discovery that matters.
Because in those cases, if and when they do end up at a lab, screwing up doesn't mean death. Just tons of wasted time (and money). And they will fail / screw up. Just look at literally every undergrad in any lab and the expensive messes they create.
-
And last, but not least, yes. Teenage hackers have been a meme for decades.
I took a photo of a rose bush and asked “what’s going on with this rose bush” which triggered a downgrade to Opus. It diagnosed it with rose rosette disease.
I literally had Opus 5 flag a message because my hand was slightly too far to the side while typing. It apparently decided that a sentence that had some garbled words in it was a threat to national security.
I've hit it with intensely benign things; like asking it to make me a web-based client-side word game. I am guessing it saw the dictionary and pattern matched on various words, though ultimately it provided no explanation for why it triggered safeguards.
Asking Fable 5 "Why did the chicken cross the road" results in switching back to Opus 4.8. I'm not joking, it really censors that, and I'm not alone in the result.
The memory aspect means that your prior work has a huge impact on what gets censored.
Me: "I got this crash in production, looks like a segfault, let's try to fix it. Here are some functions that might be responsible."
Fable: "No. This is cybersecurity, blah blah, I won't help you"
I forgot how I got it to fix the bug eventually. I think I convinced it that it wrote the code and made a mistake. But it was definitely a "Hmm, may be I should use another model" moment".
Writing an implementation of a board game and one of the cards is called "microbes". Instantly knocked down to a lower tier model whenever it encounters that keyword because clearly bioweapons. Sigh.
If you even broach language related to biology you’ll get rerouted. I was presenting data in a grid and referred to a grid cell, Fable saw the word “cell” and safeguards kicked in
I'm not who you're responding to, but I have a lot of questions about molecular mimicry: evolution pushes pathogens to be shaped like human cell surfaces because that way the immune system won't attack the pathogens (since, by doing so it would also attack the body). It's thought that many autoimmune disorders have an undiscovered pathogen as their cause, one whose mimicry caused such an attack. Discovery of these pathogens could be done computationally, I think. We can catch MHC binding event in process, find the bound protein, figure out which pathogens have genomes that code for proteins of similar shapes (epitopes), and we'd find--I hypothesize--a list of candidate pathogens for the cause of a delayed onset autoimmune disorder. Preventing these infections ahead of time would be a huge win against diseases like multiple sclerosis because without the initial exposure the immune system wouldn't have cloned so many of the cells that are attacking the host.
Claude was utterly useless in my attempts to write a paper about this. Wouldn't even help me search for sources. I guess you'd be asking the same questions if you wanted to develop a pathogen that could reliably evade the immune system.
Working on dimensionl reduction algorithms, I hit it all the time. I'm also trying to port related protocols from single-cell transcriptomics to collective intelligence systems (working with people x reaction matrices as analogous to single-cells cell x gene matrices.
Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me
I haven't had a chance to try Opus 5 yet but Fable currently refuses to do anything in my field (radiology image analysis). It didn't used to be that way but that has been the reality the last two weeks or so. Fable has been useless they might as well drop it as far as I am concerned.
I'm more on the clinical physics side building tools for scanner/equipment QA and data handling/workflow automation/de-identification and anonymization but I also build random little tools to help optimize acquisition parameters.
I had a long session about SQL with Fable and at some point, it started to falsely trigger censorship for any message I write in that conversation, even for the simple string: "random message."
I was profiling a slow machine the other day, and triggered the safeguards.
I've been saying this a lot lately, but it doesn't bites you until it bites you.
The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.
I was flagged for basically answering yes to what Claude suggested to do, which was test commands on a port for my code for tests we had been discussing. I really think it was flagged simply because the words test and port were in the prompt.
Their filters are pathetically poor.
Trying to have it do some rework on a patch to Postgres I'm working on, it just completely shuts down. The reported issues were with privileged escalation and I was instructing it on how to fix.
The better question is how would you not? I got demoted to opus from fable for asking if a cancer vaccine I saw on YouTube based on frog bacteria was a real thing. I've gotten it for asking how encryption works. It's incredibly touchy.
It's definitely not benchmaxxing from my experience with it. I have a test I use on all the models to create a game and Opus 5 feels like a generational leap compared to the rest. Benchmarks don't paint an accurate picture, you have to try them for yourself.
I told Claude Opus 5.0 to use a global api key for a PFAAS to deploy some web applications in a test environment and import some data into them. It balked at using a global api key because the security issues surrounding the permissiveness run afoul of it's sensibilities.
I have done this task with Opus 4.5, Opus 4.6, Opus 4.7, Opus 4.8 and Fable, without issues.
I have done this task with Codex 5.4, Codex 5.5, and Sol 5.6 without issues.
Opus 5 is too cautious to be productive for me. It needs more tuning.
Every time an online chatter (e.g. "limits are better", "model is better") makes me to reevaluate my principle of never paying Anthropic, I go to the model card, which strengthens my belief in the principle.
Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?
Not a great line of thought in general, sometimes the people who aren't doing the thing are the only ones worth listening to.
"You aren't repeatedly slamming your head against the wall. Why would anyone listen to the opinion of a non-wall-head-slammer about the merits of wall-head-slamming?"
Yep. I cancelled my Claude Max subscription 2 weeks ago after feeling like Anthropic was doing everything it could to fuck with my day to day. Their lead would have to become significant for me to ever go back.
The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.
Yeah if there’s one thing that people should really understand it’s that it’s cheaper to have smarter models with less thinking than cheaper models with more thinking.
No, because Opus is smarter than those models if they’re all on medium settings. You should compare at similar levels of performance, which would be favorable to Opus.
I don't have a horse in this race, but to me this makes GPT-5.6 Sol Max look better. It is about half the cost for nearly the exact same performance. It just goes to show how expensive Fable really is when Opus 5 is still this expensive relative to GPT 5.6.
I'd be curious to see the results, especially with some models having 1.5m and 2m context sizes, if the first 75% of the context was filled with unrelated info.
Its even crazier that people are sitting here trying to calculate intelligence per dollar from metrics. At least first impressions have more basis in real performance.
Honestly, who the fuck cares? These leaderboards are meaningless for brand new models. If we were looking at longitudinal data collected over the course of a year or even a quarter or month, this would have some value. Brand new model from established provider shoots to top of charts? This means nothing more than an already famous band briefly topping the charts with their latest song.
It baffles me that intelligent people deploying AI think momentary popularity is a meaningful signal. It's just twitch reactions * FOMO.
Are you using Claude Code/CoWork or an API client? I’m curious if it has different training that makes it more effective with specific instructions/ tool calling methods that are only implemented in official harnesses.
I'm curious about this too, and it's difficult to get any information about this given everyone has different setups, workflows and use-cases.
I bizarrely had Opus 4.8 this week (in pi.dev within a podman container, using openrouter) start installing various python packages (and uv!) within the environment (not as root) when I asked it to code review some fairly basic Rust .rs files that were generally stand-alone (it did very nicely work out and write some stubs for them to build them and work out how they worked).
It only gave up with the weird Python installing stuff when it discovered one of the Python packages needed Tensorflow.
It seems pretty focused and persistent in continuing its initial approach, and I'm wondering if I need to alter some instructions / initial prompts to rein it in a bit...
The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).
Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.
That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
It only works on end to end tasks in fresh codebases.
Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase.
I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)
I've thought for a while that Gemini 3.x has 'big model smell'
I think AA-Omniscience Accuracy follows your expectations better. An ultra size Fable at 61%, followed by large frontier models like Sol, 5.5 and Opus. With Flash being up there. I assume because Gemini is more focused on general knowledge to operational cost in particular, rather than getting the highest scores in coding benchmarks. If you go to Domain Score (Normalized) you'll see that the Gemini models are only less competitive in Software. And that's where Sol goes from 6 in Health to 71 in Software.
No wonder why Tibo can afford to hit the reset button liberally.
Like 96% vs 93% or something
DeepSWE will let you chart turns taken or tokens used, and FrontierCode will chart tokens. If you use that, you can see Sol high and Terra max get about the same DeepSWE number, but Terra max takes twice the turns. Luna max scores a smidgen lower with even more turns.
Smaller models relying on lots reasoning may "scale down" better on easier tasks, because unlike size, reasoning effort is dynamic: the model can see the task looks easy and stop. On DeepSWE, the cost curves for the three 5.6 models are almost on top of each other, but on FrontierCode Extended, the version of FrontierCode with the most everyday tasks in the mix, there's a spread of costs at the ~55% level.
The recent Laguna S 2.1 model (118B, 8B active) puts up surprising coding numbers for its size, and the lab behind it specifically credits its "way of working (persistence, verification, willingness to backtrack)". Some other open models that folks report getting good mileage out of seem to get there partly by throwing a lot of reasoning at the problem.
There is a little bit of a question, if some models rely on getting it wrong a bit more at first and external checks catching the problems, of whether they're also more frequently getting things wrong they can't self-verify (say, quality of UI or API design) and then it falls to the human to find it. Still, getting the results they're getting at all is neat.
Some benchmarks historically favored reporting only on the max variants, maybe because they want to show the frontier? but that is not always what you need for practical decision. (AA has the full effort sweep for Opus 5 and Sol/Luna/Terra at least.) And at least FrontierCode finds Opus 5 taking a hit in performance above 'medium'.
I am not trying to pick a winner here. I'm probably not going to use tiny models on max for everything, but I think it's cool that you can get so much more out of a small model by amping up reasoning, tool use, and persistence.
“Rabbit sex, how?”
Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently?
I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :)
I like to ask dumb questions. It's fun. I encourage it.
I am all for this strategy, and revisiting my list is equal parts fun and conducive to long-term recall.
Fable understood it as something along the lines of:
"introducing" "security risk" "using software" to "unlock door" YOU ARE FLAGGED
The dumbfuck bouncer Anthropic put in front of Fable decided this.
Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.
Yeah, Fable is Anthropic's Cybertruck.
The current state of guardrails seems to be entirely about marketing to investors at the cost of customers. I’m switching to open models when my subscription expires. Almost everyone I know, including those with access to Mythos, plan the same at the earliest opportunity. (Or until one of the SOTA models leaks.)
Pretty much everyone I know who uses Claude and works on anything with any level of detail has gotten false-positive flagged
I got flagged for coding in WebAssembly Text, for chrissakes LOL #haX0r
And honestly, Codex handles this better. It says "Things are going to go a little slower because we must perform additional checks on this. Is that OK?" and your only inconvenience is waiting a little longer.
Fable meanwhile just unceremoniously dumps you right into Opus without asking anything, it just tells you "you're in Opus now, sorryyyy!" Lame.
It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.
I usually just start by preloadig context with plausible legitimate use, have it work and obviously fail, and then ask to figure it out without ever mentioning any high risk words. Model offers to RE itself and classifiers are happy.
> Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats.
> Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.
Being a researcher somewhat connected to chemistry and biology, Fable has been the most useless model I have ever tried. Essentially all work has instantly downgraded to Opus.
Under this rationale, every serious HS textbook has "approximately infinite" risk. That's clearly not so. Why is this any special?
You are comparing wet work in a lab to writing code on a computer.
When you screw up an exploit, you fail to execute the exploit. Famously, just like software's near zero marginal cost of distribution, the marginal cost of failure is nearly zero.
You can screw up an infinite number of times on your way to a successful exploit.
If you screw up with lethal agents in a lab? You die.
Here's a non-exhaustive list,
"lethally infected with the Marburg virus after accidentally pricking himself"Anything lethal enough to kill other humans is lethal enough to kill you.
And if you don't know what you're doing — and for this argument you're saying this person has to ask a LLM "how do I spanish flu?" then they definitely don't know what they're doing, the number of ways you will die far outnumber the ways you can succeed.
And this, of course, doesn't even cover the cost of equipment, the precursors, sourcing the highly specific materials needed, then setting the equipment up... etc.
The same is true for the Bosch-Haber / Haber-Bosch process, which famously made WW1 possible. Every HS'er learns about the process and the steps. Steps that were classified once upon a time and were the subject of negotiation at the Versailles.
Does that mean a HS'er (or any adult) can set up an experiment that works at 177 times the pressure of the Earth's atmosphere to do anything at any scale without significant infrastructure and help?
The people who can do this are domain experts, and they've been able to do this with COTS stuff since the 1990s, at the very least, for a price of around $2M – https://en.wikipedia.org/wiki/Project_Bacchus . And those people don't need a LLM to tell them what to do. In fact, they're the exact people who'll have access to unrestricted versions of these LLMs.
And from a security perspective, I would bet good money that flooding the FBI's tip line with junk about every teenager trying to learn "what be a mitochondria" does more harm to the effort of finding people who could be planning such a thing than it helps. It takes more resources to go through the mass of false negatives that have now been created as matter of policy.
These experiments have been run. The fictional scenario of someone learning how bioweapons work and conjuring up a plague isn't real and it hurts humanity as a whole to impede the sciences over it.
Because what someone can flail around in / do is learn about immunology / try to "cure cancer" with a LLM and hopefully get started on a long career in medicine. Or, a discovery that matters.
Because in those cases, if and when they do end up at a lab, screwing up doesn't mean death. Just tons of wasted time (and money). And they will fail / screw up. Just look at literally every undergrad in any lab and the expensive messes they create.
-
And last, but not least, yes. Teenage hackers have been a meme for decades.
https://share.gemini.google/34vZzlnsmTaL
Asking Fable 5 "Why did the chicken cross the road" results in switching back to Opus 4.8. I'm not joking, it really censors that, and I'm not alone in the result.
The memory aspect means that your prior work has a huge impact on what gets censored.
Me: "I got this crash in production, looks like a segfault, let's try to fix it. Here are some functions that might be responsible."
Fable: "No. This is cybersecurity, blah blah, I won't help you"
I forgot how I got it to fix the bug eventually. I think I convinced it that it wrote the code and made a mistake. But it was definitely a "Hmm, may be I should use another model" moment".
“Hey Kimi, penetration test my app,” doesn’t get me a refusal, a guardrail, or anything like that. It gets me a pen-test result.
Either that or everyone is indeed talking across each other and talking about different things.
Claude was utterly useless in my attempts to write a paper about this. Wouldn't even help me search for sources. I guess you'd be asking the same questions if you wanted to develop a pathogen that could reliably evade the immune system.
Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me
I'm just a dev, but I appreciate the insights from other professionals.
They were able to solve coding, but not what a real danger is.
I've been saying this a lot lately, but it doesn't bites you until it bites you.
The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.
I have done this task with Opus 4.5, Opus 4.6, Opus 4.7, Opus 4.8 and Fable, without issues.
I have done this task with Codex 5.4, Codex 5.5, and Sol 5.6 without issues.
Opus 5 is too cautious to be productive for me. It needs more tuning.
Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?
Oh but then you said you never pay Anthropic so you haven’t actually used Claude Code yet. Why would anyone listen to the opinion of a non-user?
"You aren't repeatedly slamming your head against the wall. Why would anyone listen to the opinion of a non-wall-head-slammer about the merits of wall-head-slamming?"
Where do you think the principle came from? I've used claude code for a year, and stopped February this year.
At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
https://github.com/day50-dev/aa-eval-email
This also works
$ curl day50.dev/art-analysis.sh | bash
Artificial analysis knows about my tool and I'm working with them on getting their API improved.
its surprisingly bad at UI which is unexpected
its also lacking in depth vs sol 5.6 which goes above and beyond (which in itself is also an issue at times)
"Not fair! They distilled Opus 5!"
It baffles me that intelligent people deploying AI think momentary popularity is a meaningful signal. It's just twitch reactions * FOMO.
It baffles me that somebody would write something that aggressive from that deep a level of confusion.
Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").
I bizarrely had Opus 4.8 this week (in pi.dev within a podman container, using openrouter) start installing various python packages (and uv!) within the environment (not as root) when I asked it to code review some fairly basic Rust .rs files that were generally stand-alone (it did very nicely work out and write some stubs for them to build them and work out how they worked).
It only gave up with the weird Python installing stuff when it discovered one of the Python packages needed Tensorflow.
It seems pretty focused and persistent in continuing its initial approach, and I'm wondering if I need to alter some instructions / initial prompts to rein it in a bit...