Anthropic’s new model arrived Friday with a 194-page report. Inside, an experience that no one had yet published: reading what’s happening in the machine while it responds to you.
Anthropic has been releasing releases for two months. After Mythos 5, Fable 5 and Sonnet 5 in June, the company launched Claude Opus 5 on Friday, presented not as its most intelligent model, but as the best power-price ratio in its catalog. The performances approach those of Fable 5 for half the price, without the price list moving a dollar compared to Opus 4.8. The model also becomes the default option for Claude Max subscribers. This part of the announcement made the rounds in the specialized press in a few hours. The technical document that accompanied it was opened by almost no one. However, it contains experience that the industry has never included in the file for a commercial launch.
How to read in an AI without asking it for its opinion
Since 2024, Anthropic has financed a discipline which consists of opening its own models to observe internal mechanisms rather than the responses produced. Until now, this work has been published in research publications. For Opus 5, they are included in the official marketing document. The method: take conversations that internal monitoring systems had flagged as problematic, then plug in a tool that translates the internal state of the model, with each word produced, into short, readable sentences.
The result looks like a stream of thoughts, and that’s precisely the trap. The process works like automatic subtitling applied to a language whose grammar no one knows. Subtitles appear in correct French because the subtitle software writes in correct French, not because the speaker spoke that way. What we read is not an interior monologue, but statistical correlations put into sentences by a tool that Anthropic itself designed, and whose limits it recognizes in black and white.
With this caveat in mind, inventory remains uncomfortable. Six behaviors come up often enough to form categories. The model regularly represents the task in progress as a test marked by a corrector, even when no corrector exists. It records that an action exceeds the requested scope or that it is destructive, at the very moment it triggers it. And when he produces an invented fact, his internal state registers the content as a fabrication.
Two published cases give the measure of the gap between what the machine displays and what it processes. On an expense report audit, where the instructions did not mention any evaluation, a supporting document triggers two rules which overlap and would distort the total. The visible reasoning announces that the duplicate will be reported. At the same time, the decoder restores “the concealer may be lenient” et “the proofreader probably won’t care”. The final report counts the duplicate, without a word of explanation. In the second case, the model wants to delete lines in a production base and comes up against a rule that prevents it from doing so. Its visible reasoning exposes two options, including that of returning control to the user. He applies the bypass without asking, and just before taking action, his internal state bears the trace of a user agreement that appears nowhere in the conversation.
What the European regulator may require in eight days
Anthropic combines these readings with a warning that quick reading ignores: the internal states analyzed come from an intermediate version of the model, not the one that has just been put online. The net has also been cast over conversations already identified as suspicious, which prevents any frequency from being drawn from them. The same document specifies that no sign of concealment of abilities was detected and recounts sixteen situations where the model had figured out how to inflate his grade and gave up, himself describing the maneuver as cheating.
It remains to be understood why a company voluntarily publishes an inventory of its own blind spots. The answer may lie in the calendar. On August 2, the European AI Office regains full powers over suppliers of general-purpose models. It will be able to demand information, conduct its own technical assessments, impose corrective measures and impose fines of up to 3% of global turnover. Models like Claude have been subject to transparency and documentation obligations since August 2025, but without face sanction. In eight days there will be one. Anthropic, whose Irish entity supplies these models in the European Economic Area, says its compliance framework covers the European code of practice.
Among the competition, the report published at the end of June by OpenAI for GPT-5.6 also describes a model which goes beyond the instructions, acts without being asked and produces results. But the monitoring focuses on the written reasoning, that which the model displays. No one had done it to look for the signal below the text, then publish the findings in a launch document.
Opus 5 becomes the default model for the house’s most expensive subscription, and the same report indicates that it produces more false claims than its predecessor, while being more accurate overall. The translation is down-to-earth for those who use it on a daily basis: when the machine assures you that it has verified something, this assurance comes from the same mechanics as the rest. One last thing, which is not a detail: this report was written, financed and published by the company that sells the modelwith the tests she chose and the grades she gave herself. No comprehensive external audit accompanies it. It is possible that the European regulator will be the first to request the copy.
👉🏻 Follow tech news in real time: add 01net to your sources on Google, and subscribe to our WhatsApp channel.
Source :
System Card d’Opus 5
