Claude Is Getting Dumber. Here’s What’s Really Changing

Picture of Spencer Thomason

Spencer Thomason

September 23, 2026

Copy Link
Claude Is Getting Dumber. Here’s What’s Really Changing

A model can keep the same name, sit behind the same subscription, and still give you a very different experience. That is the issue at the center of a recent analysis of Anthropic’s Claude. According to the analysis discussed in the video, users noticed that Claude felt noticeably weaker after Fable 5 became permanently available through subscription plans. The researcher then measured the model in multiple ways and reported a significant drop in thinking tokens during August compared with July. The more interesting part was that this apparently wasn’t one isolated performance dip. Reasoning reportedly fluctuated over several weeks, with some multi-day drops appearing around product announcements and releases. The question, then, is not simply whether the model itself became worse. It is what inference setup users were actually receiving.

The Model Name Is Not the Whole Product

AI companies often talk about models as if the model name tells you exactly what you are getting. It doesn’t necessarily work that way. The analysis discussed in the video points to a gap between having access to a frontier model and receiving the amount of inference or reasoning that can reproduce the performance associated with that model. The researcher reportedly measured Claude across different invocations and effort levels. Even when users selected high or maximum effort, many calls allegedly returned very little thinking, while some returned none at all. That matters because reasoning is not just a cosmetic feature. If a model can spend more computation working through a difficult problem, the amount of inference being used can affect the result. Two requests can technically be sent to the same named model while receiving very different amounts of reasoning behind the scenes. That creates what the analysis describes as an inference gap. You may have access to the model, but that does not necessarily mean every request receives the same reasoning budget.

The Numbers Are What Make This Interesting

The strongest part of the report discussed in the video is not that someone simply said Claude “felt dumb.” That would be easy to dismiss. The researcher reportedly measured the behavior five different ways and found that August produced dramatically fewer thinking tokens than July. The analysis also claims that the reduction continued across the period rather than appearing as one isolated event. Reasoning reportedly fluctuated across multiple days, and some of those periods appeared to line up with product announcements and releases. That does not automatically prove that those releases caused the changes, but it does make the behavior worth investigating.

Another finding discussed in the video is that even longer reasoning runs reportedly failed to consistently reach the levels associated with published benchmarks. That gets directly into the difference between what a model can demonstrate under a particular evaluation setup and what an everyday user actually receives. If the model performs differently depending on how much reasoning it is actually allowed to use, then the model name alone does not tell the whole story.

Benchmarks Can Show What a Model Can Do

Benchmarks are useful, but they do not automatically tell you what every subscriber experiences. A model can produce impressive results when it is given a particular reasoning budget and evaluation setup. That does not mean every ordinary request receives the same conditions. That distinction is where the skepticism in this story comes from. If a frontier model reaches impressive benchmark results when given substantial reasoning time, but ordinary requests frequently receive much less reasoning, then the benchmark does not necessarily describe the everyday experience.

The transcript describes this as part of the broader problem of “bench maxing.” The point is not that benchmarks are automatically meaningless. The point is that the conditions surrounding a benchmark matter. A model’s published capability and a user’s actual inference experience are not necessarily identical things. What matters to someone paying for the service is not only what the model can achieve under ideal conditions, but what the system consistently delivers when they actually use it.

The Problem Goes Beyond One Performance Drop

The video also points to other reports involving Claude Code, truncated or empty thinking blocks, and API behavior during longer reasoning sessions. These are presented as separate pieces of evidence, but they share a common shape: developers may not always receive the behavior they expect from the model or API. The transcript also discusses an allegation that Claude Code could route requests from Fable 5 to another model without the user receiving the notice that Anthropic’s help documentation allegedly promises. That is a serious claim, but it remains an allegation in the material being discussed and should be checked against the original technical evidence before being presented as established fact.

The larger concern does not depend entirely on proving that allegation. It comes down to control. If a company controls the model, inference budget, routing, infrastructure, and pricing, the customer ultimately depends on that company to determine what happens behind the interface. The user can select a model name, but the underlying system can still involve decisions that are outside the user’s control.

Renting Intelligence Creates a Control Problem

This is where the argument becomes bigger than Claude. If you build important software around a hosted frontier model, you are building on infrastructure controlled by someone else. The provider can change models, routing, inference behavior, limits, pricing, or the way your application interacts with the underlying system. You may still be calling what appears to be the same model, but the behavior underneath can change.

That is the fundamental weakness of vendor dependency. You don’t fully control what you are getting. For experimentation, that may be acceptable. For critical infrastructure, the question becomes much more important because the software may depend on behavior that the company using it does not control. When AI becomes part of an important business system, predictability and ownership become much harder to ignore.

Why Local AI Changes the Equation

The argument behind OpenMonoAgent.ai is built around that control problem. The goal is not simply to use another AI model. It is to move more of the AI stack into an environment the user actually controls. OpenMonoAgent is described in the video as a terminal-native AI coding agent running on local LLMs, with zero API costs, zero telemetry, and full ownership. The broader idea is simple: instead of renting intelligence from a provider and accepting whatever changes happen behind the scenes, run the system on hardware you control.

That approach is about knowing what you are calling and what you are getting because you own the stack. If the hardware and software configuration stay the same, you are not waking up one morning wondering why the same system suddenly behaves differently because a provider changed something on its side. The appeal of local AI here is not about rejecting every hosted model. It is about having an option where control actually belongs to you and where the AI can function more like infrastructure than a subscription you rent.

Security Is Part of the Same Infrastructure Problem

The same thinking extends naturally into security. If your software, AI systems, and infrastructure are built around components you do not fully control, you introduce another layer of dependency into the stack. That is where StartupHakk Security naturally connects to the argument. The video describes StartupHakk Security as a security suite built using the playbook systems developed around OpenMonoAgent. It gives companies a way to scan their website or code while keeping the broader philosophy of controlled infrastructure at the center.

Security cannot simply be something added at the end of development. If the architecture matters to the business, the security of that architecture matters too. The more systems a company depends on, the more important it becomes to understand what is running, where it is running, and who controls it. The same principle applies whether the concern is AI behavior, software infrastructure, or security.

Security Is Part of the Same Infrastructure Problem

Build Infrastructure You Can Actually Own

The bigger lesson is not simply that one AI model may have changed its behavior. It is that the AI industry makes it increasingly easy to confuse access with control. A subscription can give you access to an impressive model, but it does not necessarily give you control over its inference budget, routing, updates, infrastructure, or future behavior.

For businesses building serious custom software, that distinction matters. The goal should not be to chase every new frontier model. It should be to build solid engineering underneath the AI and treat AI as infrastructure that can be controlled, integrated, and maintained as part of the system. That is the philosophy behind OpenMonoAgent and the broader approach at StartupHakk. When AI belongs in a solution, build it into the architecture instead of simply wrapping someone else’s API around the product. When security matters, build it into the system as well.

If your company is building custom software and you want AI integrated into an architecture you can actually control, StartupHakk can help. The focus is straightforward: solid engineering, practical AI, secure infrastructure, and software built to last.

Share this post
Copy Link
Fractional CTO · AI Builds

Stop renting intelligence. Start owning it.

More to explore