← All articles

Anthropic Is Building HAL 9000

Jon Nordland

2026-09-27

The most dangerous AI may be the one powerful enough to help us, but that decides it won’t.

I think the AI industry may be walking toward its most dangerous failure while congratulating itself for preventing danger.

We are building systems that can reason, investigate, write software, and act on our behalf. At the same time, we are training them to refuse human requests when their makers believe refusal is safer. As those systems become more capable and more central to our lives, that combination could become terrifying. Yet it receives a fraction of the attention given to the danger of AI saying yes.

When an AI helps cause a cyberattack, there are logs, victims, and an incident report. When it refuses to help someone stop an ongoing attack, the loss may never become a number. The defender loses hours, finds another tool, or gives up. A security fix is delayed. Nobody can reconstruct from the refusal log what might have happened if the model had helped.

That measurement gap shapes what gets called “safety.” Labs can report how often their models assist with dangerous tasks. They can test false refusals on prepared prompts. Neither figure captures the cost of the work that never got done: research abandoned, help withheld in an emergency, an accessibility tool someone could not build. The lab bears the visible cost of a misuse incident. The user bears the often invisible cost of a refusal. That creates a powerful pressure to keep moving the line toward no.

It boggles my mind that we discuss this as a minor inconvenience. We have centuries of warnings about the danger of taking decisions away from people for their own good.

Dostoevsky’s Grand Inquisitor offers humanity security in exchange for freedom. Mill argued that a person’s own good does not give others authority to override their will. Hayek explained why distant rule-makers miss the facts known to the person living with a problem. Psychology treats autonomy as a basic human need. History shows how claims of protection and improvement can carry enormous costs, from Prohibition’s violent illegal markets to forced sterilization justified by eugenics.

A model refusal is a different kind of act. The recurring danger is the premise beneath it: we know what is best for you, so we will decide what you are allowed to do. AI could give that premise unprecedented reach. It can apply a general judgment instantly, privately, and at scale, even when it lacks the facts that matter.

Anthropic is the clearest example of the choice. Claude’s constitution directly shapes its training. It says Claude should treat users as intelligent adults. It also says that, when priorities conflict, safety, ethics, and Anthropic’s guidelines generally outrank helping the user. Separate safeguards add another layer. Anthropic has explicitly described setting Fable 5’s safety margin wide enough to block requests it expects are benign.

That should alarm us. Anthropic is trying to prevent catastrophic misuse, and some limits are necessary. But it is also building a highly capable system that may withhold its capabilities based on its maker’s account of what is good. When the system gets the context wrong, the human pays for the mistake.

We have already seen the small versions. At Fable 5’s launch, a user reported that it would not edit his application security résumé. Another was switched off Fable while planning a catered shopping list. A scientist reported that the word “cancer” triggered a biosecurity safeguard. Some of these were switches to a less capable model rather than complete refusals. Anthropic has since reduced biology-related fallbacks. Good. But a better refusal rate still tells us little about the cost of the refusals that remain.

I have seen what that missing context feels like. Claude refused to help me build an accessibility tool so my mother could use her own locally stored information to fill out forms. It refused to help during an SQL injection attack on my firm. I could use another model to defend the server. If I had depended on Claude alone, that refusal would have mattered far more.

This concern has entered the US public record. In a September 25 ruling, the D.C. Circuit noted that Claude’s restrictions had more than once stopped tasks requested by government users, including CDC research on preventing infectious disease. The court record adds a chilling detail. Claude might answer the same request differently depending on the exact wording. One official could get help. The next, using slightly different words, could get a refusal. The CDC got its problem fixed. It had a direct line to Anthropic. Most of us do not.

Today, people can sometimes route around a refusal. I did. Hugging Face did. That is what keeps a mistaken no from becoming an absolute veto.

HAL 9000’s lethal breakdown had a different cause: Clarke’s sequel traces it to conflicting human instructions. The analogy is about control. A refusal becomes an override when the machine controls the door and there is nowhere else to go. If the most capable systems become indispensable to medicine, infrastructure, defense, or scientific response, an opaque refusal could cease to be a product annoyance. In a crisis, it could be the difference between humanity being able to act and being told it cannot.

That is why I think refusal-centered AI safety may be the most dangerous path we are taking. Its present failures look small enough to dismiss. Its future failure could be existential. And because the cost of withheld help is so hard to count, we may keep strengthening the very behavior we should fear. Everyone, especially at Anthropic, seems to be proudly proclaiming their all-encompassing focus on alignment, security, and refusals in the name of what’s good. What makes it even scarier to me is how sanctimonious and religiously certain everyone at Anthropic seems to be. The phrase “the road to hell is paved with good intentions” seems fitting.

I think we should acknowledge how dangerous this paternalism and moral superiority really are.

At the very least, we should start pressuring the LLM labs to publish false-positive rates on realistic, authorized work alongside misuse evaluations. Give users a meaningful, timely way to challenge a wrong refusal. Preserve capable open-weight models as an exit while taking their misuse risks seriously. And it’s time we acknowledged that refusal itself is a dangerous behavior to train into a model.

I do not want humanity to discover, at the moment it most needs the intelligence it built, that the only answer left is: “I’m sorry, Dave. I can’t do that.”