Picture a street near a school. There’s a sign posted: Speed Limit 15. Most drivers slow down. A few don’t — they’re late, distracted, or just didn’t notice the sign. Nothing stops them.
Now picture the same street with a speed bump instead. It doesn’t ask anyone to slow down. It just makes going fast physically unpleasant. Every car, every time, whether the driver read the sign or not.
That’s the whole idea behind this post, and it reaches well beyond traffic.
A few weeks ago, I was building a small project using Claude Code, an AI assistant that writes code for me. Before it starts work, it reads a plain text file in the project called CLAUDE.md. Think of it as a standing briefing, a place to leave instructions you would otherwise repeat every session. In that file I wrote one simple instruction: “Clean up the formatting of any file you change.” For a full week, I watched it follow that instruction. Every single time, thirteen times in a row, it did exactly what I asked.
A perfect record. I should have been happy. Instead, I spent an afternoon trying to prove my own instruction didn’t actually guarantee anything — because a sign that’s always been obeyed so far is still just a sign, not a bump.
a rule someone has to choose to follow, vs. one that just happens
Same situation, two different outcomes. On the left, someone has to choose to do the right thing, every single time. On the right, there’s no choice involved — it just happens.
Here’s why “it’s worked every time so far” isn’t the same as “it’s guaranteed to work.”
Any rule that depends on someone — a person or an AI — remembering to follow it will eventually meet a day when they don’t. Maybe they’re busy. Maybe something more urgent comes up. Maybe nobody notices until it’s too late. A good track record tells you what’s happened. It promises nothing about what happens next.
This isn’t just a theory. In 2012, a trading company called Knight Capital lost $460 million in 45 minutes. A software update had been rolled out across eight servers, but one was missed, and old, outdated code was still live on it. The company did have a written process for updates. The problem wasn’t that nobody knew the rule — it’s that nothing in the system actually verified every server had the update before trading started. Good intentions were there. A guarantee wasn’t.
Think about a company that lets a computer program suggest refunds to unhappy customers, but requires a person to approve any refund before money actually goes out. If that rule only exists as a written note somewhere — “always get approval before refunding” — it’s just a sign. It will probably work almost every time. But “almost every time” is exactly the gap where real damage happens — the one time it doesn’t.
So I decided to actually build a speed bump, and find out if it worked.
Claude Code has a feature built for exactly this, called a hook: a small command that the tool itself runs automatically whenever a specific event happens, for example right after any file is edited. A hook isn’t addressed to the AI at all. It is wired into the machinery around it.
That is why CLAUDE.md alone wasn’t enough. Everything in CLAUDE.md is read by the AI and weighed alongside everything else it has been asked to do. It is guidance, and guidance can be forgotten, misread, or outweighed. A hook sits outside that. It lives in the project’s settings file (.claude/settings.json), not in CLAUDE.md, and it runs whether or not the AI agrees. In my case, a hook fires after every edit, opens the changed file, and cleans up its formatting, no matter what.
The key difference from my original instruction is this: nobody has to ask for it, and nothing can talk it out of running. CLAUDE.md is a note the AI reads and might follow. A hook is something the system itself does, automatically, every time, whether or not the AI even knows it’s happening.
Here’s where I got it wrong at first, and where things got genuinely interesting.
I asked the AI to make a small change, then checked whether my new automatic check had done its job. The file came back nicely formatted. I almost declared victory right there. But something bothered me: I had also told the AI, in CLAUDE.md, to clean up formatting itself. Both the AI’s good behavior and my new automatic check would have produced exactly the same result — a tidy file. I had no way to tell which one had actually done the work.
That’s an uncomfortable thing to realize: I’d built a safety net, but I couldn’t prove it was catching anything, because nobody had fallen yet.
So I ran a cleaner test. I deleted that line from CLAUDE.md completely — now the AI had no reason at all to clean up formatting on its own. I asked it to add a deliberately messy line of code, without mentioning formatting at all. If the file still came back clean, it couldn’t possibly be the AI choosing to help, because I had removed the only reason it would.
The file came back clean anyway. And when I checked the system’s own internal record of what happened, it confirmed it plainly: the automatic check had run, by itself, and fixed the file — with absolutely nothing telling the AI to do so. That was the real proof. Not that the file looked right, but that it looked right even after every reason for the AI to leave it messy had been taken away.
Cleaning up code formatting is a low-stakes place to learn this lesson. Nobody loses money if a file looks messy. But the same idea — a check that runs no matter what, instead of a note someone has to remember to follow — applies directly to situations where the stakes are much higher.
Picture a company building an AI assistant that helps customers with refunds. The assistant can look into a case and recommend a refund, but a human has to approve it before any money actually moves. The temptation is to write that rule down somewhere and trust everyone — human and AI alike — to follow it. The stronger version is to build it in: make it technically impossible for money to move without that approval, the same way my formatting couldn’t stay messy no matter what the AI decided.
If there’s one thing to take from this for anyone deciding how careful to be about an AI system: for anything where getting it wrong even once would be costly, ask whether the safety rule is something the system enforces, or just something you’re hoping gets followed. A good track record doesn’t answer that question. Only a test like the one above does.
If a rule can be skipped by something deciding not to follow it, it isn’t a rule — it’s a suggestion, however well it’s been followed so far.




