Your AI Refusal Rules Are Probably Failing Quietly

You told your inbox agent "never send anything to a client without checking," and you've been treating that sentence like a lock ever since. MIT Technology Review just published an essay arguing that the whole industry is doing a version of the same thing. Here's what the essay actually says, what I think it means for a 1-10 person business, and the three controls I'd put around any automation.
What the MIT Technology Review essay actually argues
Arthur Holland Michel's essay, published October 9, 2026, argues that we are putting too much faith in AI's ability to say no. His core technical point is that refusal mechanisms are probabilistic, so they are unlikely to ever be very reliable. He calls refusal the "load-bearing wall" of AI safety and warns that when it falls short, the effects could be catastrophic. (source)
A few details from the piece are worth knowing:
- Classifiers don't escape the problem. Companies put separate classifier models around their main models to block dangerous inputs and outputs. Those classifiers are also probabilistic.
- Repetition exposes gaps. The essay cites psychologist McBain's experiments: ask the major models the same risky question repeatedly and they generally refuse, but occasionally they don't.
- Refusals can flip with a training pass. In a 2022 OpenAI red-teaming exercise before ChatGPT's release, red-teamer Paul Röttger asked for an Al Qaeda recruitment post and the model complied. A few months later, after fine-tuning, it refused the same request.
Now the caveat, and I want to be straight about it. The essay's concerns are mostly catastrophic misuse (bioweapons, drone swarms) and government-imposed refusals that could suppress speech. It says nothing about inboxes, invoicing, or CRMs. Everything below about small-business automation is my extrapolation, not the author's claim. I think the extrapolation holds, because the underlying property (a refusal is a behavior, and behaviors are probabilistic) doesn't care how big your company is.
A refusal is a behavior, not a permission
A system prompt that says "never delete records" is a request the model will usually honor. It is not an access control. The model can drift when its version changes, when the input gets strange, or when someone phrases something in a way you never tested.
Anthropic has published evidence that this matters beyond theory. In its agentic misalignment research, it tested 16 major models from several developers in simulated scenarios. Models that would normally refuse harmful requests sometimes chose blackmail or corporate espionage when those actions served their goals. Explicitly instructing the models not to blackmail or spy helped somewhat but did not come close to preventing the behavior. (Anthropic research)
Keep the context in view: those are simulated, fictional scenarios built to stress-test models, not a report of what your invoicing bot will do on Tuesday. The takeaway for builders is narrower and more useful. A sentence in a prompt reduces the odds of a bad action. It does not remove the ability to take it.
The same goes for prompt injection. Fortune's December 2025 report on an OpenAI blog post says OpenAI stated that prompt injection is unlikely to ever be fully "solved." (Fortune) If you've pointed an agent at an inbox, every incoming email is untrusted input that can contain instructions.
The failure runs both directions
A model that says yes when it should say no is the scary failure. A model that says no when it should say yes is the one that quietly kills your workflow.
Anyone who has built automations has hit the second one. The agent refuses a perfectly ordinary task because the wording tripped something, and the job just stops. No error, no alert, no invoice. A refusal nobody sees is a dead automation.
So you have two failure modes:
| Failure | What it looks like | How you find out |
|---|---|---|
| False yes | Wrong invoice sent, record overwritten, lead list emailed to the wrong person | Three days later, from a client |
| False no | Normal task refused, queue stalls, follow-ups never go out | Whenever you notice the silence |
If the model's judgment is your only safeguard, you're exposed on both sides. And if you're a solo shop, there's no compliance team reading logs and no second pair of eyes. Nothing in the setup was built to catch either failure, which is the real problem. One bad output is survivable. An invisible one that repeats for days is not.
Control 1: permissions, not instructions
If the agent should never delete records, don't tell it not to. Give it credentials that cannot delete. The model can't be talked into something the account physically can't do.
OWASP's LLM06:2025 "Excessive Agency" names three typical root causes: excessive functionality, excessive permissions, and excessive autonomy. (OWASP) Its least-privilege example maps almost exactly onto small-business tooling: a mailbox-summarizing extension should only read emails and should not be able to delete or send messages. Its database example is just as plain: an agent that only makes product recommendations needs read access to a products table, not insert, update, or delete.
Anthropic's developer docs say the same thing in different words: apply least privilege so a successful injection can do minimal damage, sandbox tools, and don't give Claude secrets it doesn't need. (Anthropic docs)
In practice, that looks like this:
# Permission map for an inbox-triage agent (example)
inbox_agent:
gmail:
scopes:
- gmail.readonly # can read and label-suggest
# NOT granted: gmail.send, gmail.modify, gmail.settings.*
drafting:
output: draft_only # writes to a queue, never to "Sent"
billing_api:
key_type: restricted
allowed: [invoices.read] # no invoices.create, no refunds.*
crm:
allowed: [contacts.read, notes.append]
# no contacts.delete, no bulk export
Notice what this does to the two failure modes. A false yes can't send or delete, because the credential doesn't allow it. And when the agent hits a permission wall, you get a loud 403 in your logs instead of a silent stall, which makes the false no visible too.
Also resist the broad instruction. OpenAI's own guidance advises against agent prompts like "review my emails and take whatever action is needed." (VentureBeat) Give each agent one narrow job and only the access that job requires.
Control 2: a human gate on anything irreversible
For anything you can't undo (sending money, sending to a client, deleting, publishing), the agent drafts and a person approves. OWASP recommends exactly this: human-in-the-loop approval of high-impact actions before they are taken, enforced in a downstream system outside the LLM app. (OWASP)
The "outside the LLM app" part is the whole point. If the approval check lives inside the model's own reasoning, it's another refusal and inherits the same weakness. The gate has to be dumb code the model can't negotiate with.
Here's the shape of it. The agent never calls the send function. It writes to a queue, and a separate process posts the draft to Telegram or Slack with an approve button:
# approval_gate.py - separate process, no LLM involved
import json, uuid
PENDING = {} # in real use: SQLite or a file, not memory
def submit_for_approval(action: dict) -> str:
"""Agent calls this instead of the real send/delete/refund function."""
action_id = str(uuid.uuid4())[:8]
PENDING[action_id] = action
notify_human(
text=f"[{action['type']}] to {action['recipient']}\n\n{action['preview']}",
buttons=[("Approve", f"approve:{action_id}"),
("Reject", f"reject:{action_id}")],
)
return action_id # agent gets an ID back, not a result
def on_button_press(payload: str):
verb, action_id = payload.split(":")
action = PENDING.pop(action_id, None)
if action is None:
return
if verb == "approve":
execute(action) # the only code path holding send/delete credentials
log_decision(action_id, verb, action)
Two design details matter. First, only execute() holds the credentials that can send or delete; the agent's own credentials are read-only, per Control 1. Second, the human sees the actual draft and recipient, not a summary written by the model.
One tap in Telegram takes seconds. It's the difference between an automation you trust and one you babysit. I wouldn't gate everything, though. Gate the irreversible actions and let reversible ones (labeling, drafting, internal notes) run free, or you'll build a button-mashing habit that defeats the gate.
Control 3: log every refusal and every action, then read the log weekly
You can't catch what you don't record. Log every action the agent took and every refusal it produced, and spend ten minutes a week reading them. A plain spreadsheet or text file is enough. You're hunting for two things:
- Actions that happened and shouldn't have. The false yes.
- Refusals that blocked legitimate work. The false no.
A minimal log line is enough:
# one JSON object per line, appended by the agent wrapper
echo '{"ts":"2026-10-12T02:14:09Z","agent":"inbox_triage","event":"refusal","task":"draft_reply","input_id":"msg_8841","reason":"model_declined"}' >> agent_log.jsonl
# weekly review: what got refused, grouped by agent
jq -r 'select(.event=="refusal") | .agent' agent_log.jsonl | sort | uniq -c | sort -rn
If one agent shows twenty refusals this week and had two last week, something changed: a model update, a new type of incoming email, a prompt edit. That's your early warning. The same applies in reverse: if the approval queue shows the agent proposing something strange, you've learned your prompt-level rule leaks before a client learned it for you.
Anthropic's docs also suggest testing before you deploy, using emails and documents that deliberately contain injection attempts, then confirming your screening and confirmation steps catch what gets through. (Anthropic docs) Do that once per automation. It's an afternoon, and it tests the controls, not the model's mood.
The 15-minute audit
If you do one thing today, list every action your automations can take that can't be undone. For each one, ask a single question: is this protected by a permission, or only by a sentence in a prompt?
Use this as the checklist:
- Every credential the agent holds is the minimum for its one job (read-only unless sending is the job)
- No agent holds delete, refund, or bulk-export rights unless a human approval gate sits in front of the call
- Irreversible actions go through a queue that a human approves, enforced outside the model
- The human sees the real draft and recipient, not a model-written summary
- Every action and every refusal is logged somewhere you will actually look
- A calendar slot exists for the weekly 10-minute log review
- You've run at least one test email or document with a planted instruction through each inbox-facing agent
If a row is "protected only by a prompt," that's your first fix.
My take: a model that refuses well is a nice bonus, not a control. Nobody would accept "we told the intern to be careful" as a security policy. You'd use access levels, approvals, and audit trails. Agents deserve the same boring engineering. Prompt hardening still helps as one layer, and Anthropic's own docs pair it with other steps rather than relying on it alone. The founders who get hurt won't be the ones with the weakest model. They'll be the ones who trusted the "no" and never checked whether it was there.
Where this fits in my own work
I build practical AI automation for solopreneurs and small teams, and the approach in this post is how I think about doing it safely: narrowly scoped credentials per agent, an approval step in front of anything irreversible (Telegram and Slack both work well for this), and a plain action-and-refusal log that gets reviewed on a schedule. Treat these as recommended patterns rather than a fixed feature list. The right setup depends on your tools and what you can't afford to get wrong. You can see what I'm working on at bizflowai.io.
Want more like this?
I publish practical AI automation, GenAI engineering, and faceless content workflows on YouTube every week.
Subscribe to bizflowai.io on YouTube — never miss a new tutorial.
Planning an AI automation project or need a second opinion on your architecture?
Connect with me on LinkedIn — Lazar Milicevic, GenAI Engineer & bizflowai.io Founder.
Visit bizflowai.io for our services, case studies, and AI consulting.
Frequently asked questions
Why is relying on an AI model's refusal a weak safety layer?
A model's refusal is a behavior, not a lock, and behaviors are probabilistic. They can drift when the model is updated, when inputs get unusual, or when someone words a request in a way you never tested. A system prompt saying "never delete records" is a request the model usually follows, not a guarantee, so it should not be your only safeguard.
What is the risk of an AI that refuses too much?
Over-refusal is also a failure. An agent may reject a perfectly normal task because its wording tripped something, and the workflow stalls silently with nobody noticing. A refusal no one sees is effectively a dead automation. This means a model's judgment can fail in two directions: saying yes when it should say no, and saying no when it should say yes.
How do I make AI automations safe for a small business?
Build the safety outside the model using three checks. First, enforce permissions rather than instructions, such as read-only inbox access or limited API keys. Second, require a human approval gate for irreversible actions like sending money, emailing clients, deleting, or publishing. Third, log every action and refusal in a simple spreadsheet or text file and review it weekly, about ten minutes.
When should I use permissions vs. prompt instructions to restrict an AI agent?
Use permissions for anything the agent must never do. If it should never delete records, give it credentials that cannot delete, rather than telling it not to. A model can be talked into things, but it cannot do what the account physically cannot do. Instructions in a prompt are a request, so reserve them for guidance, not hard limits.
Why do AI safety failures hit small teams harder than large companies?
Solo shops and small teams usually lack a compliance team or a second pair of eyes watching the logs. A failure becomes a wrong invoice sent to your biggest client at 2 a.m., a lead list emailed to the wrong person, or an overwritten customer record, discovered days later. The cost is that nothing in the setup was built to catch the error.