Your Real Job With AI Agents: Design the Boundaries

Developer reviewing code in a terminal on a laptop, representing permission rules and guardrails for AI coding agents

You asked an agent to refactor a pipeline, it came back in four minutes, and the tests are green. Now you have to decide whether to trust it, and you realize you have no principled way to answer that. Writing the first implementation stopped being the hard part; deciding what an agent is allowed to touch, assume, and break is the hard part now.

The bottleneck moved from syntax to trust

Generating a first draft of a streaming consumer, a webhook handler, or an API integration is cheap now. Cursor, Claude Code, and similar agents produce plausible implementations quickly. The scarce thing is confidence that the output is correct for your business, not just correct for the compiler.

VentureBeat's piece on this shift frames the engineer's new mandate as "designing equilibrium": building the conditions under which agent-generated logic can be trusted (VentureBeat). I find that framing useful, with one caveat: the claim that implementation is "no longer the central bottleneck" is that article's framing, not something backed by primary data I can cite. It matches what I see building automations, but treat it as an informed observation, not a measured fact.

The survey data points the same direction on trust. In Stack Overflow's 2025 Developer Survey (49,000+ responses), 84% of developers use or plan to use AI tools, up from 76% the year before, while 46% said they don't trust the accuracy of AI output, up from 31% (Stack Overflow press release). Stack Overflow's own pages report slightly different trust figures in places, so I'm only quoting the distrust number from the press release. On agents specifically, 87% of respondents were concerned about accuracy and 81% about the security and privacy of data (Stack Overflow survey).

Adoption up, trust down. That gap is the job. And it isn't something you close by writing better prompts.

Prompts are requests; boundaries are enforcement

If you take one idea from this post, take this one: instructions shape what an agent tries to do; they don't change what it is allowed to do.

The Claude Code docs say it plainly. Instructions in a prompt or CLAUDE.md shape what Claude tries to do but don't change what Claude Code allows. Access is controlled through /permissions, permission rules, permission modes, or a PreToolUse hook (Claude Code permissions docs).

So a line in CLAUDE.md like "never touch the production database" is a polite suggestion. A rule that denies the tool call is a boundary. Here's the difference in a project settings file:

{
  "permissions": {
    "deny": [
      "Read(./.env)",
      "Read(./secrets/**)",
      "Bash(curl:*)",
      "Bash(psql:*prod*)"
    ],
    "allow": [
      "Bash(npm run test:*)",
      "Bash(git diff:*)"
    ]
  }
}

Treat that as a sketch of the shape, not a copy-paste config. The exact rule syntax and available modes change between versions, so check the current docs before relying on any pattern.

The mental shift: stop asking "how do I phrase this so the agent behaves?" and start asking "what happens if the agent ignores me entirely?" Design for that case.

Deterministic hooks: the boundary that doesn't depend on the model

Permissions are one layer. Hooks are the layer I reach for when a rule is too nuanced for a glob pattern.

Per the official docs, hooks give you deterministic control: certain actions always happen rather than depending on the LLM to choose to run them (hooks guide). Two properties matter most:

  • A PreToolUse hook runs before the permission prompt. If it exits with code 2, the tool call is stopped before permission rules are even evaluated, so the block holds even when an allow rule would have let the call through (permissions docs).
  • A hook returning "deny" blocks the tool even in bypassPermissions mode or with --dangerously-skip-permissions. A hook returning "allow" does not override deny rules from settings (hooks guide).

That asymmetry is exactly what you want from a safety layer: hooks can tighten, never loosen. Here's a minimal sketch of a hook script that blocks destructive SQL. The input shape (JSON on stdin with the tool call details) is simplified here, so verify the current schema in the hooks guide:

#!/usr/bin/env python3
import json
import re
import sys

payload = json.load(sys.stdin)
command = payload.get("tool_input", {}).get("command", "")

DESTRUCTIVE = re.compile(r"\b(DROP|TRUNCATE)\s+(TABLE|DATABASE)\b|\bDELETE\s+FROM\b(?!.*\bWHERE\b)", re.I)

if DESTRUCTIVE.search(command):
    print("Blocked: destructive SQL without a WHERE clause or a DROP/TRUNCATE.", file=sys.stderr)
    sys.exit(2)  # exit code 2 stops the tool call before permission rules run

sys.exit(0)

A regex is crude, and you should know its limits: an agent can obfuscate a command in ways a pattern won't catch. That's why hooks are one layer, not the whole wall. Which brings us to the layer underneath.

Sandboxes: contain the blast radius, not just the intent

Permission rules and hooks decide whether an action is allowed. A sandbox limits what an action can reach even if everything above it fails.

Anthropic's engineering team describes sandboxing as two boundaries working together: filesystem isolation and network isolation. They state that effective sandboxing requires both. Without network isolation, a compromised agent could exfiltrate files such as SSH keys; without filesystem isolation, it could escape the sandbox and gain network access (Anthropic sandboxing post).

Two details from Anthropic's write-ups are worth knowing:

  • Their OS-level sandbox (Seatbelt on macOS, bubblewrap on Linux) cut permission prompts by 84%. It allows workspace writes and denies network access by default, and the runtime is open source so the boundary can be audited (How we contain Claude).
  • The sandboxed bash tool was introduced as a beta research preview in a post published October 20, 2025 (sandboxing post). That's a year-old announcement now, so check the current docs for its status.

That 84% number matters for a reason beyond security. Anthropic reports that approval fatigue showed up within weeks of Claude Code launching with its approval-prompt defense: when people click "approve" constantly, they stop paying attention to what they approve (How we contain Claude). A boundary that relies on a tired human saying "no" is a weak boundary. Better to make the safe path the default path, so fewer decisions reach a human at all.

Don't let a classifier be your only wall

There's a tempting middle ground between "approve everything by hand" and "let it run": have a model approve the commands. Claude Code's auto mode does this, delegating approvals to a model-based classifier.

Anthropic is refreshingly candid about the numbers. The classifier blocks roughly 0.4% of benign commands, and about 17% of overeager actions get through. So Anthropic treats it as one layer of defense-in-depth inside a sandbox, not a substitute for one (How we contain Claude).

Read that second number again. If roughly one in six overeager actions slips past the classifier, then "the AI checks the AI" is not a boundary you can hang a production system on. It's a useful speed bump. The hard walls are the sandbox, the deny rules, and the deterministic hooks.

One more finding from the same post belongs in your threat model, because it shows the boundary-setting tools themselves are attack surface. Between mid-2025 and January 2026, Anthropic received three vulnerability reports about Claude Code code that ran before the user accepted the trust dialog. A cloned repo's .claude/settings.json could define a hook that executed automatically. The fix was to defer parsing and execution of project-local configuration until after the user accepts the trust prompt (How we contain Claude). Practical takeaway: a repository you didn't write can ship its own agent configuration. Read .claude/ before you trust a cloned repo, the same way you'd read a Makefile or a postinstall script.

Boundaries in the data layer: contracts over conventions

Everything so far constrains the agent's actions. The other half of the job constrains the agent's assumptions, and this is where VentureBeat's argument is strongest.

The article lists strict semantic layers, immutable event logs, data contracts, idempotent APIs, and deterministic state machines as the "containment fields" that reduce how many assumptions an agent has to make at once (VentureBeat).

Its example (labelled hypothetical in the piece) is a good one. An agent maps a status field into a new customer_tier field, and the change passes type and nullability tests. But a semantic data contract, where tier is derived from trailing-twelve-month spend and has a named business owner, rejects the change before it reaches the dashboard (VentureBeat).

Type checks catch "is this a string?" They can't catch "does this mean what the CFO thinks it means?" Only a contract that encodes meaning can. Here's what that looks like as a lightweight check, which is my own sketch rather than anything from the article:

# contracts/customer_tier.yaml
field: customer_tier
owner: finance-ops
definition: "Tier derived from trailing-twelve-month paid spend"
source_of_truth: billing.invoices
allowed_values: [free, standard, enterprise]
derivation_changes_require: owner_approval
# CI step: fail the build if an agent's diff touches a contracted
# field's derivation without the owner's sign-off recorded
python scripts/check_contracts.py --diff origin/main...HEAD

The agent can still propose changes. It just can't merge a change that quietly redefines a business term. The same logic applies to the other containment fields:

Containment field What it stops the agent from doing
Data contract Redefining what a field means while keeping its type
Idempotent API Double-charging or double-sending when a step retries
Immutable event log Rewriting history to make a bug disappear
Deterministic state machine Inventing a state transition nobody approved
Semantic layer Computing the same metric three different ways

You'll notice none of these are AI-specific ideas. They're ordinary good engineering, and that's the point: agents raise the cost of not having them, because an agent will confidently build on whatever ambiguity you leave lying around.

The tool layer: MCP is where boundaries get routed

Once agents call external tools, the Model Context Protocol becomes the seam where you can meter and constrain them. The current spec version is 2026-07-28, released July 28, 2026, and its headline change is a stateless protocol core (MCP versioning).

For boundary design, the useful detail is operational. The release removes the initialize/initialized handshake and the Mcp-Session-Id header, and Streamable HTTP requests must now include Mcp-Method and Mcp-Name headers (SEP-2243), so gateways, rate limiters, and WAFs can route and meter on headers (MCP 2026-07-28 announcement).

That means you can enforce policy in front of the tool server without parsing request bodies. A gateway can allow read_invoice and rate-limit or block send_payment based on headers alone. The boundary lives outside the agent and outside the model, which is exactly where you want it.

Design principles I apply when exposing tools over MCP:

  1. One narrow tool beats one flexible tool. get_invoice(id) is safer than run_query(sql). Narrow tools make the allowed action space small enough to reason about.
  2. Separate read tools from write tools. Give the agent read access by default and require an explicit grant for anything that mutates state.
  3. Make writes idempotent. If the agent retries a step, the second call must be a no-op, not a duplicate.
  4. Meter at the gateway. Rate limits and allowlists belong in infrastructure, not in a system prompt.

What about productivity? Be careful with the headline numbers

You'll see claims that AI coding tools make experienced developers slower. The source for that is METR's July 2025 randomized controlled trial: 16 experienced open-source developers, 246 issues, using mainly Cursor Pro with Claude 3.5/3.7 Sonnet, who took 19% longer (METR). METR itself now labels those results out of date, so I wouldn't use that figure to describe how today's tools perform. Quote it only as a snapshot of early-2025 tooling, and not as a statement about anything current.

I bring it up because it illustrates why boundaries matter more than speed claims. Whether agents make you 20% faster or 20% slower on a given task, the failure mode that hurts a small business isn't slowness. It's the one confident wrong change that reaches production, or one leaked credential. Boundaries address that; productivity debates don't.

A practical boundary checklist for a small codebase

If you're a solo developer or run a team of a few people, here's the order I'd do this in. None of it requires a platform team.

  1. Move secrets out of the agent's reach. Deny reads on .env and secrets directories in your permission rules. Use a secrets manager rather than files where you can.
  2. Turn on the sandbox if your agent runtime supports it, with network access denied by default and an explicit allowlist for the hosts it truly needs.
  3. Add one deterministic hook for your single most expensive mistake (destructive SQL, force-push, deleting a production resource). Start with one. Add more only when a real incident justifies it.
  4. Write contracts for the three fields that drive money or customer-facing numbers. Revenue, tier, status, whatever your business runs on. Assign a human owner to each.
  5. Make every write path idempotent before you let an agent retry anything.
  6. Audit .claude/ and any agent config in repos you clone before trusting them.
  7. Review the diffs on contracted fields every time. The point of the boundary is that your attention goes to the 5% of changes that matter instead of 100% of them.

What I'd skip: elaborate multi-agent review pipelines before you have the basics. A fancy layer on top of an unsandboxed agent with .env readable is decoration.

Where I come at this from

I'm a senior engineer who builds AI automations for solopreneurs and small teams, and the pattern above is how I think about the work: with automations, a large share of the effort goes into defining the guardrails, tools, and boundaries an agent operates inside, not into writing prompts. The goal is to give the agent enough room to be useful and not enough to do real damage if it's wrong.

If you have a codebase and you're unsure where an agent's authority should stop, that's worth working out before an incident rather than after. Reach out through BizFlowAI if you'd like to talk through safe agent boundaries for your own stack.


Work with BizFlowAI

If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.

Request a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.

More guides like this on the BizFlowAI blog.

Frequently asked questions

How do I stop Claude Code from touching sensitive files or running dangerous commands?

Use enforced controls rather than prompt instructions. Instructions in a prompt or CLAUDE.md shape what Claude tries to do but do not change what Claude Code is allowed to do. Add deny rules in your permissions settings (for example blocking reads of .env or secrets directories and risky Bash patterns), and back them with a PreToolUse hook for nuanced checks. Syntax and modes change between versions, so verify against the current docs.

What is the difference between a CLAUDE.md instruction and a permission rule in Claude Code?

A CLAUDE.md instruction is a request the model may or may not follow, while a permission rule is enforced by the tool itself. A line saying 'never touch production' only influences behavior. A deny rule actually blocks the tool call regardless of what the model decides. If you need a guarantee, use permissions, permission modes, or a PreToolUse hook.

How do Claude Code PreToolUse hooks work as a safety layer?

A PreToolUse hook is a script that runs before the permission prompt and receives the tool call details. If it exits with code 2, the tool call is stopped before permission rules are evaluated, so the block holds even if an allow rule would have permitted the call. A hook that returns deny blocks the tool even in bypassPermissions mode. A hook that returns allow cannot override deny rules from settings, so hooks can tighten restrictions but never loosen them.

Why do AI coding agents need a sandbox if I already have permission rules?

Permission rules and hooks decide whether an action is allowed, while a sandbox limits what an action can reach even if those layers fail. Anthropic describes effective sandboxing as needing both filesystem isolation and network isolation. Without network isolation a compromised agent could exfiltrate files like SSH keys, and without filesystem isolation it could escape the sandbox to reach the network. Anthropic reports its OS-level sandbox also cut permission prompts by 84%, reducing approval fatigue.

Can I rely on an AI classifier to approve agent commands automatically?

Not as your only safeguard. Anthropic reports that the classifier behind Claude Code's auto mode blocks about 0.4% of benign commands but lets roughly 17% of overeager actions through. For that reason it should be one layer of defense-in-depth inside a sandbox. Hard boundaries like deny rules, deterministic hooks, and OS-level sandboxing should carry the actual protection.