AI That Touches Machines Needs Different Rules

Abstract tech illustration: AI That Touches Machines Needs Different Rules

Your first wave of AI drafted emails and scored leads, and nothing happened until a human clicked send. The second wave sends the invoice, issues the refund, and updates the CRM on its own. That is the same category of risk industrial teams are now wrestling with, and their safety thinking transfers directly to a five-person business.

What Actually Changed in Industrial AI

Industrial AI is moving from predicting what will happen to acting on physical systems, and that changes what a failure looks like. MIT Technology Review Insights published "Building a safer path to autonomous industrial AI" on October 8, 2026, an interview with Arti Garg, chief technologist at AVEVA.

A fair caveat first: the piece carries an "In partnership with AVEVA" label. AVEVA is an industrial software vendor, so read it as an informed practitioner view, not independent reporting. The ideas hold up on their own, though, and I'm only going to claim what the article supports.

What it says:

  • Industrial AI differs from purely digital AI because it can interact directly with physical systems. An unexpected decision there can affect safety, reliability, and critical infrastructure.
  • Garg names foundation models, physical AI, and agentic AI as the recent shift in the type of AI being applied.
  • She says newer models are harder to explain, and their behavior can change over time as they are tuned.

That last point matters more than it looks. A rule-based script does the same thing on Tuesday as it did on Monday. A tuned model might not. If you can't fully explain a system and its behavior can drift, you can't rely on "it worked in the demo" as your safety argument. You need structure around it.

The Same Shift Is Happening in Your Business

If your agent can send, pay, delete, or book, you are in the same risk category as the factory, minus the physics. The stakes are lower, but the structure is identical: an action in the world that you can't take back.

Here is the difference between the two waves in plain terms:

Wave 1: assist Wave 2: act
Typical tasks Summarize inbox, score leads, draft replies Send invoice, issue refund, update CRM, book appointment
Who executes A human clicks send The agent
Cost of a wrong guess A bad draft you delete An external message, a money movement, a deleted record
Reversible? Almost always Often not

An email to a client can't be unsent. An invoice with the wrong amount creates an accounting and compliance problem. A refund issued twice is real money gone. None of these require a robot arm to hurt you.

The article's core governance idea is to augment people rather than replace them in critical decision loops, with guardrails that define where automated systems may act and where human supervisors stay responsible. You can adopt that sentence as a design spec today.

The Design Lesson: Permitted, Not Just Capable

Software that touches the real world needs three things most demos skip: a permission boundary, a suggest-versus-execute split, and a human checkpoint where the action becomes hard to undo. This is the pattern I recommend.

1. A boundary on what the agent is permitted to do. Not what it is capable of. An LLM with API credentials is capable of almost anything those credentials allow. Permission is a separate, explicit layer you write down.

2. A split between suggesting and executing. The agent does the reading, matching, drafting, and preparation. Executing is a different function with different rules.

3. A checkpoint exactly where the action gets hard to undo. Not everywhere. Approval fatigue kills oversight faster than any bug, so you place the gate precisely.

Here is a minimal sketch of the boundary plus the suggest/execute split:

from dataclasses import dataclass
from enum import Enum

class Risk(Enum):
    FREE = 1        # reversible, low cost
    LOGGED = 2      # reversible, but annoying
    GATED = 3       # irreversible or expensive

# Explicit allowlist: the agent can ONLY do what appears here.
PERMISSIONS = {
    "draft_reply":      Risk.FREE,
    "tag_ticket":       Risk.FREE,
    "update_crm_field": Risk.LOGGED,
    "move_file":        Risk.LOGGED,
    "send_invoice":     Risk.GATED,
    "send_external_email": Risk.GATED,
    "issue_refund":     Risk.GATED,
    "delete_record":    Risk.GATED,
}

@dataclass
class Proposal:
    action: str
    params: dict
    rationale: str   # agent explains itself; you read this when reviewing

def dispatch(p: Proposal, approver, executor, log):
    risk = PERMISSIONS.get(p.action)
    if risk is None:
        log.write("DENIED_UNLISTED", p)      # not permitted, full stop
        return "denied"

    if risk is Risk.GATED and not approver.approve(p):
        log.write("REJECTED_BY_HUMAN", p)
        return "rejected"

    before = executor.snapshot(p) if risk is Risk.LOGGED else None
    result = executor.run(p)
    after = executor.snapshot(p) if risk is Risk.LOGGED else None
    log.write("EXECUTED", p, before=before, after=after, result=result)
    return result

Two details worth copying. An action that isn't on the list is denied by default, not allowed by default. And the agent produces a Proposal, never a side effect; only dispatch executes.

The Three-Bucket Audit You Can Run This Afternoon

Sort every step of a workflow into reversible-cheap, reversible-annoying, or irreversible-expensive, then match autonomy to the bucket. No new tools required. Take one workflow you want to automate and classify each step.

  • Bucket 1: reversible and low-cost. Drafting, tagging, summarizing, sorting. Let the agent run these fully autonomously.
  • Bucket 2: reversible but annoying. Updating records, moving files. The agent does them, but you log every action with before and after values so you can roll back.
  • Bucket 3: irreversible or expensive. Sending money, sending external messages, deleting data. Put an approval step here, even if it's a single tap on your phone.

Then run it that way for two weeks and review the log.

A worked example, purely illustrative. Say your invoicing automation has six steps:

workflow: monthly_client_invoicing
steps:
  - name: pull_hours_from_tracker      # bucket 1: read-only
    autonomy: full
  - name: match_hours_to_contract      # bucket 1: pure computation
    autonomy: full
  - name: draft_invoice_pdf            # bucket 1: nothing leaves your system
    autonomy: full
  - name: write_invoice_to_accounting  # bucket 2: reversible, log before/after
    autonomy: full_with_log
  - name: send_invoice_to_client       # bucket 3: external, can't unsend
    autonomy: approval_required
  - name: apply_late_fee               # bucket 3: money + client relationship
    autonomy: approval_required

Four of six steps run with no human involvement. The agent still does nearly all the work; your part is a quick approval on the two steps that can embarrass you, which should take a few seconds each (an estimate, not a measured figure). That ratio is the point. Gates cost little when they sit in the right places.

Reviewing the log

After two weeks, read your approval log with one question: what did I reject, and why?

  • If you rejected a bucket-3 action, the gate earned its keep. Find out what the agent got wrong and fix the cause.
  • If you approved a bucket-3 action every time, with no edits and no hesitation, it becomes a candidate to loosen.
  • If your log is thin, you don't have enough data yet. Keep the gate.

That is earned autonomy: your own real numbers instead of a vendor's promise. Treat loosening as a per-action decision, not a global setting. "Send invoices under a fixed amount to clients with a clean payment history" is a defensible rule. "Trust the agent now" is not.

Why Logging and Rollback Matter More Than the Model

The model will occasionally be wrong; what decides the damage is whether you can see what it did and undo it. Given Garg's point that newer models are harder to explain and can change as they are tuned, you shouldn't plan around predicting the model's behavior. Plan around observing it.

A minimum viable log entry for any bucket-2 or bucket-3 action:

{
  "ts": "2026-10-08T14:02:11Z",
  "agent": "invoicing-agent",
  "action": "write_invoice_to_accounting",
  "rationale": "Hours matched contract line 2; total differs from last month by +12%",
  "before": {"invoice_id": null},
  "after":  {"invoice_id": "INV-0412", "total": 1840.00},
  "approved_by": null,
  "result": "ok"
}

Three habits that pay off:

  • Store the rationale. When something goes wrong, "why did it do that" is the first question.
  • Capture before and after values. Rollback without a before-state is guesswork.
  • Pin your model and prompt versions in the log. If behavior shifts after a tuning change or provider update, you want to see the day it started.

If you're curious how formal security bodies are approaching the same problem, NIST has been working on standards and guidance for AI agents, including the Cloud Security Alliance's research note on NIST's AI agent standards effort. The status of these efforts changes, so check NIST directly for current documents before citing them. The goals, though, line up with the permission boundary and containment idea above.

A Note on Regulation (Check Your Jurisdiction)

If you actually control physical machinery, regulation exists, but whether it applies to you depends on where you are and what you're doing. I won't pretend to give legal advice, and several of these points are unsettled.

For readers in the EU, a few points worth knowing:

  • The EU Machinery Regulation (EU 2023/1230) replaces the earlier Machinery Directive. Its application date and transitional rules are covered in this overview, and you should confirm the dates in the official text.
  • It explicitly addresses machine learning in safety functions, and cybersecurity (protection against accidental and intentional corruption of safety software and data) becomes an essential requirement.
  • It treats certain digital changes that create a new hazard or increase an existing risk as substantial modifications, which can shift manufacturer responsibilities and trigger a new conformity assessment for the affected aspects. If you retrofit AI control logic onto an existing machine, read that carefully, and see this explainer for context.
  • The EU AI Act's high-risk timelines were adjusted by the Digital Omnibus amendments. Dates differ between stand-alone high-risk systems and AI embedded in regulated products such as machinery. See Morgan Lewis's summary and Inside Global Tech's update, and verify current dates against the official texts.

What I could not confirm: whether a small business that only uses an AI agent on its own existing machine is directly in scope, and exactly how the Omnibus changed the machinery treatment (secondary sources describe it differently). Check the official texts or ask a compliance professional. If you operate in the US, confirm directly with OSHA or a safety professional what applies to you before relying on any assumption.

On the robotics side, ISO 10218 was revised in 2025, with updated functional safety and cybersecurity content, and the collaborative-robot material that previously lived in a separate technical specification was folded in. See ISO's catalogue entry and this FAQ on the update. These are standards, not automatically law; whether one is legally binding depends on your jurisdiction.

If your agents only touch software, none of this binds you directly, but it's a useful preview of where scrutiny of autonomous action is heading.

Putting the Pattern to Work

The approach I recommend is simple: the agent handles reading, matching, drafting, and preparation, while irreversible steps like sending an invoice or an external message sit behind an explicit rule or a one-tap approval. I'd suggest logging every action that changes something with before-and-after values, so rollback is a lookup, not a forensic project. And I'd expand autonomy step by step, based on what your own approval log shows after a couple of weeks of real use.

My hot take: the winners over the next couple of years won't be the teams with the most autonomous agents. They'll be the teams whose agents have the best brakes. Autonomy is easy to demo. Trustworthy autonomy is an engineering discipline, and it's boring on purpose: logging, permissions, rollbacks, approval gates. If a vendor pitches you an agent and can't tell you what it's not allowed to do, walk away.


Want more like this?

I publish practical AI automation, GenAI engineering, and faceless content workflows on YouTube every week.

Subscribe to bizflowai.io on YouTube — never miss a new tutorial.

Planning an AI automation project or need a second opinion on your architecture?

Connect with me on LinkedIn — Lazar Milicevic, GenAI Engineer & bizflowai.io Founder.

Visit bizflowai.io for our services, case studies, and AI consulting.

Frequently asked questions

What is the difference between traditional industrial AI and the new phase of industrial AI?

Industrial AI has spent decades in predictive analytics and narrow, specialized applications, such as models that forecast when a machine will fail. The new phase is driven by advances in foundation models, physical AI, and agentic AI, which make it possible to automate more complex tasks across industrial environments. Unlike purely digital AI, industrial AI can interact directly with physical systems, which changes what failure looks like.

Why does industrial AI matter for software builders and founders who don't own a factory?

The same shift is happening in business software. The first wave of AI was analysis and drafting, where nothing happens until a human clicks send. The second wave is agents that act: sending invoices, updating CRMs, issuing refunds, booking appointments. These actions aren't physical, but they are irreversible. An email can't be unsent, a wrong invoice creates compliance problems, and a double refund is real money lost.

What safeguards do AI agents need before they act in the real world?

Software that touches the real world needs three things demos usually skip. First, a clear boundary on what the agent is permitted to do, not just what it is capable of doing. Second, a distinction between suggesting and executing. Third, a human checkpoint placed exactly where an action becomes hard to undo. Typically the agent handles reading, matching, drafting, and preparation, while the irreversible step sits behind a rule or approval.

How do I decide which AI agent actions need human approval?

Sort each workflow step into three buckets. Reversible, low-cost steps like drafting, tagging, summarizing, and sorting can run fully autonomously. Reversible but annoying steps like updating records or moving files should run with logs of before and after values for rollback. Irreversible or expensive steps like sending money, sending external messages, or deleting data should require an approval, even a single tap on a phone.

How do I decide when to give an AI agent more autonomy?

Run the agent with approval steps on irreversible actions for about two weeks, then review the log. Any action in the irreversible bucket that is never rejected becomes a candidate for loosening the approval requirement. This is earned autonomy, based on your own real numbers rather than a vendor's promises. It requires no new tools and takes roughly an afternoon to set up.