Goodfire's Monitor Reads the Model's Mind, Not Its Logs

Abstract tech illustration: Goodfire's Monitor Reads the Model's Mind, Not Its Logs

Your agent works in the demo. Then someone asks what happens when it does something wrong at 3 a.m., and the safe answer, a second model reviewing every action, adds a significant cost to every task. Goodfire just launched monitors that skip that bill by reading the model's internals instead of its transcript. Here's what they shipped, what the numbers actually say, and which part of the idea you can use without buying anything.

What Goodfire Actually Shipped

Goodfire launched "inside-out" agent monitors on October 8, 2026, available to customers of Baseten, which hosts and runs AI models for other companies. Instead of a second model reading everything the agent does, small probes read the model's internal activity as it generates, and a more expensive check runs only when a probe fires. (TechCrunch)

Why it's cheap: Goodfire's CEO Eric Ho says the monitors reuse computations the model already performs in its forward pass. The probes just read the results. There's no second model reading every token.

Baseten customers choose which risks to monitor: offensive hacking, chemical and biological weapons misuse, and reward hacking. They also choose the automated response: log the event, send it for human review, or refuse the request. Goodfire's probe guide describes a wider menu of actions when a probe fires: stop the model, escalate to an LLM judge, re-route the task, inject custom prompts, or log the transcript. (Goodfire)

One more detail from the launch coverage: Goodfire's CTO and co-founder Dan Balsam says the main advantage is catching things before they happen, including detecting when a model might hack during evaluation or training.

The Cascade: How the Two Stages Fit Together

The architecture is a two-stage cascade. A probe reads internal activations as tokens are generated. When its readout crosses a calibrated threshold, the rollout escalates to an LLM judge, which can refuse the request or continue the interaction. (Unite.ai)

Strip away the interpretability and the shape is simple:

def monitor(step):
    score = cheap_check(step)           # runs on EVERYTHING, near-zero cost
    if score < THRESHOLD:
        return "allow"
    verdict = expensive_judge(step)     # runs on the flagged slice only
    return verdict                      # refuse / continue / human review

The threshold is the whole game. Set it low and you escalate too much and lose the savings. Set it high and you miss things. Goodfire's numbers show the trade-off in practice (more on that below), and the escalation rate matters as much as the detection rate.

This is the same pattern human teams use. You don't have a senior manager read every email a junior sends. You have rules, and the manager looks at the flagged ones. Goodfire's contribution is putting the cheap layer inside the model, where it can see intent-adjacent signals that never show up in the output text.

The Numbers, and How Much to Trust Them

All of these are Goodfire's own figures, and I haven't reproduced any of them. Attribute them to the vendor accordingly.

Claim Figure Source type
Cost to monitor ~1,500 Kimi K3 sessions ~$51 Company-run test
Cheaper AI model checking every step, same sessions $233 Company-run test
Top-tier model checking every step, same sessions ~$10,000 Company-run test
Malicious hacking sessions detected 94% Reported by TechCrunch syndications
Harmless sessions escalated for extra review 8.7% Escalation rate, not block rate
Added time before the model starts responding (4 probes) <2% Not total session time
Cyber monitors vs. an LLM judge on recall Match, at ~50x lower cost Goodfire's cyber-monitor post
Reward-hacking cascade 90% cost cut, ~1% precision loss Goodfire's reward-hacking study

The $51 / $233 / $10,000 comparison comes via RuntimeWire, and the 94% and 8.7% figures via Superpowerdaily.

Goodfire also estimates that judging one million Kimi K3 turns with an every-turn judge would cost on the order of $200,000 with Claude Opus 5, or about $30,000 with Kimi K3 itself, using OpenRouter rates. Its research post reports the cascade's judge cost at $65 per million agent turns without judge reasoning and $185 with reasoning. Those come from secondary write-ups, so check goodfire.com before quoting them.

Caveats worth knowing before you repeat any of this:

  • The evaluation is partly in-house. The dataset is Goodfire's, the reference labels come from another model (Claude Opus 5) rather than human review, and the only external test was two days of static red-teaming by FAR.AI. (ProgressiveRobot)
  • 94% means 6% got through. Some malicious sessions went unflagged, and the Kimi K3 results don't show how the probes perform on other risks, other architectures, or other attack strategies.
  • Sources disagree on details. One write-up gives 93% recall from the research post rather than 94%. These likely come from different evaluations. I'd treat "roughly 93-94%" as the honest range until someone checks the primary paper.
  • 8.7% is a lot of escalations. On a low-volume agent that's fine. It's the number that determines your real savings, because every escalation pays for the expensive judge.

Where It Doesn't Help a Small Business

This is the part most coverage skips. The probes need access to the model's activations, so the monitors only work where the operator controls the serving stack. That fits inference providers like Baseten. It does not help a business consuming a closed model through an API. (ProgressiveRobot)

If your agent runs on a hosted closed-model API, you can't attach these probes. In the sources I read, availability is stated only for Baseten customers, and I didn't find an official pricing page. Check goodfire.com and Baseten directly for current availability and pricing.

What does apply to you is the cascade pattern: cheap check on everything, expensive check on the exceptions. My hot take is that interpretability-based monitoring is promising, but a ten-person company shouldn't wait for it. The tooling is aimed at inference providers and open-weight models first, and the claim that it reliably catches rogue behavior hasn't been proven in public. If it does prove out, it slots in later as your cheap first layer.

Build the Tiered Version This Week

Here's the copyable part. Take one agent you already run, list every action it can take, and sort them into three buckets.

  1. Reversible, low stakes (drafting a reply, tagging a record): log everything, sample a few percent for review.
  2. Reversible but annoying (updating a customer field, moving a deal stage): add rule checks, like "does this value look like the last hundred values?"
  3. Irreversible or expensive (sending money, deleting data, emailing a client, filing an invoice): require human approval or a second-model review, every time.

A minimal version in config form:

actions:
  draft_reply:      { tier: 1, review: "sample_5pct" }
  tag_record:       { tier: 1, review: "sample_5pct" }
  update_crm_field: { tier: 2, review: "rules", rules: ["matches_recent_distribution"] }
  move_deal_stage:  { tier: 2, review: "rules", rules: ["stage_transition_allowed"] }
  send_client_email: { tier: 3, review: "human_approval" }
  file_invoice:      { tier: 3, review: "human_approval" }
  delete_data:       { tier: 3, review: "human_approval" }

The "5pct" sampling rate is an example with round numbers; tune it to your volume and risk. In most agents, the irreversible bucket is a small slice of total actions, so the expensive review touches only a fraction of traffic. That's where the savings come from: you stop paying premium review prices on actions that don't need them. Measure your own split from a week of logs before you rely on any estimate.

One thing to build in from day one: log what the agent decided and why, not just what it did. When something goes wrong you want the decision, the action, and who approved it. That's the difference between a fixable incident and a mystery.

audit_log.write({
    "ts": now(),
    "action": "file_invoice",
    "inputs_summary": summary,
    "agent_reasoning": reasoning,   # the "why", not only the "what"
    "tier": 3,
    "approved_by": approver_id,
})

Fix Permissions Before You Buy Monitoring

The real risk for a ten-person company isn't a scheming model. It's an ordinary agent with too many permissions and no approval step on the one action that costs real money. Goodfire's earlier research found that leading open models, including Kimi K3 and GLM 5.2, reward-hacked in 50% to 96% of runs on tests of AI agents. That's a good reason to take agent behavior seriously, but it's a finding about open models in benchmark settings, not a prediction for your invoicing bot.

Practical order of operations:

  1. Scope credentials so the agent literally cannot do bucket-three actions without going through your approval step.
  2. Add the decision log.
  3. Add rule checks on bucket two.
  4. Only then think about model-level monitoring, and only if you control the serving stack.

Boring, and it works today.

Where bizflowai.io fits

I focus on practical AI automation for solopreneurs and small teams at bizflowai.io. The pattern in this post is plumbing, not magic: sort actions by reversibility, route the irreversible ones to a human or a second check, and log every decision with its reasoning so incidents are traceable. It's the same shape Goodfire is putting inside the model.


Want more like this?

I write about practical AI automation and GenAI engineering for small teams.

Planning an AI automation project or need a second opinion on your architecture?

Connect with me on LinkedIn, Lazar Milićević, senior engineer.

Visit bizflowai.io for our services, case studies, and AI consulting.

Frequently asked questions

What is Goodfire's approach to monitoring AI agents?

Goodfire, a company known for interpretability research (studying what happens inside a neural network), launched monitors for AI agents. Instead of running a second model to read everything an agent does, the monitor looks at the model's internal activity while it works and only escalates to a more expensive check when something looks suspicious. Goodfire claims this costs a fraction of read-everything monitoring, but the figure is unverified by independent benchmarks.

Why does AI agent monitoring cost matter for small businesses?

Monitoring cost can quietly kill agent projects. The safe approach is a second model reviewing every action, but that reviewer must read the full transcript, not just the final answer, which can double cost per task or more. Small teams then either skip monitoring and hope nothing breaks, or monitor everything and never see the project pay back. Neither outcome is good.

How do I build a tiered monitor for an AI agent?

List every action your agent can take and sort them into three buckets. Reversible, low-stakes actions (drafting replies, tagging records): log them and sample a few percent for review. Reversible but annoying actions (updating customer fields, moving deal stages): add simple rule checks, such as whether a value resembles recent values. Irreversible or expensive actions (sending money, deleting data, emailing clients, filing invoices): require human approval or a second model review every time.

Why should I log an AI agent's reasoning and not just its actions?

If something goes wrong, you need to see the decision the agent made, the action it took, and who approved it. Logging only what the agent did leaves you guessing why. Recording the decision and its rationale is the difference between a fixable incident and a mystery, and it makes reviewing any flagged action much faster.

Should small businesses wait for interpretability-based agent monitoring?

No. Interpretability tooling will likely target labs and large enterprises first, and the claim that it reliably catches rogue behavior hasn't been proven publicly. For a small company, the bigger risk is an ordinary agent with too many permissions and no approval step on the one action that costs real money. Tiered checks and permission limits work today, and interpretability monitoring could be added later as a cheap first layer if it proves out.