Agent Infrastructure Gaps: Talk, Trust, Audit

You have an agent that drafts invoices, another that triages email, and a third that updates your CRM. Now you want them to hand work to each other, and you realize nobody can tell you what each one is allowed to touch or what it did last Tuesday. That gap is real, and at VB Transform 2026 (July 14-15, Menlo Park), several startups showed up claiming to fill it.
This post breaks the problem into three parts (communication, permissions, audit), looks at what the startups are actually building, and then gives you a practical stack you can assemble as a small team without a platform engineering department.
The three gaps, stated plainly
The gap is not capability. It is plumbing. Agents can already do the work. What is missing is the infrastructure that lets them (1) exchange messages reliably, (2) act only within narrowly granted permissions, and (3) leave a trail you can read after something goes wrong.
VentureBeat's July 29, 2026 write-up of the VB Transform startups frames it the same way: orchestration, observability, connectivity, and security are the four areas where early infrastructure is being built.
Here is how the three gaps show up for a small business:
| Gap | What it looks like in a 5-person company | What breaks |
|---|---|---|
| Communication | Your lead-intake agent and your quoting agent run in different tools and cannot pass context | You copy-paste between them, or each one re-asks the customer |
| Permissions | The agent uses your personal API key or a shared service account | One bad prompt or injected instruction has the same access you do |
| Audit | Logs live in three vendor dashboards, or nowhere | "Why did it email that client?" has no answer |
The permissions one is the most dangerous. A 2026 VentureBeat survey, cited in the VB Transform agenda, reportedly found that 88% of enterprises had experienced an AI agent security incident, and the agenda says the common cause was a legitimate agent using a valid but over-scoped credential. Treat that figure as a conference-agenda claim: I could not find the survey's sample size or methodology. The pattern it describes (valid credential, too much scope) is the one I see in small-business setups constantly.
Gap 1: Agents that cannot talk to each other
Two protocols matter here, and they solve different problems. MCP (Model Context Protocol) connects an agent to tools and data. A2A (Agent2Agent) is for communication and interoperability between agent systems. They are complementary, not competing.
A quick orientation:
- MCP: agent to tool. "Read this spreadsheet," "create this invoice."
- A2A: agent to agent. "I've qualified this lead; you take the quote." A2A is an open protocol hosted by the Linux Foundation and originally contributed by Google. In April 2026 the Linux Foundation said the project had passed 150 supporting organizations at its one-year mark.
- Governance: Anthropic donated MCP to the Agentic AI Foundation at the Linux Foundation on 2025-12-09, co-founded by Anthropic, Block and OpenAI (per a third-party protocol timeline citing Anthropic's announcement).
What BAND is doing
BAND (also known as Thenvoi AI Ltd.) is the startup VentureBeat highlights for orchestration. It exited stealth with a $17 million Seed round, and VentureBeat describes it as a deterministic communication layer, a "Slack for agents." Calcalist reports the round was led by Sierra Ventures, Hetz Ventures and Team8, and that the company was founded in mid-2025 by Arick Goomanovsky (CEO) and Vlad Luzin (CTO).
Two design choices are worth understanding even if you never buy the product:
- No LLM in the routing path. VentureBeat says BAND uses a patent-pending multi-layer architecture rather than an LLM to route messages, because LLM routing would introduce non-deterministic errors. This is the right instinct. If the thing deciding who gets which message can hallucinate, you cannot reason about your system.
- Long-running workflows. CTO Vlad Luzin says BAND supports autonomous workflows that run for eight to 20 hours, and that it is compatible with both A2A and MCP.
The lesson for a small team: keep message routing boring and deterministic, and let the LLM do the thinking inside each agent. You can apply that without buying anything.
A deterministic router in about 30 lines
This is a sketch, not a product. The point is that routing is a lookup table, not a model call.
# router.py - deterministic agent-to-agent routing (illustrative)
from dataclasses import dataclass
from typing import Callable
import json, time, uuid
@dataclass
class Message:
id: str
sender: str
topic: str # e.g. "lead.qualified"
payload: dict
ts: float
# Explicit routing table. No LLM decides who receives what.
ROUTES: dict[str, list[str]] = {
"lead.qualified": ["quote_agent"],
"quote.sent": ["crm_agent", "followup_agent"],
"invoice.overdue": ["collections_agent"],
}
HANDLERS: dict[str, Callable[[Message], None]] = {}
def register(agent: str, fn: Callable[[Message], None]):
HANDLERS[agent] = fn
def publish(sender: str, topic: str, payload: dict):
msg = Message(str(uuid.uuid4()), sender, topic, payload, time.time())
recipients = ROUTES.get(topic, [])
# Append-only log BEFORE delivery, so audit exists even if a handler crashes
with open("agent_messages.jsonl", "a") as f:
f.write(json.dumps({**msg.__dict__, "recipients": recipients}) + "\n")
for r in recipients:
HANDLERS[r](msg)
Two things to notice. Unknown topics go nowhere (fail closed), and every message is written to an append-only log before any handler runs. That second property is your audit trail starting for free.
Gap 2: Agents you cannot safely give permissions to
The core rule: an agent should hold a credential scoped to one task, not a copy of your access. Most small-business agent setups fail this. The agent runs as "you," or as a shared service account, and anything it can be talked into doing, it can do.
What Arcade is doing
Arcade.dev says its secure agent runtime provides an authentication and authorization layer plus observability, with actions attributable and issued under least-privilege scopes. It can be deployed on-prem as an installable plugin, and companies keep using their own sign-in and security tooling: whatever runs in Arcade is gated by existing RBAC, identity provider policies and entitlements.
The design principle matters more than the vendor: the agent platform should consume your existing identity and access rules, not invent a parallel set. If your permissions live in two places, they will drift.
(I have seen one aggregator quote Arcade funding figures, but I could not confirm them in a primary source, so I am leaving them out.)
What MCP itself does and does not give you
Do not assume adopting MCP solves permissions. The NSA's May 2026 MCP security guidance (Version 1.0) says authorization in MCP is optional, and that MCP currently lacks support for exchanging RBAC permissions at instantiation. In plain terms: the protocol gives you hooks, but enforcement is your job. The guidance is published as a CSI on MCP security.
The current MCP specification, version 2026-07-28, tightens several things relevant here:
- A stateless protocol core, Multi Round-Trip Requests, authorization hardening, and a formal extensions framework.
- Issuer validation: clients must validate the RFC 9207
issparameter before redeeming an authorization code (SEP-2468). - Audience-bound tokens: per Cloudflare's write-up, clients send the canonical server URI as the RFC 8707
resource, and tokens must be issued for, and accepted only by, that audience. A token minted for one MCP server is useless against another. - Client registration: Dynamic Client Registration is formally deprecated in favor of Client ID Metadata Documents (CIMD). The official MCP blog says DCR will be removed "in a future version"; I'd go with that wording rather than a specific date.
Sources: the official MCP 2026-07-28 spec announcement and Cloudflare's next generation of MCP post.
A practical permission policy for a small team
You do not need an enterprise IAM program. You need a per-agent policy file that a tiny gateway enforces. Illustrative example:
# agent_policies.yaml - one entry per agent, deny by default
agents:
invoice_agent:
credential: svc-invoice-agent # its own identity, not yours
allow:
- tool: accounting.invoice.create
- tool: accounting.invoice.read
deny_all_else: true
limits:
max_invoice_amount_usd: 5000 # example threshold
requires_human_approval_above_usd: 1000
email_triage_agent:
credential: svc-email-triage
allow:
- tool: mail.read
- tool: mail.label
- tool: mail.draft # drafts only
deny_all_else: true # note: no mail.send
The key decisions in that file: each agent has its own identity (so logs are attributable), the email agent can draft but not send, and money-moving actions have an approval threshold. The dollar figures are examples; set yours from your own risk tolerance.
Gap 3: Agents you cannot audit
An audit trail answers four questions: who acted, on whose behalf, with what authority, and what changed. If your logs cannot answer all four for a given action, you do not have an audit trail; you have debug output.
The ingredients:
- Per-agent identity (from the permissions section) so "who" is never "the service account."
- A correlation ID that follows a task across agents, so you can reconstruct a chain like lead, quote, CRM update.
- Tool-call records, not just chat transcripts: tool name, arguments, result status.
- Append-only storage that the agents themselves cannot edit.
A minimal record format:
{
"ts": "2026-10-09T14:22:07Z",
"correlation_id": "c9a1-4f20",
"agent": "invoice_agent",
"on_behalf_of": "owner@example.com",
"tool": "accounting.invoice.create",
"args_hash": "sha256:ab12...",
"decision": "allowed",
"policy_rule": "invoice_agent.allow[0]",
"approval": {"required": false},
"result": "ok"
}
Log a hash of sensitive arguments rather than the raw values if the log itself could leak customer data, and keep the raw payload in a separate, access-controlled store.
Security tooling you already own
Conifers is the third startup in VentureBeat's piece for which I could verify details: it connects to a company's existing security tools such as EDR and SIEM. Again, the pattern is the useful part. If you already have any log aggregation, even a basic one, send agent audit records into it rather than standing up a separate "AI audit" island. Alerts, retention, and access control then come along for free.
A note on the other two startups from the article: the text I could verify only covers BAND, Arcade and Conifers, so I am not going to guess at the other two. Read the full VentureBeat article for the complete list.
Putting it together: a reference architecture at SMB scale
You can get most of the value with four components and no enterprise budget. Here is the shape:
[Agents] --MCP--> [Policy gateway] --> [Tools: CRM, mail, accounting]
| |
| +--> [Append-only audit log] --> [Your existing log/alert tool]
|
+--(topic messages)--> [Deterministic router] --> [Other agents]
Build order I recommend, in priority sequence:
- Per-agent credentials first. This is the cheapest control and the one that limits blast radius. Do this before anything else.
- Policy gateway with deny-by-default. Even a thin proxy that checks
agent + toolagainst a YAML file beats trusting each agent's prompt to behave. - Audit log with correlation IDs. Start writing it on day one; you cannot retroactively create history.
- Deterministic routing between agents. Only needed once you actually have more than one agent handing work to another.
- Human approval gates on anything irreversible or money-moving.
Where MCP fits and where it does not
MCP is a good standard for the tool-access leg. With the 2026-07-28 spec's audience-bound tokens and issuer validation, you get meaningfully better token hygiene than earlier versions. But per the NSA guidance, authorization is optional and RBAC exchange at instantiation is not supported, so your gateway and policy file carry the real enforcement burden. Use MCP as the connector, not as the security model.
A note on regulation
If you sell into the EU, timelines recently shifted. The EU Digital Omnibus on AI is Regulation (EU) 2026/1744, published in the Official Journal on July 24, 2026 and in force from July 27, 2026. It moves Annex III high-risk obligations to December 2, 2027 and Annex I obligations to August 2, 2028. Deferred is not cancelled, and audit trails are the kind of evidence these regimes tend to ask for. This is not legal advice; check the regulation text or a qualified professional for what applies to your business. The Cloud Security Alliance research note on the high-risk deadline is a useful starting point.
What to do this week
- Inventory your agents and write down, for each, which credential it uses. If the answer is "mine," that is your first fix.
- Split credentials so each agent has its own identity with the narrowest scopes that work.
- Turn on append-only logging with a correlation ID before adding any new agent.
- Remove send/delete/pay permissions from any agent that can be talked into things by external text (inbound email, web pages, uploaded docs). Draft-only is a fine default.
- Revisit when you cross two agents. The moment one agent hands off to another, add deterministic routing and shared correlation IDs.
The startups at VB Transform are building the enterprise-grade versions of these controls. The principles (deterministic routing, least privilege, attributable actions, logs that plug into what you already run) scale down just fine.
A practical starting point
The three gaps above are the ones to design around when you build agent systems for a small team: a separate identity for each agent, a policy check in front of the tools, an append-only audit log, and human approval on anything that moves money or sends external messages. That is my recommendation as a builder, not a description of a packaged product. MCP works well as the connector layer for tool access, with the enforcement logic living in your own gateway rather than being delegated to the protocol, for the reasons in the NSA guidance above.
Work with BizFlowAI
If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.
Request a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.
More guides like this on the BizFlowAI blog.
Frequently asked questions
What is the difference between MCP and A2A for AI agents?
MCP (Model Context Protocol) connects an agent to tools and data, such as reading a spreadsheet or creating an invoice. A2A (Agent2Agent) handles communication and interoperability between separate agent systems, such as one agent handing a qualified lead to another. They solve different problems and are complementary, not competing. Most multi-agent setups end up using both.
How do I give an AI agent safe permissions without giving it my own access?
Give each agent its own credential scoped to a single task instead of running it under your personal API key or a shared service account. Use deny-by-default policies that list exactly which tools the agent may call, and add limits such as maximum transaction amounts and human approval thresholds. Where possible, have the agent platform consume your existing identity provider and RBAC rules so permissions are not defined in two places that can drift apart.
Does adopting MCP automatically secure my AI agents?
No. The NSA's May 2026 MCP security guidance notes that authorization in MCP is optional and that the protocol lacks support for exchanging RBAC permissions at instantiation. MCP gives you hooks, such as audience-bound tokens and issuer validation in the 2026-07-28 spec, but enforcement is your responsibility. You still need a gateway or policy layer that restricts what each agent can do.
How can I keep an audit trail of what my AI agents did?
Write every inter-agent message and tool call to an append-only log before delivery or execution, so a record exists even if a handler crashes. Include the sender, topic, payload, recipients, and timestamp in each entry. Centralize these logs in one place rather than scattering them across vendor dashboards, so you can answer questions like why an agent emailed a specific client.
Should an LLM decide how messages are routed between agents?
Generally no. If an LLM decides which agent receives which message, routing can become non-deterministic and hard to reason about or debug. A safer pattern is an explicit routing table that maps message topics to recipient agents and fails closed for unknown topics. Let the LLM do the reasoning inside each agent, while the routing layer stays simple and predictable.