Only 2% Picked an Agent Platform for Its Model

Your agent prototype works. Now someone asks where it runs, who can stop it, and what happens when it loops overnight on a pay-per-token API. The model vendor's platform is the obvious default, but a recent enterprise survey suggests the buyers who've been through this are choosing on other criteria entirely.
This post walks through what that survey found, what transfers to a 1-10 person business, and how to build an orchestration layer you can govern without a platform team.
What the survey actually says (and what it doesn't)
Among enterprises running an orchestration platform, only 2% said alignment with a leading AI model most influenced their choice. That's 3 of 162 platform users, according to VentureBeat's August VB Pulse.
What drove choices instead:
| Selection factor | Share of choices |
|---|---|
| Flexibility across models and tools | 22% |
| Production reliability | 20% |
| Ease of development | 19% |
| Control over agent execution | 18% |
| Security and permissions | 10% |
| Total cost of ownership | 9% |
The first four add up to 78%. Model alignment barely registers.
Three caveats before you quote any of this:
- It's not a probability sample. VentureBeat surveyed a self-selected audience of its own readers: 221 responses, 169 passed the qualifying questions, 166 remained after removing three contradictory respondents.
- It covers organizations with 100+ employees. Most small businesses are below that cutoff. Treat the percentages as a signal about how larger buyers think, not as a statistic about your peers.
- Waves aren't comparable month to month. Different respondents answered each month, and the sample composition shifted: technology and software companies were 53% of July respondents and 10% of August respondents. Earlier waves showed different numbers, including a higher "model gravity" share. Don't build a trend line out of them.
Also note what the data does not say: cost was only 9% as a selection factor. It wasn't a leading reason people chose a platform. Where cost shows up is later, as a control problem, which is covered below. So the honest framing is: model loyalty is a weak reason to pick a platform, and the things that matter in practice are flexibility, reliability, and control.
Why the model is the wrong thing to anchor on
Models change faster than infrastructure. If your orchestration layer is welded to one vendor's agent runtime, every model swap becomes a migration project.
The survey's own numbers point the same way. When asked about the top risk of keeping orchestration control inside a model provider's platform, respondents ranked:
- Inflexibility across models and tools: 28%
- Security and permissioning limitations: 22%
- Limited visibility and observability: 22%
- Vendor lock-in: 16%
Notice that "lock-in" itself is last. The bigger worry is practical: you can't route work to a better or cheaper model, and you can't see what the agent did.
The market is also not settled. According to the same survey, 60% of respondents plan to adopt a new, additional, or replacement orchestration platform within 12 months, and 28% expect to do so within three months. Among multi-platform enterprises, 84% plan a change versus 46% of single-platform ones. Whatever you pick, plan for it to be replaceable.
For a small team, the takeaway is simple: keep the part you'll want to change (the model) separate from the part you need to keep stable (triggers, permissions, logging, budgets).
Where the control plane is heading
The survey found that 67% of respondents expect the agent control plane to sit at least partly outside a provider-managed service by the end of 2026. Only 27% expect a provider-managed agent service. "Hybrid" was the most common single answer at 33%.
In plain terms: buyers want the model vendor to supply the model, and want to own the thing that decides what the agent may do, what it may spend, and how it gets audited.
That's a useful design rule even at one-person scale. Split your system into three layers:
┌──────────────────────────────────────────┐
│ Control plane (you own this) │
│ triggers · approvals · budgets · logs │
├──────────────────────────────────────────┤
│ Tool layer (standardized) │
│ MCP servers · API connectors · webhooks │
├──────────────────────────────────────────┤
│ Model layer (swappable) │
│ Claude · GPT · open-weight · routed │
└──────────────────────────────────────────┘
A managed vendor platform can be a fine choice for the model layer or even parts of the runtime. The test is whether you can still enforce your own limits and read your own logs when you swap something underneath.
The tool layer: why MCP matters for portability
Model Context Protocol (MCP) is the closest thing to a standard for connecting agents to tools. Anthropic announced it was donating MCP to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded by Anthropic, Block and OpenAI, with MCP's maintainers and governance model unchanged (Anthropic announcement; the MCP project post is dated 2025-12-09).
Why that matters to you: if your CRM lookup, invoice fetch, and calendar actions are exposed as MCP servers (or plain HTTP endpoints you control), any orchestrator and any model can call them. You write the integration once. The agent runtime becomes the replaceable part.
A minimal sketch of the principle, a tool wrapped behind your own gateway rather than called directly from the agent:
# tool_gateway.py - every agent tool call passes through here
import time, json
ALLOWED_TOOLS = {"crm.lookup", "invoice.draft"} # no send, no delete
READ_ONLY = {"crm.lookup"}
def call_tool(agent_id: str, tool: str, args: dict, run_budget_left: float):
if tool not in ALLOWED_TOOLS:
raise PermissionError(f"{tool} not permitted for {agent_id}")
if run_budget_left <= 0:
raise RuntimeError("budget exhausted - halting run")
started = time.time()
result = TOOLS[tool](**args) # your real implementation
log = {"agent": agent_id, "tool": tool, "args": args,
"ms": int((time.time() - started) * 1000),
"write": tool not in READ_ONLY}
print(json.dumps(log)) # ship to your log store
return result
It's deliberately boring. The value is that permissions, budget checks, and logging live in one place you own, regardless of which model is calling.
The cost problem nobody budgets for
Cost didn't drive platform selection (9%), but the survey shows it's where enterprises are exposed afterward. Among enterprises using agent orchestration:
- 23% track agent spending only through after-the-fact logs, so they have no real-time way to stop an agent before it exceeds its budget.
- 28% built their own gateways or middleware.
- 20% route token-heavy work to low-cost models.
Those are the three real strategies: watch after the fact (weakest), enforce in a gateway, or route by cost. A small business should skip the first one.
A hard per-run cap, not a dashboard
A dashboard tells you about a loop after it happened. A cap stops it. The minimum viable version:
# budget_guard.py
class BudgetExceeded(Exception): ...
class RunBudget:
def __init__(self, max_usd: float, max_steps: int):
self.max_usd, self.max_steps = max_usd, max_steps
self.spent, self.steps = 0.0, 0
def charge(self, input_tokens: int, output_tokens: int,
usd_per_mtok_in: float, usd_per_mtok_out: float):
# prices come from your config - check the vendor's current pricing page
self.spent += (input_tokens / 1e6) * usd_per_mtok_in
self.spent += (output_tokens / 1e6) * usd_per_mtok_out
self.steps += 1
if self.spent > self.max_usd or self.steps > self.max_steps:
raise BudgetExceeded(f"${self.spent:.2f} / {self.steps} steps")
Put the per-token prices in config, not code. Vendor prices change, and published figures for current models have varied between sources, so verify against the vendor's pricing page before you rely on any number.
Understand the billing shape of your platform
Different platforms meter differently, and that changes your risk profile:
- Runtime-metered agent services. Anthropic's pricing page lists Claude Managed Agents at standard token rates plus $0.08 per session-hour of active runtime (pricing). One secondary analysis notes runtime accrues only while the session is running, and calculates roughly $58/month for an agent running around the clock (730 hours × $0.08), before tokens (source). Runtime is small; tokens are where a runaway loop hurts.
- Per-execution workflow tools. n8n bills per workflow execution (one whole run), not per node or step. A cloud plan with a fixed execution allowance makes workflow cost predictable, but token spend on whatever model you call inside it is separate and still needs a cap. Plan prices vary by billing period and change over time, so check n8n's current pricing page rather than trusting a blog's snapshot.
Illustrative math with round numbers (an example, not a quote): if an agent loops and burns 2 million tokens before anyone notices, at a hypothetical blended $5 per million tokens that's $10. Annoying. At 200 parallel runs the same bug is $2,000. The lesson is that spend scales with concurrency, so per-run caps need a global ceiling as well.
Route by task, not by loyalty
Since model choice isn't what you're locked into, use it. Cheap, fast models for classification and extraction; a stronger model only for steps that need judgment:
routes:
classify_email: { model: small-fast, max_usd: 0.01 }
extract_invoice: { model: small-fast, max_usd: 0.02 }
draft_reply: { model: mid-tier, max_usd: 0.10, approval: human }
contract_review: { model: top-tier, max_usd: 0.50, approval: human }
global:
daily_cap_usd: 15
on_breach: pause_all_and_notify
Model names here are placeholders; map them to whatever you run. The structure is what matters: a cost ceiling and an approval rule per task, plus a global kill switch.
Reliability and observability: what to log from day one
Reliability (20%) and ease of development (19%) were top-three selection factors, and "limited visibility and observability" tied for second among provider-platform risks. You can't improve what you can't replay.
Log these for every agent run, whatever platform you use:
{
"run_id": "2026-10-06-0042",
"trigger": "inbound_email",
"model": "mid-tier",
"tools_called": ["crm.lookup", "invoice.draft"],
"tokens_in": 4210,
"tokens_out": 612,
"cost_usd": 0.031,
"approval": "pending_human",
"outcome": "draft_created",
"error": null
}
Three habits pay off quickly:
- Store the input and the decision, not just the output. When an agent does something odd, you need the context it saw.
- Alert on cost per run and steps per run, not only on errors. A looping agent often returns no errors.
- Keep a "last 20 runs that touched money or customers" view. That's your daily five-minute review.
Governance without a platform team
Security and permissions ranked 10% as a selection factor but 25% as the orchestration investment expected to grow most next year (behind agent workflow tooling at 34%, ahead of monitoring and debugging at 20%; only 5% said their budget isn't increasing). People choose a platform for capability, then discover they need permissions.
A pragmatic governance checklist for a small team:
- Least privilege per agent. Separate credentials for each agent; read-only by default; no shared admin keys.
- Approval gates on irreversible actions. Sending email, issuing refunds, deleting records, and posting publicly go through a human step until you have weeks of clean logs.
- One kill switch. A single flag or workflow that disables all agent triggers. Test it.
- A weekly cost-and-error review. Fifteen minutes, same time each week.
- Written ownership. Every agent has a named person who gets paged when it misbehaves.
This is also where the "who fixes a wrong fact" question gets answered: the person named in the last item.
A decision framework: how to pick your orchestration layer
Use the survey's actual ranking as your checklist, weighted for a small team:
| Question | What good looks like |
|---|---|
| Can I swap the model without rewriting workflows? | Model is a config value, not an SDK dependency |
| Can I enforce a spend cap before the call, not after? | Gateway or workflow node checks budget per run |
| Can I replay any run? | Inputs, tool calls, outputs, cost stored |
| Are my tools portable? | MCP servers or plain HTTP endpoints I own |
| Can I turn everything off in one step? | Single kill switch, tested |
| Can a non-engineer see what happened? | Readable run log or dashboard |
Then match to your situation:
- One workflow, low volume, one model is fine: a vendor-managed agent service can be the fastest path. Add the budget guard anyway.
- Several workflows touching email, CRM, invoicing: a workflow engine such as n8n as the control plane, with models and MCP tools plugged in. Per-execution billing keeps the workflow side predictable.
- Customer-facing or money-touching agents: custom agent code behind your own gateway, with approval gates and full run logging.
- You already run two or more platforms: put a thin gateway in front of all of them for spend and logs. The survey found 59% of platform users run one platform, 23% run two, and 19% run three or more; sprawl is normal, ungoverned sprawl is the problem.
Whichever you choose, schedule a re-evaluation. With the large majority of surveyed enterprises planning platform changes within a year, "we'll revisit this in six months" is a reasonable default, not a failure of planning.
Where this fits in practice
I build AI automations that run in production for solopreneurs and small teams, and the pattern above is the one I'd recommend: keep triggers, approvals, budgets, and logs in a control plane you own, keep tools portable, and keep the model swappable. I'd also suggest giving every agent a per-run budget, a global cap, and a kill switch before it touches a real inbox or ledger.
If you're deciding between a vendor-managed platform, a workflow engine, or custom code and want a second opinion on which parts to own, reach out and we can talk through your workflows.
Work with BizFlowAI
If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.
Book a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.
More guides like this on the BizFlowAI blog.
Frequently asked questions
How do I choose an AI agent orchestration platform for a small business?
Choose on flexibility across models and tools, production reliability, ease of development, and control over agent execution, not on which model vendor runs the platform. A VentureBeat survey of enterprises found only 2% picked a platform because of model alignment, while those four factors accounted for 78% of choices. Keep the model layer swappable and keep triggers, permissions, budgets, and logs in a layer you own. That way a model change never turns into a migration project.
Why shouldn't I lock my agent system to one model vendor's platform?
Models change faster than infrastructure, so an orchestration layer welded to one vendor's runtime turns every model swap into a migration. Surveyed buyers named inflexibility across models and tools (28%) as the top risk of vendor-hosted orchestration, ahead of security limits and weak observability (22% each). You also lose the ability to route work to a cheaper or better model. Separating the swappable model layer from your stable control layer avoids this.
How do I stop an AI agent from running up a huge bill overnight?
Enforce a hard cap per run instead of relying on a spending dashboard. Track estimated cost and step count on every model call, and raise an exception that halts the run when either limit is crossed. Keep per-token prices in config because vendor pricing changes. In the survey, 23% of enterprises only tracked agent spend through after-the-fact logs, which cannot stop a loop that is already running.
What is the agent control plane and who should own it?
The control plane is the layer that decides what an agent may do, what it may spend, and how its actions are audited, covering triggers, approvals, budgets, and logs. In the VentureBeat survey, 67% of enterprises expected it to sit at least partly outside a provider-managed service by the end of 2026. The model vendor supplies the model while you own the rules around it. Even a one-person team can implement this as a gateway that checks permissions, budgets, and logging on every tool call.
Why does MCP make AI agent systems more portable?
Model Context Protocol (MCP) is a standard way to expose tools like CRM lookups or invoice actions to any agent. If your integrations are MCP servers or plain HTTP endpoints you control, any orchestrator and any model can call them, so you write each integration once. The agent runtime then becomes the replaceable part. Anthropic donated MCP to the Agentic AI Foundation under the Linux Foundation, co-founded with Block and OpenAI.