Snowflake Cortex AI Gateway: Governance For Agents

Your team just wired Claude Code into the customer database. A junior dev pointed Cursor at the production Snowflake account for a "quick refactor." Someone's custom agent is looping on a Bedrock call at 3 AM. Nobody knows what the bill looks like until Friday. This is the actual state of AI adoption inside most companies right now, and it's why Snowflake shipped Cortex AI Gateway on Tuesday.
The gateway is not a new model. It's not a new chat UI. It's a control plane — the thing your platform team should have been asking for six months ago but couldn't buy off the shelf.
What Snowflake actually shipped
Cortex AI Gateway is a centralized control layer that governs how AI agents access enterprise data, tools, and models inside Snowflake. That includes third-party agents — Claude Code, Cursor, and whatever internal agents your team is building against the warehouse. The gateway sits between the agent and the data, and it decides what runs, what gets logged, and what gets blocked.
Alongside the gateway, Snowflake announced a first wave of security integrations with 1Password, Aembit, Linx Security, SailPoint, and Saviynt. That lineup matters more than the gateway itself. Identity and secrets vendors don't co-launch with a data warehouse unless the picture they're painting is: AI agents are the next identity problem, and nobody has solved it yet.
If you've been running agents in anything larger than a hobby project, you already know why. An agent isn't a user. It's not a service account either. It's something in between — a workload that authenticates like a service, requests data like a user, and can spend money like a rogue cron job.
Why "runaway costs" is the honest framing
Snowflake's launch language leads with cost. That's rare from a company that historically leads with performance and scale. It's also the correct framing.
Here's what actually happens without a gateway:
- An agent gets a task. It plans. It calls a model to plan. It calls the model again to critique the plan. It runs a SQL query. The query returns 40k rows. It feeds those rows back into the model to summarize. The summary triggers three more tool calls.
- Multiply by a team of eight developers each running their own agent sessions all day.
- Multiply by an autonomous overnight job that retries on failure.
Nobody set out to spend $18k on inference in a week. It just happened, one plausible tool call at a time. I've seen this exact pattern in three separate engagements this year — an agent that "worked" in dev quietly burning a five-figure monthly bill in prod because token accounting was invisible to the team shipping the feature.
A gateway fixes this in a specific way: every agent call passes through one enforcement point that knows the identity, the budget, the allowed tools, and the audit destination. If Claude Code wants to query your orders table, the gateway decides whether that's allowed, logs the query, meters the tokens, and cuts the request off if the caller is over budget. No gateway, no enforcement.
The five capabilities that matter in an agent gateway
Snowflake's launch page will list more features than this. These are the ones that determine whether a gateway is real infrastructure or a dashboard with a login screen.
| Capability | What it means | Why it matters |
|---|---|---|
| Identity binding | Every agent request carries a verifiable identity (workload, human, or delegated) | Without this you cannot audit or bill anything |
| Policy enforcement | Rules about which tools, tables, models an agent can hit | Least privilege for non-human actors |
| Token + cost metering | Per-agent, per-tenant, per-workflow accounting | Turns "AI spend" into a line item, not a surprise |
| Prompt + output logging | Full trace of what the agent asked and got back | Required for debugging and for security review |
| Kill switch | Revoke or throttle a specific agent in real time | The one feature you'll wish for at 2 AM |
Notice how boring this list is. There's nothing here about reasoning, or planning, or multi-agent orchestration. That's the point. Governance is boring plumbing, and boring plumbing is what enterprises actually pay for once the demos are over.
What this means for Claude Code and Cursor users
Snowflake explicitly called out Claude Code and Cursor as agents the gateway will govern. Read that carefully: Snowflake is not blocking them. It's making them first-class citizens inside a control plane.
If you're a solo developer or a small team already using Claude Code against your Snowflake account, nothing breaks today. You keep working the way you work. What changes is the enterprise motion around you:
- Your customer's security team can now say yes to Claude Code touching their warehouse, with an audit trail, instead of blanket-blocking it.
- The token bill from your agent runs shows up on the same invoice as your Snowflake compute, which is either a feature or a nightmare depending on your finance setup.
- The "shadow agent" problem — engineers wiring up their own automations without approval — has a defensible answer: run them through the gateway or don't run them at all.
For small teams building agents that touch enterprise data as a product, this is the shape of the buyer conversation for the next 18 months. Your prospect's CISO is going to ask, "Does this go through Cortex AI Gateway?" or the equivalent from AWS Bedrock AgentCore, Databricks, or whatever their platform team standardizes on. Have an answer ready.
The identity vendor pileup tells you where the market is going
1Password, Aembit, Linx Security, SailPoint, and Saviynt on the same launch page is not a coincidence. Two of those (SailPoint, Saviynt) are legacy identity governance. Two (Aembit, Linx) are newer workload-identity plays. 1Password is the consumer/team secrets manager that has been quietly moving into developer secrets.
The bet: AI agents will be the largest population of "users" inside enterprises within a few years, and none of the existing IAM tooling was designed for a caller that spawns dozens of ephemeral sub-agents, each needing a scoped credential for eight seconds.
The practical implication for a small team building agents right now:
- Do not hardcode API keys in agent prompts or config files. This has always been bad. It's now unshippable to any customer over 500 employees.
- Assume every agent needs a workload identity — an OIDC token, a short-lived credential, something that maps back to a specific workflow.
- Log the identity on every tool call. Not the human who "started" the agent. The agent's own workload identity. Otherwise your audit log is a lie.
Here's the shape of what a compliant agent tool call looks like from the outside:
{
"timestamp": "2026-09-28T14:22:11Z",
"agent_id": "invoice-triage-v3",
"workload_identity": "wi_a8f2...",
"delegated_from": "user_j.rivera@acme.com",
"tool": "snowflake.query",
"resource": "PROD.FINANCE.INVOICES",
"policy_decision": "allow",
"tokens_in": 1840,
"tokens_out": 512,
"cost_usd": 0.023
}
Every one of those fields is a control point. Miss any of them and you have a governance gap.
A practical architecture for teams that aren't Snowflake customers
Most of the SMBs I work with don't run on Snowflake. They run on Postgres, or BigQuery, or a mix of SaaS APIs. The gateway pattern still applies. You just build a lighter version of it yourself.
Here's a minimal Python sketch of the enforcement point every serious agent deployment needs:
from dataclasses import dataclass
from typing import Callable
@dataclass
class AgentContext:
agent_id: str
workload_identity: str
delegated_user: str | None
monthly_budget_usd: float
class AgentGateway:
def __init__(self, policy, ledger, logger):
self.policy = policy
self.ledger = ledger
self.logger = logger
def call_tool(self, ctx: AgentContext, tool: str,
params: dict, run: Callable):
# 1. Policy check
decision = self.policy.evaluate(ctx, tool, params)
if not decision.allow:
self.logger.log_denied(ctx, tool, params, decision.reason)
raise PermissionError(decision.reason)
# 2. Budget check
spent = self.ledger.month_to_date(ctx.agent_id)
if spent >= ctx.monthly_budget_usd:
self.logger.log_over_budget(ctx, spent)
raise RuntimeError("agent over budget")
# 3. Execute + meter
result, cost = run(params)
self.ledger.record(ctx.agent_id, tool, cost)
self.logger.log_call(ctx, tool, params, result, cost)
return result
Three checks, one execution, three log lines. That's the whole pattern. You can host this behind an internal API, wire your agents to hit it instead of calling tools directly, and you now have 80% of what Cortex AI Gateway gives you at the enterprise tier.
The policy engine can be as simple as a YAML file to start:
agents:
invoice-triage-v3:
allowed_tools:
- snowflake.query
- email.send
forbidden_resources:
- PROD.HR.*
- PROD.FINANCE.PAYROLL
max_rows_per_query: 5000
max_tool_calls_per_run: 25
max_tool_calls_per_run is the single most under-appreciated control. It's the seatbelt that catches the runaway agent before your inference bill catches it.
What to actually do this month
If you're a solo dev or small team running agents against real business data, do these in order:
- Inventory your agents. Not the sexy ones. All of them. The scheduled job. The Slack bot. The Claude Code sessions that hit prod. If you can't list them on one page, you don't have an agent program, you have an agent problem.
- Give each agent its own identity. No shared service accounts. No personal API keys.
- Put a wrapper in front of every tool call. Even if the wrapper only logs, you now have the data to write policy against.
- Set a per-agent monthly budget. Not org-wide. Per agent. When one goes rogue you want to know which one.
- Add a kill switch. A single config flag that stops an agent from making external calls. Test it. Actually test it, at least once.
- Decide which platform gateway you'll trust. If you're on Snowflake, Cortex AI Gateway is the obvious answer. On AWS, AgentCore. On Databricks, their AI Gateway. On your own stack, you build the lightweight version above.
None of this is about being anti-agent. It's the opposite. Governance is what makes agents shippable. The teams that skip it don't move faster — they just accumulate risk faster and then hit a wall the first time a customer's security review asks a real question.
How BizFlowAI approaches this
Every agent architecture we design for clients starts with the boring layer: identity per agent, per-tool policy enforcement, cost metering at the call site, and a kill switch that actually works. We do this before we write a single planning prompt, because retrofitting governance onto an agent that's already in production is three times the work and half as effective. When a client asks us to add Claude Code or Cursor to their workflow against Snowflake, BigQuery, or a Postgres warehouse, the first question is always the same: how are we going to know what these agents did, what they cost, and how do we turn them off in a hurry?
That's the difference between an agent demo and an agent you can put in front of your customers' data. If you're weighing Cortex AI Gateway, AgentCore, Databricks AI Gateway, or a self-hosted control plane and you want a straight answer on which fits your setup, book a discovery call — we've built inside all three patterns and we'll tell you honestly which one wastes the least of your time.
Work with BizFlowAI
If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.
Book a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.
More guides like this on the BizFlowAI blog.
Frequently asked questions
What is Snowflake Cortex AI Gateway?
Cortex AI Gateway is a centralized control plane Snowflake launched to govern how AI agents access enterprise data, tools, and models. It sits between agents like Claude Code or Cursor and the Snowflake warehouse, deciding what runs, what gets logged, and what gets blocked. It handles identity binding, policy enforcement, token metering, prompt logging, and real-time kill switches. It is not a new model or chat UI, but governance infrastructure.
Why do AI agents cause runaway cloud costs?
AI agents chain multiple model calls to plan, critique, query, and summarize, and each step consumes tokens invisibly. When multiplied across a team of developers plus overnight autonomous jobs, a single feature can quietly burn five figures per month. Token accounting is usually invisible to the team shipping the agent, so bills only surface at month-end. A gateway meters every call per-agent and per-workflow to turn AI spend into a tracked line item.
How does Cortex AI Gateway affect Claude Code and Cursor users?
Snowflake explicitly supports Claude Code and Cursor as governed agents rather than blocking them. Solo developers and small teams keep working as before, but enterprise security teams can now approve these tools with a full audit trail. Token costs appear on the same invoice as Snowflake compute, and shadow agent usage becomes controllable. Vendors selling agent-based products should expect CISO questions about gateway compatibility.
What capabilities does an AI agent gateway need?
A real agent gateway needs five things: identity binding so every request carries a verifiable workload or delegated identity, policy enforcement to restrict which tools and tables agents access, token and cost metering per agent and workflow, full prompt and output logging for debugging and audit, and a kill switch to throttle or revoke a specific agent in real time. Without all five, you have a dashboard, not infrastructure.
How can small teams build an agent gateway without Snowflake?
Teams on Postgres, BigQuery, or SaaS APIs can build a lightweight enforcement point in Python that wraps every tool call with three checks: a policy evaluation, a budget check against a monthly limit, and a metered execution with logging. Agents call this internal API instead of hitting tools directly. This pattern replicates roughly 80 percent of what Cortex AI Gateway provides using about 30 lines of code.