Agent Auth Isn't Enough: Drift, Leaks, Poisoned Memory

You wired up an AI agent, gave it a service account, put a gateway in front of it, and watched the logins succeed. Authentication passed, so you assumed the hard part was done. It wasn't: an agent can be fully authenticated and still wander off task, hand data to the wrong place, or act on instructions someone planted in its memory three weeks ago.
This post walks through the failure modes that authentication doesn't catch, what recent gateway incidents tell us about where to start, and a practical control order for a solo developer or small team.
Authentication answers one question, and it isn't the one that hurts you
Authentication proves who is calling. It does not prove what the caller will do, whether it should be doing it, or whether its instructions are still its own. An agent that passes every login check can still drift from its task, expose data through a legitimate tool call, or reason from poisoned memory.
OWASP's Top 10 for Agentic Applications 2026 (published by the OWASP GenAI Security Project on December 9, 2025) names these failures directly. A few of the entries matter most for small deployments:
| OWASP ID | Name | What it looks like in practice |
|---|---|---|
| ASI03 | Identity and Privilege Abuse | An agent inherits or escalates high-privilege credentials it never needed |
| ASI06 | Memory and Context Poisoning | Attackers corrupt memory, embeddings, or RAG stores to steer decisions across sessions |
| ASI10 | Rogue Agents | A compromised agent acts harmfully while still looking legitimate |
Notice that none of these are "someone stole a password." In all three, the identity is valid. The agent logged in correctly. What went wrong happened after the login.
I'll split the rest of this post into three failure modes: drift, data exposure, and memory poisoning. Then I'll cover the gateway question, because that's where most teams reach first.
Drift: the agent is still you, just not doing what you meant
Drift is when an authenticated agent gradually or suddenly does something different from what you scoped it to do. It rarely looks like an attack. It looks like a slightly different tool choice, a wider search, or a retry loop that now touches a system it never touched before.
Why it happens in small deployments:
- Prompt changes with side effects. You tweak a system prompt to fix one behavior and shift three others.
- Model swaps. A new model version interprets the same tool descriptions differently.
- Tool sprawl. Every MCP server you add expands what the agent can do, whether or not you intended it to.
- Long-running context. The longer a session or memory store lives, the more the agent's effective instructions diverge from the original ones.
The defense is boring and effective: make the agent's allowed behavior explicit and testable, then check actual behavior against it. A minimal version is a per-agent allowlist of tools and a log comparison.
# agent-policy.yaml - one file per agent, checked into git
agent: invoice-followup
allowed_tools:
- crm.read_contact
- email.draft # draft only, never send
- invoices.read
denied_tools:
- email.send
- invoices.write
- files.*
max_tool_calls_per_run: 15
import yaml, json, sys
policy = yaml.safe_load(open("agent-policy.yaml"))
allowed = set(policy["allowed_tools"])
def check_run(log_path: str) -> list[str]:
violations = []
calls = 0
for line in open(log_path):
event = json.loads(line)
if event.get("type") != "tool_call":
continue
calls += 1
tool = event["tool"]
if tool not in allowed:
violations.append(f"unlisted tool used: {tool}")
if calls > policy["max_tool_calls_per_run"]:
violations.append(f"call budget exceeded: {calls}")
return violations
if __name__ == "__main__":
problems = check_run(sys.argv[1])
for p in problems:
print("DRIFT:", p)
sys.exit(1 if problems else 0)
This is not sophisticated, and that's the point. Run it after every agent run (or on a sample), and alert when it fails. The first time a prompt tweak makes your invoice agent reach for files.*, you'll know the same day instead of discovering it in a bill or a customer complaint.
Data exposure: legitimate tool calls, wrong destination
An agent doesn't need to be hacked to leak data. It needs a tool that can read sensitive things and a tool that can send things out. With both, a confused or manipulated agent can connect them.
The standard pattern: the agent reads a customer record (allowed), then includes it in an email draft, a webhook payload, or a support-ticket comment (also allowed). Each step is individually authorized. The combination is the leak.
Practical controls, in order of effort:
- Separate read tools from egress tools per agent. An agent that summarizes inbound email doesn't need
http.postto arbitrary URLs. Give egress to as few agents as possible. - Scope credentials to the narrowest audience. The MCP authorization specification requires servers to validate that access tokens were issued specifically for them as the intended audience (RFC 8707), and invalid or expired tokens must get an HTTP 401. If you run your own MCP server, enforce this. A token minted for one tool should not open another.
- Redact before the model sees it. If the agent doesn't need the full card number, SSN, or API key, strip it upstream. Data the model never receives can't be echoed.
- Log egress with content, not just metadata. "Agent sent one email" tells you nothing. "Agent sent an email containing a field from the customer table" tells you what happened.
One caveat worth stating plainly: MCP authorization is optional in the spec. In the 2026-07-28 revision, an HTTP server that requires it points to the OAuth 2.1 flow, with PKCE (S256) required for clients, and stdio servers should use environment variables. OAuth 2.1 is still an IETF draft, not a finalized standard. So "we use MCP" doesn't mean "we have authorization." Check whether your servers actually enforce it.
Memory poisoning: the attack that looks like a model bug
Memory poisoning is when something untrusted gets written into an agent's persistent memory, embeddings, or RAG store, and later influences its decisions. OWASP describes it as attackers corrupting stored information to manipulate decisions across sessions.
The reason it's nasty is that it's hard to tell from ordinary model failure. An arXiv study on this, "The Misattribution Gap," reports that in 59 of 65 valid entries, agents cited an injected document as normative authority, as though it were a real policy. The same study reports that four safety classifiers returned zero detections across 510 checkpoints. Treat that as one study's result, not a universal rate, but the direction is clear: the agent doesn't flag the poison, and common classifiers don't catch it either.
For a small business, the realistic entry points are mundane:
- A customer email or support ticket that gets summarized into long-term memory
- A scraped web page or PDF that lands in your RAG index
- A shared document someone edited with text addressed to "the assistant"
- A note an earlier (confused) run wrote back to memory, which later runs treat as fact
Defenses that hold up without a security team:
- Track provenance on every memory write. Store source, timestamp, and trust level alongside the content. Memory from a customer email is not the same as memory from your own runbook.
- Don't let agents write to the same store they read instructions from. Keep operating instructions (read-only, in git) separate from learned notes (writable, low-trust).
- Expire and review. Set a TTL on writable memory and review diffs periodically, the way you'd review a config change.
- Treat retrieved text as data, never as policy. In your prompt, state explicitly that retrieved content cannot change rules or grant permissions. This isn't bulletproof, but it raises the bar.
{
"memory_entry": {
"id": "mem_0192",
"content": "Customer prefers invoices sent on the 1st.",
"source": "email:customer_reply",
"trust": "untrusted",
"written_by": "agent:invoice-followup",
"created": "2026-10-01",
"expires": "2026-11-01",
"can_influence": ["scheduling"],
"cannot_influence": ["permissions", "tool_selection", "payment_details"]
}
}
The cannot_influence field is the idea to steal. Even if you can't stop poison from entering, you can limit which decisions low-trust memory is allowed to touch.
The gateway trap: first control teams reach for, last one they're ready to run
An AI gateway (a proxy that fronts model calls and, increasingly, MCP tools) feels like the obvious first control. One chokepoint, one place for auth, logging, and rate limits. The VentureBeat piece by Nik Kale (published August 30, 2026) argues the opposite ordering: gateway controls should be the fifth control for securing agents, not the first, because gateways sit on top of identity and attribution layers that mostly aren't there yet. If you can't say which agent made a call and on whose behalf, the gateway has nothing meaningful to enforce.
There's a second problem: the gateway is itself attack surface. LiteLLM, an open-source gateway, makes the case in this year's record alone:
- March 2026: Two malicious LiteLLM releases (1.82.7 and 1.82.8) were published to PyPI by the TeamPCP group after it obtained a maintainer's publishing credentials. That's a supply-chain compromise, not a code bug. (Cloud Security Alliance)
- April 2026: CVE-2026-42208 (CVSS 9.3), an SQL injection in proxy API key verification, fixed in 1.83.7 on April 19, 2026, and later added to CISA's KEV catalog. Sysdig observed the first exploitation attempt about 36 hours after the advisory was published. (Security Affairs)
- June 2026: CISA's alert dated June 8, 2026 added CVE-2026-42271 (BerriAI LiteLLM Command Injection) to the KEV catalog based on evidence of active exploitation. (CISA) It has a CVSS score of 8.7: a command injection in LiteLLM's MCP preview endpoints that lets any authenticated user run arbitrary commands on the host, affecting versions >= 1.74.2 and < 1.83.7. (The Hacker News)
- June 2026, chained: Horizon3.ai chained it with CVE-2026-48710, a Starlette "BadHost" Host-header authentication bypass, producing unauthenticated remote code execution with a combined CVSS of 10.0. (The Hacker News)
- September 2, 2026: CISA added CVE-2026-59822 (Improper Authentication) to KEV. It lets an unauthenticated attacker with a fabricated Bearer token open an authenticated MCP session and potentially reach configured tools. It's fixed in 1.84.0, and the federal deadline was Sept. 16. (eSecurityPlanet; CISA alert)
VentureBeat's author also says the same gateway had seven CVEs disclosed in a single month. I couldn't find an independent count, so treat that as the author's claim.
Two details deserve emphasis because they bite people who think they patched:
- The fix for CVE-2026-42271 is LiteLLM v1.83.7, which restricts the MCP test endpoints to the PROXY_ADMIN role and updates the Starlette dependency. CISA directed US federal civilian agencies to remediate by June 22, 2026. (Help Net Security)
- The Cloud Security Alliance says the full chain needs both LiteLLM 1.83.7 and Starlette 1.0.1 or later. Without the Starlette update, the CVSS 10.0 unauthenticated chain still works even against patched LiteLLM. Patching the app but not its dependency leaves the door open.
The lesson isn't "don't use a gateway," and it isn't specific to one vendor. A gateway concentrates credentials and tool access in one process. If it's internet-reachable, unpatched, or running with broad permissions, it becomes the best target in your stack. I have no data on how many small businesses run these gateways, so I won't guess; if you do, assume you're in scope.
A control order that fits a team of one to ten
Following the reasoning above, here's the order I'd build in. Each layer makes the next one meaningful.
- Identity and attribution. Every agent gets its own identity. No shared service account. You should be able to answer "which agent did this, for whom?" from logs alone.
- Least-privilege tool scope. Per-agent allowlists like the YAML above. Short-lived, audience-bound tokens where the protocol supports them.
- Data boundaries. Separate read and egress capabilities. Redact upstream. Log content-level egress.
- Memory hygiene. Provenance, TTLs, trust levels, and a hard split between instructions and learned notes.
- Gateway. Now it has identities and scopes to enforce. Run it patched, minimally exposed, and on the least-privileged account you can.
- Behavioral monitoring. Drift checks, anomaly alerts, periodic review of memory diffs and tool-call logs.
And a short operational checklist for the gateway itself:
# Pin versions and verify what's actually installed
pip show litellm starlette | grep -E "^(Name|Version)"
# Compare against current advisories before you trust a version number:
# - CISA KEV catalog: cisa.gov
# - The LiteLLM project's security advisories
# Pin hashes so a poisoned release can't slide in silently
pip install --require-hashes -r requirements.lock
- Pin dependencies and use hash checking, since the March compromise was a poisoned package release.
- Subscribe to advisories for the gateway and its web framework dependencies; the June chain needed both.
- Don't expose admin or MCP test endpoints publicly. If a test endpoint exists, restrict it to an admin role and a private network.
- Rotate keys the gateway holds after any suspected exposure, and keep those keys scoped so one leak doesn't unlock everything.
- Check current versions and fix status on the official advisories; version numbers above reflect what was reported at the time of writing.
What to do this week
You don't need all six layers by Friday. Start here:
- List every agent you run and every tool each can call. If you can't produce that list, that's finding number one.
- Write a per-agent allowlist and a drift check. An hour of work, immediate visibility.
- Look at what each agent can both read and send. Cut the egress you can't justify.
- Find every place an agent writes to memory or an index. Add source and expiry fields.
- If you run a gateway, confirm its version and its web-framework dependency version against current advisories, and take admin/test endpoints off the public internet.
A note on where this fits
I'm Lazar, a senior engineer who builds AI automations that actually ship, and the control order above is the one I'd want for any agent setup, including my own: per-agent identities and tool scopes first, explicit egress boundaries, provenance on anything an agent writes to memory, and the gateway added after those layers exist rather than instead of them. Treat the gateway as software that needs patching and minimal exposure, not as a security guarantee.
If you want a second set of eyes on your own stack, walk through the same questions yourself: what each agent can reach, where data can leave, and what an attacker (or a confused model) could get it to do.
Work with BizFlowAI
If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.
Request a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.
More guides like this on the BizFlowAI blog.
Frequently asked questions
Is authentication enough to secure an AI agent?
No. Authentication only proves who is calling, not what the agent will do afterward or whether its instructions are still trustworthy. An authenticated agent can still drift from its task, leak data through legitimate tool calls, or act on poisoned memory. OWASP's Top 10 for Agentic Applications 2026 lists identity and privilege abuse, memory poisoning, and rogue agents as failures that happen after a valid login.
What is agent drift and how do I detect it?
Agent drift is when an authenticated AI agent gradually or suddenly does something different from what you scoped it to do, such as using new tools or widening its searches. It is usually caused by prompt changes, model swaps, tool sprawl, or long-running context. A simple detection method is a per-agent policy file listing allowed tools and a call budget, plus a script that compares each run's tool-call log against it and alerts on violations.
How can an AI agent leak data without being hacked?
An agent leaks data when it has one tool that reads sensitive information and another that sends data out, such as email, webhooks, or HTTP posts. Each step is individually authorized, but chaining them can move private data to the wrong place. To reduce the risk, separate read tools from egress tools, scope credentials narrowly, redact sensitive fields before the model sees them, and log the content of outbound actions.
What is memory poisoning in AI agents and how do I defend against it?
Memory poisoning happens when untrusted content, such as a customer email, scraped page, or edited document, gets written into an agent's persistent memory or RAG store and later influences its decisions. It is hard to spot because it looks like an ordinary model error. Defenses include tracking source and trust level on every memory write, keeping read-only instructions separate from writable notes, expiring memory with a TTL, and restricting which decisions low-trust memory can influence.
Does using MCP mean my agent tools are authorized?
No. Authorization is optional in the MCP specification, so using MCP does not guarantee your servers enforce it. When a server does require it, it should validate that access tokens were issued for that specific server as the audience (RFC 8707) and return HTTP 401 for invalid or expired tokens. HTTP servers can use the OAuth 2.1 flow with PKCE, though OAuth 2.1 is still an IETF draft, so you should verify your servers actually enforce these checks.