Cohere North 2: Model-Agnostic Agents, Minus the Lock-In

You run a small business and your AI tools can't talk to your Gmail, your CRM, or your billing system. Cohere just shipped North 2, an agent platform that claims to work with any model, so I read the launch coverage to see what's confirmed, what's vendor claim, and what you can steal today without buying anything.
Short version: North 2 is aimed at large organizations, not solopreneurs. But the questions it forces you to ask apply to every agent platform you'll ever evaluate.
What Cohere actually announced
Cohere announced North 2 on October 5, 2026, as the latest version of its North enterprise agent platform (Unite.AI). The confirmed feature list is: cross-session agent memory, reusable Skills, shared Libraries, Automations (templates plus a drag-and-drop workflow builder with real-time performance monitoring), and Applications (decks, dashboards and lightweight apps generated from natural language).
Cohere's own launch post lists "model agnostic" as a headline feature: you can use Cohere's models or bring your own (Cohere). Trade press has called it a "control room," but that framing comes from coverage, not from Cohere. Cohere calls it an agent platform or harness.
What I'm not going to do is invent pricing, benchmarks, or customer results. Here's what's actually public, and what isn't:
| Question | What's public |
|---|---|
| Pricing | Not published. Cohere doesn't publish standard prices; cost depends on deployment scale and complexity, and new customers contact sales (BERI) |
| Compatibility matrix | Not published (AlphaSignal) |
| Memory retention rules | Not published (BERI) |
| Independent benchmarks | None listed in the launch coverage I reviewed |
| Feature availability dates | Some features have already shipped to customers, per Cohere's Joelle Pineau in the Globe and Mail; which ones is unspecified |
If you're evaluating it, those gaps are your question list for the sales call.
Model-agnostic: the architecture is right, the proof is missing
The architectural idea is the one I care about most: keep the orchestration layer separate from the model so you can swap the model and keep the workflow. Model prices and quality shift every few months. The best model for your invoice parsing today may not be the best in six months, and a workflow welded to one vendor means a rebuild every time.
Here's what's claimed. Cohere's VP of product for AI, Paul Teyssier, said North 2's model, harness, search and embedding components are interchangeable, positioned as a way to reduce lock-in (VentureBeat). An administrator can assign models per team, including models brought from outside Cohere (SiliconANGLE). According to The Logic, customers can plug in models from Anthropic, OpenAI and Google, or cheaper open-source models from Alibaba, DeepSeek and Z.ai, and unlike Claude Code or Codex, North 2 is open to other developers' models (The Logic).
Here's the catch: there is no published compatibility matrix covering model providers, tool-calling formats, context limits, or supported deployment combinations (AlphaSignal). And in the launch coverage I reviewed, I found no independent tests of tool-calling, memory, or guardrails with third-party models. "Works with any model" and "works well with any model" are different claims. Tool-calling formats differ between providers, and a platform that accepts a model can still handle its function calls badly.
You can build the same separation yourself in a few lines. This is the pattern I use in my own automations: the workflow only knows a role name, and a config file maps roles to models.
# models.yaml - the only file that changes when you swap a model
roles:
invoice_parser:
provider: your-provider-a
model: your-model-a
email_triage:
provider: your-provider-b
model: your-model-b
import yaml
CONFIG = yaml.safe_load(open("models.yaml"))
def run_step(role: str, prompt: str) -> str:
cfg = CONFIG["roles"][role]
client = get_client(cfg["provider"]) # your thin adapter layer
return client.complete(model=cfg["model"], prompt=prompt)
# workflow logic never mentions a vendor
parsed = run_step("invoice_parser", raw_email_text)
If swapping a model means editing one line in models.yaml and re-running your test set, you have the lock-in protection that platforms are selling. If it means touching workflow code, you don't.
Memory is a feature and a liability
Cross-session memory is the difference between a tool and a colleague. An agent that remembers a client always pays late, or that you already refused their discount request last month, saves you the re-explaining tax that generic chatbots impose.
It's also a risk, because remembered context can be wrong, stale, or something you never meant it to keep. And here's the gap in North 2's launch: according to a BERI analysis, Cohere's post gives no retention period, deletion mechanism or storage location for agent memory, and no admin control specific to memory (BERI). Cohere may well have these controls and just not have published them. That's exactly the thing to ask about before signing.
My memory test works on any platform, and takes ten minutes:
- Tell the agent a false fact ("client X always gets 30% off").
- Open the memory view. Can you find that entry?
- Delete it. Can you?
- Start a fresh session and ask a question that depends on it. Is it gone?
If you can't complete step 2 or 3, you don't control the agent's memory, whatever the marketing says.
Multi-step workflows break at the handoffs
Most small-business automation breaks at step three, not step one. An agent can read an email fine. Trouble starts when it has to read the email, look up the client, draft an invoice, wait for your approval, send it, and log it. Every handoff is a place where things stall or quietly go wrong.
North 2's Automations feature (templates plus a drag-and-drop builder with real-time performance monitoring) is aimed at exactly this. Whether it holds up on messy data is something I couldn't confirm from the coverage, and the customer references (Bell Cyber, LG CNS, CoreWeave) were picked by Cohere with no independently measured results.
One honest limitation for small teams: the launch connector list covers Slack, SharePoint, OneDrive, Microsoft Outlook, Microsoft Exchange, Jira, Linear, Notion and GitHub, with financial-data connectors (PitchBook, Crunchbase, Daloopa, FiscalAI, S&P Global, FactSet) listed as planned (Unite.AI). I found no Gmail, CRM, or billing connector in that launch list. If your stack is Gmail plus a CRM plus an invoicing tool, that's the first thing to verify.
When you test any workflow platform, don't accept a demo with clean inputs. Run this instead:
# Build a stress set from your real last 30 days
mkdir test_set
# Pull 10 real emails: 5 normal, 3 messy (typos, forwarded threads,
# attachments), 2 that should NOT trigger any action
# Feed all 10 through the full chain: read -> lookup -> draft -> approve -> send -> log
# Score each one: correct / wrong-but-caught / wrong-and-sent
The number that matters is "wrong-and-sent." Aim for zero, and treat anything above it as a reason to add an approval gate.
Cost caps, guardrails, and what's still vendor claim
The most practical part of North 2 for any team size is spend control. North Admin tracks token use down to the individual user and agent. Admins can define consumption tiers based on request and token rates, cap usage across the company, and set alerts that warn teams before a limit is reached (SiliconANGLE). It can also downgrade to cheaper tools for simpler tasks, though per The Logic, Cohere hasn't built its own model router into the platform yet (The Logic).
Per-agent guardrails screen prompts and responses for personally identifiable information and prompt-injection attempts, and autonomy policies restrict each agent to authorized actions. Treat those as vendor claims: the launch materials I reviewed don't state detection rates.
The same goes for cost. Cohere and Nvidia claim Cohere's models produce more tokens per second per node on Blackwell and Hopper GPUs, but no benchmark tasks, baselines, latency targets or pricing assumptions were published, so it can't be turned into a cost per query (AlphaSignal). Don't budget around "more tokens for less" until someone measures it.
You can steal the spend-cap idea today, even on a DIY stack. A hard cap in front of your model calls takes minutes:
DAILY_TOKEN_CAP = 500_000 # example round number; set from your own bills
usage = load_usage_today()
def guarded_call(role, prompt):
if usage.tokens >= DAILY_TOKEN_CAP:
alert("Token cap hit - agent paused") # Telegram, email, whatever you read
raise RuntimeError("cap reached")
result = run_step(role, prompt)
usage.add(estimate_tokens(prompt, result))
return result
Pausing an agent at a cap is annoying. A runaway loop on a Saturday night is worse.
The honest fit question, and a test for any platform
Is North 2 for a five-person team? Probably not. Coverage frames it as a platform for large organizations, there's no self-serve pricing, and new customers have to contact Cohere sales. Cohere doesn't publish standard prices; cost depends on the scale and complexity of the deployment, so ask for a quote. It deploys on-premises, in a customer's VPC, hybrid, self-hosted, or fully air-gapped, and Cohere lists SOC 2 Type 2, ISO 27001 and ISO 42001 certifications (Unite.AI). That's an enterprise buyer's checklist, not a solopreneur's.
Context worth knowing: North 2 arrived less than three weeks after Cohere signed a definitive merger agreement with Aleph Alpha on September 16, 2026, subject to regulatory approval and expected to close later in 2026 (Unite.AI). I'm not quoting company valuation figures. For a buyer, the practical point is that the company is mid-merger, so ask how the roadmap and support commitments are affected.
Whatever platform you consider, run this test on one real workflow, the task that eats two hours of your week. Write down its steps, then ask:
- Can you swap the underlying model in under an hour without touching workflow logic? If the answer involves a consultant and a quote, it isn't model-agnostic in any way that helps you.
- Where does a human approve? Pick the step where a mistake costs real money, like sending an invoice or replying to a client, and find out how the platform pauses there. An agent that can't stop and ask will eventually do something you have to apologize for.
- Can you see what it remembers? Open the memory, read it, delete one entry.
- Can you cap spend per agent? And does it alert you before the limit, not after?
- How fast can you catch it being wrong?
That last one is the metric I'd judge North 2, and everything like it, on. A dashboard full of green checkmarks tells you the agent finished, not that it was right. For a small business, a wrong answer delivered quickly is worse than a slow human, because nobody double-checks the automation they trust. If Cohere nails fast error detection, it's a serious option for the organizations it targets. If it's mostly orchestration polish, you're paying enterprise prices for a nicer flowchart.
Where I'd start instead
If a platform like this is out of reach, you don't need one to apply the lessons above. Keep the model behind a config file, put a human approval step before anything that sends money or messages to a client, make memory readable and deletable, and cap spend with an alert. All of that works on a plain script today. I'm Lazar Milićević, a senior engineer who builds practical AI automation for solopreneurs and small teams, and bizflowai.io is where I share what I ship.
Want more like this?
I share practical AI automations I've built.
Planning an AI automation project or need a second opinion on your architecture?
Connect with me on LinkedIn — Lazar Milićević, senior engineer for AI automation.
Visit bizflowai.io for our services, case studies, and AI consulting.
Frequently asked questions
What is Cohere North 2?
North 2 is the next version of Cohere's enterprise AI platform, pitched as a control center for AI agents. According to The Decoder, it handles multi-step workflows on its own, retains context across sessions, and works with any model, so users are not locked into Cohere's own models. Pricing, benchmarks, and customer numbers were not confirmed, so check Cohere's documentation before buying.
Why does model-agnostic design matter for small teams using AI agents?
Model prices and quality change every few months, so the best model for a task today may not be the best in six months. If a workflow is tied to one vendor, switching means rebuilding it. When the orchestration layer is separate from the model, you can swap the model and keep the workflow. Cohere claims North 2 works with any model, but that needs testing.
How do I evaluate an AI agent platform before adopting it?
Run a three-question test on one real workflow. First, can you swap the underlying model in under an hour without changing workflow logic? Second, where does a human approve, especially at steps where mistakes cost money? Third, can you view what the agent remembers and delete an entry? Pick a task that takes about two hours weekly and test it in an afternoon.
Why do multi-step workflows matter when automating small-business tasks with AI agents?
Most small-business automation breaks at step three, not step one. Reading an email is easy, but reading it, looking up the client, drafting an invoice, waiting for approval, sending it, and logging it involves many handoffs where things can stall or quietly go wrong. A platform claiming to manage the whole chain targets that problem, but you should ask for a live run on your own messy data.
What are the risks of AI agents that remember context across sessions?
Persistent context lets an agent recall details like a client's payment habits or a declined discount, so you don't re-explain things each time. The risk is that remembered information can be wrong, stale, or something you never meant it to keep. Any agent with memory should let you see what it remembers and delete entries. If you can't, you don't control it.