Nous Hermes for Business: What Makes an Agent

Developer working on a laptop with code on screen, building and monitoring an AI agent workflow

You tried an agent demo last month. It booked a meeting, drafted an invoice, summarized a thread, and you thought: this could take 10 hours a week off my plate. Then you pointed it at your real inbox and real accounting data, and the questions started: what can it touch, who approves what, and how will I know when it breaks at 3 a.m.? Nous Research's new funding is a good excuse to answer those questions properly.

What Nous Research announced, and what's still fuzzy

Nous Research, the startup behind the open-source Hermes Agent, confirmed a $90 million Series B at a $1.5 billion valuation, according to TechCrunch's report. Robot Ventures led, with participants including Nvidia, Union Square Ventures, Menlo Ventures, Samsung and 1789 Capital. Total funding is reported at $158 million, per TechCrunch-derived coverage. Investor lists vary by outlet, so don't treat any one list as complete.

The money funds an enterprise push called "Hermes for Businesses": customized agents that handle multi-step workflows while keeping data private. Some context on the product line:

  • Hermes Agent launched in February 2026 under the MIT license, with persistent memory across sessions and self-generated reusable skills.
  • Hermes Business, a Nous Portal team tier, was announced around September 14, 2026. It has one shared credit balance, per-member spending caps and shared skills.
  • Hermes Enterprise offers the same capabilities on-prem or in the customer's chosen cloud, with single sign-on, SLAs and onboarding support, per Tao Media.
  • Hosted tiers run $20 to $200 a month; self-hosting is free, per Startup Fortune.

A few things I'd treat carefully:

  • Adoption numbers are company-supplied. Nous says Hermes Agent has been cloned more than 24 million times and drives roughly 2.5% of global AI token usage. Cloned is not the same as installed or used, and the token-share figure hasn't been independently verified.
  • Revenue is reported, not audited. The Wall Street Journal reported about $36 million in annualized revenue by mid-September 2026, with an expectation of passing $100 million before year-end.
  • It's unclear what's new on October 7. The business and enterprise tiers appeared in September, so the headline "launch" may be the same product. Check Nous's pricing page for current availability.

None of this tells you whether the agent will survive your Tuesday. For that you need to look at what happens after the demo.

The demo-to-production gap

Gartner predicted in June 2025 that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls. The same release says only about 130 of the thousands of agentic AI vendors are real, and describes "agent washing": rebranding chatbots, RPA and assistants without substantial agentic capability.

Notice the three cancellation reasons: cost, unclear value, risk controls. None of them is "the model wasn't smart enough." That matches what I see building these systems. A demo agent is a model plus a prompt plus a happy path. A production agent is everything around the model:

Concern Demo agent Production agent
Permissions Whatever the API key allows Explicit tool whitelist per integration
Writes (send, pay, delete) Executes immediately Approval gate, logged
Failures Silent or a stack trace Retries, timeouts, alerts
Cost Unknown until the bill Caps per member and per workflow
State Lost on restart Persistent memory, versioned skills
Visibility Chat log Audit trail you can search

The rest of this post walks through those rows, using Hermes's documented features where they exist, because they're a good concrete reference for what "production-minded" looks like in an open-source agent.

Guardrail 1: Scope the tools, not the prompt

The most common mistake I see is trying to enforce safety with instructions: "never delete anything." Instructions are suggestions. Tool lists are walls.

Hermes's MCP docs let you filter each MCP server's tools with include/exclude lists, so the agent only sees what you expose. The docs call a tight whitelist "usually the best default for sensitive systems". I'd go further: whitelist by default for every integration, not just sensitive ones.

An illustrative config shape (check the current MCP config reference for exact key names before copying):

mcp_servers:
  accounting:
    # Illustrative: expose read tools plus one controlled write
    tools:
      include:
        - list_invoices
        - get_invoice
        - create_draft_invoice
      # No send, void, or delete tools are ever visible to the agent

The point is that delete_invoice doesn't exist as far as the agent is concerned. It can't be talked into calling a tool it can't see, whether the instruction came from a confused user or a prompt injection hidden in an email body.

The practical rule: for each integration, write down the five actions you actually want, and expose only those. Everything else is opt-in later, after you've watched the agent run for a couple of weeks.

Guardrail 2: Put humans on the writes

Reading is cheap to get wrong. Writing is not. A bad summary wastes a minute; a wrong invoice sent to a client or a payment fired twice costs money and trust.

Hermes supports a trust tier on MCP servers: set a server to untrusted, and every write-capable tool call on it requires user approval before it runs. Its documentation also lists a Security section covering command approval, DM pairing and container isolation, per the project README.

You don't need Hermes to apply the pattern. The design I use with clients is three tiers:

  1. Auto: read-only and reversible actions (search, summarize, draft, tag).
  2. Approve: anything that leaves the building or moves money (send email, issue invoice, charge card, post publicly).
  3. Never: destructive or irreversible actions the agent has no business doing (bulk delete, changing credentials, altering bank details).

Hermes's own user stories page describes this exact shape for a small business: a service company ran lead to payment over Telegram, with the owner approving each step. Estimates went out after owner approval, Stripe collected payment, and QuickBooks stayed reconciled (source). That is a vendor-published example, not an independent case study, but the structure is right: the agent does the legwork, the human owns the commitments.

Approval gates have a cost, and you should be honest about it. Every gate is friction. If you gate everything, the owner stops reading the prompts and taps approve reflexively, which is worse than no gate. Gate the few actions that matter and automate the rest.

Guardrail 3: Treat MCP integrations as dependencies that fail

MCP (Model Context Protocol) is how agents plug into your tools: accounting, CRM, email, calendars, databases. Every MCP server is a dependency, and dependencies go down, hang, expire tokens, and change behavior.

Hermes's defaults are instructive. Per the MCP config reference, the default tool-call timeout is 300 seconds and the default connect timeout is 60 seconds, and remote servers can authenticate with OAuth 2.1 with PKCE. Five minutes is a long time for an agent to sit on a hung call while a customer waits. For anything customer-facing, I tighten timeouts and decide up front what happens on failure: retry once, then escalate to a human, never loop.

Authentication deserves the same attention. Prefer OAuth with scoped permissions over pasting a full-access API key into a config file. When a token expires, the failure should be loud, not a silent no-op.

Questions to answer for every integration before it goes live:

  • What's the narrowest scope that does the job?
  • What's the timeout, and what happens when it fires?
  • If this server is down for an hour, does the workflow queue, skip, or alert?
  • Where do credentials live, and who can rotate them?

Guardrail 4: Monitor like it's a service, because it is

An agent that runs unattended is a production service. It needs health checks, and the lowest-effort version is a cron job.

Hermes ships a health check for exactly this: hermes mcp test <server> exits 0 on a completed connect, 1 on connection failure, and 3 when the server isn't in the config, so a watchdog can branch on the exit code (docs). A minimal watchdog:

#!/usr/bin/env bash
# Run from cron every 10 minutes. Alerts only on failure.
SERVER="accounting"

hermes mcp test "$SERVER"
code=$?

case $code in
  0) exit 0 ;;                                   # healthy
  1) MSG="MCP server '$SERVER': connection failed" ;;
  3) MSG="MCP server '$SERVER': not found in config" ;;
  *) MSG="MCP server '$SERVER': unexpected exit $code" ;;
esac

# Replace with your alert channel (Slack webhook, email, SMS)
curl -s -X POST "$ALERT_WEBHOOK_URL" \
  -H 'Content-Type: application/json' \
  -d "{\"text\": \"$MSG\"}"

Connection health is only the first layer. The monitoring I add for production agents:

  • Run logs: every tool call with inputs, outputs, duration and who approved it. When a client asks "why did it email that vendor?", you need the answer in 30 seconds.
  • Cost tracking: spend per workflow and per day, with a hard cap. Hermes Business exposes per-member spending caps for this reason; if you self-host, build the equivalent.
  • Outcome checks: a weekly sample of agent outputs reviewed by a human. Silent quality drift is the failure no uptime check will catch.
  • Dead-man's switch: if a scheduled workflow doesn't run at all, alert. Absence of errors is not evidence of success.

Hosted, business tier, or self-hosted?

Nous's tiering gives a useful framing for the build-versus-buy decision, whichever agent you pick.

Option Good for What you take on
Self-host (free, MIT license) Technical owners, tight data control Hosting, upgrades, monitoring, security
Hosted ($20 to $200/month per reported pricing) Solo operators, fast start Less control over where data runs; check the current pricing page
Business/Enterprise tiers Teams needing shared credits, spending caps, SSO Vendor dependence; verify availability and terms with Nous

Self-hosting isn't free in the way that matters: the license is free, the operations aren't. If nobody on your team will watch the watchdog, a hosted tier may be cheaper in real terms. If your data can't leave your infrastructure, self-hosting or the enterprise option is the point.

Also keep Gartner's "agent washing" warning in mind when evaluating any vendor, Nous included. Ask for a live walkthrough of one real workflow with approval gates, logs and a forced failure. A vendor who can only show the happy path is showing you a demo.

A pre-launch checklist for any business agent

Before an agent touches live customer or financial data, I want yes on all of these:

  1. Every integration uses a tool whitelist, not full access.
  2. Every write action that leaves the building has an approval gate.
  3. Timeouts are set deliberately, with a defined failure behavior.
  4. A health check runs on a schedule and alerts a human.
  5. Every tool call is logged with enough detail to reconstruct a decision.
  6. There's a cost cap, and someone gets notified before it's hit.
  7. Credentials are scoped, stored properly and rotatable.
  8. There's a documented way to switch the agent off in under a minute.
  9. A human reviews a sample of outputs weekly for the first month at least.

If you can't tick all nine, you have a demo, not a system. That's fine as a stage, just don't confuse the two.

How BizFlowAI approaches this

Most of what I build for clients is the unglamorous list above: tool whitelists per integration, approval gates on anything that sends or spends, MCP connections with deliberate timeouts and scoped auth, and monitoring that pages a human when something stops working. The model is the easy part to swap. The guardrails and the run logs are what let a small team trust the agent with real work.

If you're weighing whether to adopt Hermes, another agent framework, or a custom build, and want a straight answer about what will hold up in your workflows, book a discovery call. I'll tell you plainly where an agent fits, where plain automation is the better tool, and where it isn't worth building at all.


Work with BizFlowAI

If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.

Request a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.

More guides like this on the BizFlowAI blog.

Frequently asked questions

What is Nous Hermes Agent and is it free to use?

Hermes Agent is an open-source AI agent from Nous Research, released in February 2026 under the MIT license. It offers persistent memory across sessions and self-generated reusable skills. Self-hosting is free, while hosted Nous Portal tiers reportedly run from $20 to $200 per month. A Business tier adds shared credits, per-member spending caps and shared skills, and an Enterprise tier adds on-prem deployment, SSO and SLAs.

Why do so many AI agent projects fail to reach production?

Gartner predicted in June 2025 that over 40% of agentic AI projects will be canceled by the end of 2027. The cited reasons are escalating costs, unclear business value and inadequate risk controls, not weak models. A demo agent is just a model, a prompt and a happy path. A production agent also needs permissions, approval gates, failure handling, cost caps, persistent state and audit trails.

How do I limit what an AI agent can do in my business systems?

Enforce limits with tool access rather than prompt instructions. With MCP servers, use include/exclude lists so the agent only sees the specific tools you expose, such as list and draft actions but not send or delete. An agent cannot call a tool it cannot see, even if a confused user or a prompt injection in an email tells it to. Start with a tight whitelist for every integration and expand only after watching the agent run for a couple of weeks.

Which AI agent actions should require human approval?

Use three tiers. Let the agent act automatically on read-only and reversible tasks such as searching, summarizing, drafting and tagging. Require human approval for anything that leaves your organization or moves money, such as sending emails, issuing invoices or charging cards. Block destructive or irreversible actions entirely, such as bulk deletes or changing credentials. Gate only the few actions that matter, because too many prompts lead people to approve reflexively.

How should I handle MCP server failures and timeouts in an AI agent?

Treat every MCP server as a dependency that can hang, expire tokens or go down. Set timeouts well below long defaults for customer-facing workflows, since Hermes defaults to 300 seconds per tool call. Decide in advance what happens on failure: retry once, then escalate to a human and never loop. Prefer scoped OAuth over full-access API keys, and make expired tokens fail loudly rather than silently doing nothing.