n8n Email Graph: 3 Routes, Zero Auto-Replies

Abstract tech illustration: n8n Email Graph: 3 Routes, Zero Auto-Replies

One test email arrives. n8n reads it, picks a route, and Telegram pings you. If Gmail eats two hours of your day, that's the piece you want back — but the interesting engineering isn't "add an AI agent." It's the connections between nodes, and what happens when the model guesses wrong.

Most tutorials stop when a classifier spits out a label. They never show the fallback path when the JSON is malformed, the API times out, or the model returns "maybe_action" instead of one of your two allowed labels. This post walks through the full graph — Gmail in, one deterministic gate, one AI classifier, three Telegram routes — including the human-review lane that catches everything the model gets wrong. Zero customer replies. Zero silent invoice filing.

The shape of the graph (and why it's not one big agent)

Here's the whole thing in one picture:

Gmail Trigger
   └─> Normalize Email (Edit Fields)
          └─> If: finance keywords?
                 ├─ yes ─> Telegram: "Review finance email"
                 └─ no  ─> HTTP Request: classify
                             ├─ success ─> Code: validate JSON ─> Switch(label)
                             │                                     ├─ action  -> Telegram: "Reply needed"
                             │                                     ├─ routine -> (end)
                             │                                     └─ review  -> Telegram: "Needs review"
                             └─ error   ─> Edit Fields: label=review -> Telegram: "Needs review"

Five real decisions, three exit points, one hard rule: AI never touches finance, and AI never sends anything to a customer. Its only job is to sort ordinary messages into action or routine. If it fumbles, the message goes to a human.

I've seen too many "AI inbox assistant" builds collapse the whole thing into one super-node that fetches, classifies, drafts, and sends. That works in a demo. In production, one bad classification autoreplies to a lawyer, and you're rewriting your automation policy on a Saturday.

Gmail Trigger and the Normalize Email node

Start with the Gmail Trigger. Connect a test Gmail account (not your main one for the first week), watch incoming messages, and fire the test event so you can inspect a real payload before wiring anything downstream.

Send yourself an email:

  • Subject: Can we discuss a project?
  • Body: Hi, would you have 20 minutes this week for a quick call? Thanks.

Open the trigger output and find four fields: sender, subject, the plain-text body or preview, and message_id (Gmail calls this id). Field names shift a little between n8n versions and between Gmail's messages.get shapes, so use the field picker against your actual test item — do not copy an expression blindly from a screenshot.

Add an Edit Fields node right after, name it Normalize Email, and produce exactly four keys:

{
  "message_id": "{{ $json.id }}",
  "sender":     "{{ $json.from.value[0].address }}",
  "subject":    "{{ $json.subject }}",
  "excerpt":    "{{ ($json.text || $json.snippet || '').substring(0, 500) }}"
}

Why 500 characters? Two reasons. First, the classifier doesn't need a whole thread to tell "can we jump on a call" from a newsletter — subject plus the first paragraph is enough. Second, token cost is a real line item. A 500-char excerpt is roughly 120-150 tokens; a full email with quoted history is often 3,000+. On a cheap classifier model that's the difference between $0.0001 and $0.002 per email. At 200 emails a day, that's $12/month vs $0.60/month for the same accuracy.

Run the node and check every field. If sender is blank, fix it here — every downstream branch depends on this shape.

Fields that must be non-null before you continue

  • sender — used in every Telegram notification
  • subject — used in the finance rule and the classifier prompt
  • excerpt — used in the classifier prompt
  • message_id — for future dedup, threading, or Gmail label-write callbacks

The finance gate: deterministic, not clever

Add an If node before you go anywhere near AI. Check whether subject or excerpt contains any of your finance keywords. For a US SMB running QuickBooks or Xero, a workable starter list:

invoice, invoice #, receipt, payment, wire transfer, ACH,
routing number, past due, remittance, purchase order, PO #,
statement, W-9, 1099

Matches go straight to Telegram with the message "Review finance email" plus sender and subject. Do not forward the body or attachments. A human opens the original email in Gmail and cross-checks it in the accounting system.

Why deterministic here? Because the failure modes are asymmetric. A missed newsletter costs nothing. A misrouted invoice — silently labeled "routine" and never paid, or worse, auto-acknowledged and marked as handled — can turn into a late fee, a broken vendor relationship, or an audit finding. Keyword rules will over-catch (some harmless emails will trip them) and under-catch (a vendor writing "our records show an outstanding balance" without the word "invoice"). That's fine. This gate is a conservative first filter, not an accounting control. The control is the human reading the original email.

If you're in a regulated context — bookkeeping for clients, anything touching payroll — this is where you also write to an audit log node (Postgres, Airtable, or an append-only Google Sheet) so every finance-flagged email leaves a trail.

The classifier: two labels, JSON only, own the failure path

Everything that isn't finance goes to an HTTP Request node. Use your AI provider's API through an n8n credential, never a key pasted into the node. The prompt is small on purpose:

You are an email triage classifier. Read the subject and excerpt.
Return a JSON object with exactly one key: "label".
Allowed values:
  - "action"  : the sender expects a human response (question, meeting request, lead, complaint)
  - "routine" : informational only (newsletter, receipt confirmation, notification, digest)

Return only the JSON. No prose.

Subject: {{ $json.subject }}
Excerpt: {{ $json.excerpt }}

Two labels, not five. Every extra label roughly doubles the cases you need to test and halves your confidence in each one. Two is the floor that actually reduces inbox load: "the sender is waiting on me" vs "the sender is not."

Now the part most tutorials skip. Add a Code node after the HTTP Request to validate the response:

const raw = $input.first().json.choices?.[0]?.message?.content
         ?? $input.first().json.content?.[0]?.text
         ?? '';

let label = 'review';
try {
  const parsed = JSON.parse(raw.trim());
  if (parsed.label === 'action' || parsed.label === 'routine') {
    label = parsed.label;
  }
} catch (e) {
  label = 'review';
}

return [{
  json: {
    label,
    sender:  $('Normalize Email').item.json.sender,
    subject: $('Normalize Email').item.json.subject,
  }
}];

Three things this does that a naive setup does not:

  1. Whitelist labels. Anything other than action or routine becomes review. If the model returns "Action", "action_needed", or a helpful paragraph explaining its reasoning, it goes to a human.
  2. Catch parse errors. Malformed JSON → review. Empty response → review.
  3. Re-attach identity fields. The HTTP node's output does not carry sender and subject. I pull them from the Normalize Email node by name so the Telegram notification always knows who the email was from.

Also configure the HTTP Request node's error output (settings → "Continue on fail" or the dedicated error branch, depending on your n8n version) and route it into a small Edit Fields node that sets label = review and re-attaches sender/subject the same way. Rate limits, 5xx from the provider, network blips — none of them should silently drop an email.

Switch node and the three Telegram routes

A Switch node on label gives you three exits:

Label Route Telegram message
action Ops chat Reply needed — <sender> — <subject>
routine (none) no notification
review Ops chat Classification needs review — <sender> — <subject>

routine deliberately ends without a ping. If the whole point is to reclaim attention, don't buzz the phone for a Stripe receipt. The reader can still find those messages in Gmail with a saved search.

For Telegram setup: create a bot via BotFather, store the token as an n8n credential (not in the node), start a chat with the bot, and copy that chat's numeric ID into the Telegram nodes. Send one harmless test notification first — "n8n hello" — before touching the classifier. If Telegram rejects it, the problem is the chat ID or the bot never being /start-ed, not your prompt.

Keep tokens out of workflow exports. When you share a .json of the graph, credentials should serialize as references, not values. Check the export before you send it to anyone.

What each route actually costs you in attention

  • action — one notification, one glance, one decision (reply now, snooze, delegate)
  • routine — zero notifications, zero attention
  • review — one notification, but you open the email and possibly correct the classifier's behavior over time (keyword additions, prompt tweaks)

Track the review rate for two weeks. If it sits above ~10%, your prompt is under-specified or your two labels are too broad for your actual inbox. Fix the prompt, not the graph.

Testing the graph before you point it at your real inbox

Send five test emails through the test Gmail account, one for each expected path:

  1. Project inquiry — "Can we discuss a project?" → expect action
  2. Newsletter — a real newsletter forwarded in → expect routine
  3. Invoice — subject "Invoice #4471 from Acme" → expect finance gate, never reaches AI
  4. Ambiguous — subject "Following up", body "per our conversation" → expect action or review
  5. Deliberately broken classifier — temporarily change the prompt to return prose instead of JSON → expect review via the Code node whitelist

If all five land where you expect, then point the trigger at your real inbox — but keep the classifier's output as notifications only for the first week. No Gmail label writes, no auto-archive, no auto-reply drafts. You want to build trust in the routing before the graph starts modifying anything.

The first week is where you'll find the edge cases: the client who signs off with "invoice attached" as a joke, the vendor whose newsletter subject line reads like a support ticket. Every one of those is a keyword to add or a line to tighten in the classifier prompt.

Why bizflowai.io helps with this

Wiring an n8n graph like this is a weekend if you already know the platform, and a month if you don't — and most of that month is spent on the failure paths, not the happy path. At bizflowai.io I build email triage and operations workflows for small US teams where the deterministic gates, the classifier prompt, the human-review lane, and the audit trail are all sized to what the business actually handles. The interesting work is never "add another AI agent." It's deciding which decisions the AI is allowed to make, and where a person still has to look.


Want more like this?

I publish practical AI automation, GenAI engineering, and faceless content workflows on YouTube every week.

Subscribe to bizflowai.io on YouTube — never miss a new tutorial.

Planning an AI automation project or need a second opinion on your architecture?

Connect with me on LinkedIn — Lazar Milicevic, GenAI Engineer & bizflowai.io Founder.

Visit bizflowai.io for our services, case studies, and AI consulting.

Frequently asked questions

What is the purpose of this n8n Gmail-to-Telegram workflow?

It's a tutorial workflow that helps a small business owner sort incoming Gmail messages by triggering on new emails, normalizing key fields, applying keyword rules for finance messages, and using AI only to classify ordinary emails as action or routine. It never replies to customers, moves money, or files invoices automatically—risky cases are routed to a human via Telegram.

How do I safely handle invoice emails in an n8n automation?

Add an If node after normalization that checks whether the subject or excerpt contains invoice or payment keywords. Route matches directly to a Telegram notification saying 'Review finance email' with sender and subject only—do not forward invoice contents, attachments, or send them to AI. A person then checks the original email and accounting system before anything is recorded or paid.

Why should I normalize email fields before branching in n8n?

Every downstream branch depends on a consistent data shape. Use an Edit Fields node called Normalize Email to extract message_id, sender, subject, and an excerpt limited to roughly 500 characters from the Gmail trigger output. This gives enough context to sort messages without pulling entire threads or attachments, and fixing missing fields here prevents failures in later If, AI, and notification steps.

How do I validate AI classification output in n8n?

Send only the normalized subject and excerpt to your AI provider's HTTP endpoint and require a JSON object with one key, label, restricted to action or routine. Add a Code node that parses the response and accepts it only if label matches exactly. On parsing errors, empty answers, or other labels—and on the HTTP node's error output—set label to review so a human handles ambiguous cases.

When should I use keyword rules vs AI classification for email sorting?

Use conservative keyword rules for high-risk categories like invoices and payments, where a data-entry mistake could cause compliance or audit problems, and always route those to human review. Use AI classification for ordinary messages where phrasing varies widely, such as distinguishing a call request from an informational update. AI provides a useful category but should not be the final authority on financial decisions.