Cloudflare Clef: A 39ms Decision Isn’t an Approval

Abstract tech illustration: Cloudflare Clef: A 39ms Decision Isn’t an Approval

If your inbound leads land in Gmail and someone on your team still reads each one, updates the CRM, and decides who replies, a model that classifies in milliseconds sounds like the missing piece. It is part of the answer. But a fast label is not permission to send an email or change a record, and the gap between those two things is where small-business automations break.

Here's exactly how I'd test Clef on a lead-routing workflow before letting it touch anything real.

What Cloudflare actually released

Cloudflare released Clef and Clef-flash on October 1, 2026. They're the first models trained in-house by the Workers AI team, and they're hosted on Workers AI (changelog). Clef is a 27B model and Clef-flash is a 9B model, exposed as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. Both are open-sourced on Hugging Face under Apache 2.0, and Cloudflare says the API is compatible with TypeSafe AI's Jev (launch post).

The important part is what a decision model is. It doesn't write text. It reads an input state plus a set of typed questions, then returns a probability for every allowed answer. Your code gets a structured result it can branch on: route the ticket, escalate, or defer to a human.

That matters for the headline claim. The launch post says a human does not necessarily need to be in the loop, and that agents can act on their own or defer to a human when needed. Press coverage compressed that into "humans no longer need to be in the loop." Those are different statements. The first is about what the model makes possible. The second sounds like a safety conclusion nobody has demonstrated for your business.

One more caveat: Cloudflare's documented examples are support triage (is this ticket urgent, which team owns it) and agent guardrails (a "should I take this action?" check before calling a tool). I found no Cloudflare source that says Clef is suited to lead routing, sending emails, or editing CRM records. The lead-routing workflow in this post is my scenario, not theirs.

What "39 milliseconds" does and doesn't mean

The 39 ms figure is a median from Cloudflare's own benchmarks, not a guarantee. Here are the numbers from their latency table (launch post):

Model Median latency p95 latency
Clef-flash 38.8 ms 122.4 ms
Clef 209.3 ms 238.6 ms
Jev 524.1 ms 536.0 ms

Cloudflare reports these across 43 benchmark runs and says Clef-flash is 13x faster than Jev at the median, with Clef at 2.5x (changelog).

Three things to keep in mind before you quote "39 ms" to anyone:

  • It's a median. Clef-flash's p95 is about three times higher. If your workflow runs per email, a tail latency of ~122 ms is still small, but it's the honest number.
  • It's the model call only. Network time, your prompt length, the Gmail fetch, and the CRM write are not in that figure. In a lead workflow, the CRM API will almost certainly dominate total time.
  • It's vendor-reported. Cloudflare's numbers are vendor figures; independent measurements are rare, small, and have produced different values, so verify on your own traffic. One developer's independent benchmark reported a median of roughly 315 ms for Clef-flash running locally on a consumer GPU, with Jev at roughly 225 ms over the network, and another write-up reported roughly 661 ms during the beta. Those setups differ from Cloudflare's, which is exactly why you should measure your own.

For lead follow-up, speed was never really the bottleneck. A human reads a lead in minutes or hours. Going from "an LLM call takes a couple of seconds" to "a decision takes 39 ms" doesn't change your response time in any way a customer will notice. What changes is that the output is structured, which is the real value.

Where the fast model is weaker

The speed comes with trade-offs that show up in Cloudflare's own tables, as reported by The Decoder and DEV Community:

  • Decision Index (self-reported): Clef scores 61.2, Clef-flash 57.1, and Jev 57.9. So Clef-flash roughly matches Jev's accuracy at a fraction of the latency, per Cloudflare.
  • When2Call: Clef-flash scores 65.58, versus 72.37 for Clef and 80.97 for Jev. On this benchmark, Jev beats both Clef models.
  • CLINC150 with an out-of-scope category: Clef scores 97.43, Clef-flash 66.77. The fast model is much weaker at recognizing inputs that don't belong in any of your categories.

That last one is the critical one for inbound email. Real inboxes are full of out-of-scope messages: vendor pitches, forwarded newsletters, a customer asking about an invoice sent to your sales address. A model that forces everything into your four labels and does it confidently is more dangerous than a slow one.

Also note that all of these Jev comparisons are Cloudflare's own. According to the Hugging Face model card, the Jev figures come from Cloudflare's internal run of the Decision Index 0.2.1 package, with the workflow evals drawn from Typesafe Evals, not from an independent party. These are vendor claims, not independent results. And "without hallucinations," the marketing line for decision models like Jev, only means the model stays within predefined answer options. It doesn't stop the model from choosing the wrong option.

The three-part design: decide, act, verify

The system I'd build separates three jobs that people tend to collapse into one:

  1. Decide: the model returns probabilities over a fixed label set.
  2. Act: deterministic rules map labels (and thresholds) to allowed actions.
  3. Verify: the integration records what actually happened.

Here's a sketch of the middle layer. The point isn't the specific code, it's that the model never holds the keys:

ALLOWED = {
    "spam":              {"action": "tag_spam",        "min_p": 0.95, "auto": True},
    "qualified_lead":    {"action": "assign_and_draft", "min_p": 0.85, "auto": False},
    "existing_customer": {"action": "route_to_support", "min_p": 0.90, "auto": False},
    "needs_review":      {"action": "queue_for_human",  "min_p": 0.00, "auto": True},
}

def route(probs: dict, message_meta: dict) -> dict:
    label, p = max(probs.items(), key=lambda kv: kv[1])
    rule = ALLOWED[label]

    # Hard overrides: topics where a wrong action has a real cost
    if message_meta.get("mentions_contract_or_payment"):
        return {"label": "needs_review", "reason": "high_cost_topic"}

    if p < rule["min_p"]:
        return {"label": "needs_review", "reason": f"low_confidence:{label}:{p:.2f}"}

    return {"label": label, "action": rule["action"], "auto": rule["auto"], "p": p}

Notice what's happening. A clear spam message with high probability gets tagged automatically. A qualified lead gets assigned to a person with a draft reply, not a sent one (auto: False). Anything touching a contract, payment, or ambiguous customer identity goes to a human regardless of what the model says.

The verify step is the part most demos skip:

def act_and_verify(decision, lead):
    result = crm.add_tag(lead.id, decision["label"])   # lowest-risk action
    confirmed = crm.get(lead.id).tags
    log.append({
        "lead": lead.id,
        "decision": decision,
        "crm_write_ok": decision["label"] in confirmed,
        "ts": now(),
    })
    if not decision["label"] in confirmed:
        alert("CRM write did not stick", lead.id)

A 39 ms classification tells you nothing about whether that write succeeded or whether an email reached the right person. Logging the outcome is what lets you trust the system later.

One caution on probabilities: Cloudflare describes calibrated probabilities, but whether they're well calibrated on your leads is unverified. One source describes a confidence field that measures how concentrated the probabilities are rather than how likely the answer is to be correct, so don't read a high value as a guarantee of accuracy. Don't treat 0.85 as meaning "right 85% of the time" until your own test says so. The thresholds above are placeholders you set from data.

How to test it with your own data

You can run this without replacing your stack. Here's the test I'd do, in order:

  1. Export 50 recent inbound messages. Strip personal information you don't need.
  2. Label each with what your team actually did. Not what you think the model should say. What happened.
  3. Define four classifications plus an explicit "needs review". Qualified lead, existing customer, spam, other-or-unclear, and review.
  4. Run all 50 through candidate models. Clef, Clef-flash, and whatever you use now. Cloudflare's pricing page lists per-million-token rates for both models, with no output-token billing since the models generate no text. Prices can change, so check the page before you budget. For a short email, the cost per decision is tiny. Cost is not the constraint here, accuracy is.
  5. Count disagreements by type, not as one accuracy number.

That last step is where the real insight comes from. A missed sales lead and a spam message escalated to a human have completely different costs. Build a small confusion table:

Your team chose Model said Cost
qualified_lead spam High: lost revenue
spam qualified_lead Low: a human glances at it
existing_customer qualified_lead Medium: wrong team, delayed reply
anything needs_review Low: extra human time

If the model's worst errors are in the high-cost row, that tells you what to gate. Include some deliberately out-of-scope messages in your 50, given the CLINC150 gap above.

One independent test with 42 decisions (Clef-flash 66.7% vs Jev 71.4%) was published by a developer on DEV Community; a sample that small is too small for a serious conclusion either way. Your 50 messages are worth more to you than any of the published numbers, because they're your distribution.

Also note that fine-tuning Clef on your own data isn't self-service yet. Cloudflare's forward deployed engineers handle it with customers first, with a self-service platform planned for later (The Decoder). For a small team, assume you're using the off-the-shelf models.

Earn autonomy one action at a time

Once the test results are in, don't connect everything at once. Order actions by cost of being wrong:

  • Start with tagging. Adding a CRM tag is reversible and invisible to customers. Wire this first and watch it for a couple of weeks.
  • Then assignment and drafts. Route to a person with a draft reply they approve. The model saves reading time; a human still owns the send.
  • Keep record changes and billing behind approval. Anything affecting invoices, contracts, or customer identity stays gated until your logs justify more.

A commentary piece on Clef frames this well: the human role shifts from reviewing each case to setting the probability threshold at which software may act alone, and that's a policy decision about risk (Adyog Pulse). That's the right mental model. The threshold is your call, informed by your own error costs.

A note on the sources: Cloudflare's own internal uses of fine-tuned Clef (reviewing abuse reports, sorting support requests, separating good bots from bad) are described as planned uses, not reported results. Don't read them as proof it works for your inbox.

My take: "no humans in the loop" is the wrong goal. The useful goal is no humans stuck checking routine decisions that have already been tested. Clef could make the decision step faster and easier to structure. Whether it saves your team hours depends on which cases it gets wrong and the handoffs you build around it.

Why bizflowai.io helps with this

The hard part of this workflow is rarely the classification model; it's the plumbing around it. I build practical AI automations for solopreneurs and small teams, and the pattern I'd apply here is simple: a decision step feeds a defined action, approval gates sit where a wrong move costs money, and a log records what actually happened. If you want to run the 50-message test and wire only the low-risk actions first, that's the kind of build I do.


Want more like this?

I publish practical AI automation and GenAI engineering content every week.

Planning an AI automation project or need a second opinion on your architecture?

Connect with me on LinkedIn: Lazar Milićević, senior engineer building practical AI automation for solopreneurs and small teams.

Visit bizflowai.io for our services, case studies, and AI consulting.

Frequently asked questions

What are Cloudflare's Clef and Clef-flash models?

Clef and Clef-flash are Cloudflare models built to make structured decisions, such as returning a classification label, without generating a written response. Both are built on Qwen and licensed under Apache 2.0. According to the report, Clef-flash delivers classifications in about 39 milliseconds, more than ten times faster than TypeSafe AI's Jev decision model.

How do I use a structured decision model to handle inbound leads?

Split the workflow into three parts: decide, act, and verify. The model proposes a narrow label, such as qualified lead, existing customer, spam, or needs review. Rules determine which labels may trigger an action. The integration then records what actually happened. Contracts, payments, or ambiguous customer identity should go to human review instead of triggering automatic actions.

How do I test a decision model like Clef on my own messages?

Export 50 recent inbound messages, remove unneeded personal information, and label each with the outcome your team actually chose. Define four allowed classifications plus an explicit "needs review" option. Run the messages through the candidate model and count disagreements by type, since a missed sales lead and a spam message sent to a human have different costs. Then connect only a low-risk action, like a CRM tag.

Why doesn't a fast classification speed prove a decision model is safe to automate?

Speed only shows how quickly a label is returned. The reported 39-millisecond figure does not show how often the model picks the correct label on your data, whether a CRM update succeeded, or whether an email reached the right person. Accuracy, action success, and verification must be measured separately before granting more autonomy.

Should AI agents run with no humans in the loop?

Cloudflare's headline claim is that Clef-style models can remove the need for humans in the loop, but that describes what the models could enable, not proof every business process is safe without review. A better goal is removing humans from routine decisions that have already been tested, while keeping review for cases where a wrong action has real cost.