Consumer AI's Cost Problem: Build for the Workflow Instead

Abstract tech illustration: Consumer AI's Cost Problem: Build for the Workflow Instead

You run a small business, and you're looking at AI to clear email, chase invoices, and follow up on leads. You open a chat window, paste in an email, get a decent answer, and then still copy the result into three other tools by hand. A polished answer is not a finished job, and nobody tells you what a finished job costs or who fixes it when it's wrong.

TechCrunch just published a piece on exactly the economics underneath this, and the lesson for a solopreneur isn't the one the headline suggests. Here's exactly how I think about turning it into a test you can run on your own process.

What TechCrunch Actually Reported (and What It Didn't)

TechCrunch's "The ugly economics of consumer AI," published 30 September 2026, argues that frontier labs have become cautious about consumer AI products, and that the caution isn't a verdict on whether the technology works. The problem is cost: per TechCrunch, AI is unusually expensive to operate, and even hundreds of millions of paying customers don't guarantee break-even. The piece also says the industry is shifting toward the Anthropic model: enterprise contracts and vertical-by-vertical expansion. (source)

The adoption numbers are small. Citing the a16z State of Markets report, which draws on PNC data, TechCrunch says 2.2% of consumers were paying for AI as of May, at an average of $31 a month. The sources disagree on dates and exact percentages (a16z's own text says barely about 2% of US households as of April, and TechCrunch cites Bank of America at roughly 3% in March), so treat it as "roughly 2%, give or take a point" rather than a precise figure. (a16z)

What the piece does not give you is a cost-per-job number for any business process. I'm not going to invent one. That gap is the point: if the labs themselves find per-user serving costs hard to pin down, you shouldn't trust a vendor demo to tell you what your invoice workflow will cost.

Why a Chat Window Is the Expensive Way to Buy AI

A chat interface is open-ended by design. Every message can wander, every answer can trigger a follow-up, and nothing in the interaction defines "done." That is cheap to demo and hard to cost.

A workflow flips that. It has a trigger and a deliverable:

  • Trigger: a new invoice email lands in a specific Gmail label.
  • Deliverable: a prepared accounting entry, or a named exception assigned to a human.

Between those two points, the model does a bounded job: extract fields, compare them to the customer record and purchase order, flag conflicts. You are no longer paying for a conversation. You are paying for a fixed-shape task you can measure, retry, and cap.

This is the same logic behind the enterprise-first shift TechCrunch describes. Enterprises buy outcomes inside a defined process, not unlimited chat. A five-person company can buy the same way.

The Model Is Only One Line on the Bill

When people ask what AI will cost, they price the tokens. Tokens matter, but they are one cost among several. For a completed job, count all of these:

  • Model calls, including retries when output fails validation
  • Tool fees on top of tokens
  • Integration and hosting costs
  • Human review time on flagged cases
  • The cost of a mistake that slips through

A wrong invoice number can become an audit problem even if the model call cost a fraction of a cent. And a pricier call can be the cheaper choice if it prevents several manual handoffs. The unit you want is cost per completed, checked job.

Tool fees are easy to miss. OpenAI's pricing page, for example, lists built-in web search at $10.00 per 1,000 calls, with search content tokens billed at model rates on top. Tool charges add to token costs rather than replacing them. (OpenAI pricing)

Anthropic does something similar for hosted agents: Claude Managed Agents are billed at standard token rates plus $0.08 per session-hour of active runtime. (Anthropic pricing) If your workflow keeps an agent session open for a long time waiting on slow steps, the runtime line starts to matter.

The Levers That Move Cost, Using Published Rates

You can't compute your cost per job from memory, but you can see which levers exist. These are the published rates from the vendors' own pricing pages, per million tokens (input / output):

Model Standard Batch / Flex
Claude Haiku 4.5 $1 / $5 $0.50 / $2.50 (Batch)
Claude Sonnet 5.5 $2 / $10 $1 / $5 (Batch)
Claude Opus 5.5 $4 / $20 $2 / $10 (Batch)
gpt-6.1-sol $2 / $10 $1 / $5 (Flex)
gpt-6-luna $0.10 / $0.50 $0.05 / $0.25 (Flex)

Sources: Anthropic pricing docs, OpenAI pricing. Prices change often; check the official pages before you budget.

Three things I take from this table when designing a workflow:

  1. Tiering is the biggest lever. The spread between a small and a large model is large. Most steps in an invoice workflow (classify, extract, compare) don't need the top tier. Reserve the expensive model for the ambiguous cases.
  2. Async work is half price. Anthropic's Batch API is half the standard rate for work that doesn't need an immediate answer, and OpenAI's Flex tier is similarly cheaper than Standard. A nightly lead-follow-up sweep doesn't need a real-time response.
  3. Repeated context is cheaper to re-read. Anthropic's prompt-caching read rate is $0.10 per million tokens for Haiku 4.5 and $0.20 for Sonnet 5.5. If every run starts with the same long instructions and customer-record schema, caching that prefix is real money at volume.

Watch for the opposite trap too: OpenAI charges more for long-context requests. For gpt-6.1-sol the long-context rate is $4 input / $15 output per million tokens versus $2 / $10 short-context. Stuffing an entire email thread plus the full customer history into every call can double your input rate. Pass in only what the step needs.

One more: US-only inference on the Claude API is priced at 1.1x for input and output tokens. If a data-residency requirement applies to you, bake that multiplier into your math from day one.

The Twenty-Case Test

Here is the test I'd run before committing to any AI workflow. It takes a couple of weeks of ordinary work, not a project plan.

Step 1: Log twenty real cases of one recurring workflow. Pick something concrete: inbound invoices, lead replies, or booking requests. For each case, record:

case_id: 07
trigger: "Invoice email from supplier, PDF attached"
human_minutes: 6
tools_touched: [Gmail, accounting software, CRM]
final_deliverable: "Bill entered, matched to PO"
judgment_points:
  - "PO number missing from PDF"
  - "Amount differs from PO by a line item"

The judgment points are the most valuable column. Every place a human had to decide something is a place the automation needs a rule, a review queue, or a stop.

Step 2: Prototype only the repeatable path. Don't automate the weird cases. Build the path that twelve of your twenty cases follow, and route everything else to a human.

Step 3: Make the exception path explicit. The system must never silently guess. A routing rule like this is the whole philosophy in a few lines:

def route(extraction):
    required = ["supplier", "invoice_number", "amount", "po_number"]
    missing = [f for f in required if not extraction.get(f)]

    if missing:
        return {"status": "review", "reason": f"missing: {missing}"}

    if not matches_purchase_order(extraction):
        return {"status": "review", "reason": "PO conflict"}

    return {"status": "ready", "entry": build_entry(extraction)}

A "review" result lands in a queue with a named owner. A "ready" result becomes a prepared entry a human approves. Nothing posts on a guess.

Step 4: Meter it per workflow. Anthropic offers a Usage and Cost Admin API with a cost_report endpoint that returns USD cost breakdowns grouped by workspace or description, so you can track spend per workflow rather than per account. (docs) Put each workflow in its own workspace and you get the model-cost line of your cost-per-job math from the vendor directly.

Step 5: Compute cost per completed job. For the twenty cases, add up:

  • Actual model and tool charges
  • Retries
  • Minutes of human review, valued at your or your staff's hourly cost
  • Corrections after the fact

Compare that to the original human minutes. If the automated path plus review isn't clearly cheaper, you've learned that before building anything permanent. If it is, you now have real numbers from your own operation instead of a savings estimate from a demo.

What Breaks This Math

A few failure modes I'd watch for, because they quietly erase savings:

  • Review that never shrinks. If a human has to re-check every output because trust never builds, you've added a step instead of removing one. Track the review rate over time; it should fall.
  • Retries that hide in the average. A cheap model that fails validation 30% of the time and gets retried may cost more than a mid-tier model that passes the first time. Log retries separately.
  • Scope creep into chat. The moment someone adds "and also answer questions about this email," you've reintroduced the open-ended conversation you were avoiding. Keep the deliverable fixed.
  • Treating one good week as a trend. Twenty cases is a sample, not a guarantee. Re-measure after a month of real volume.

My hot take: the strongest small-business AI products won't look like consumer chat apps. They'll look like boring, dependable handoffs between tools that don't talk to each other today. The interface matters less than whether the work gets completed, checked, and recorded at a cost you can sustain. That's also, not coincidentally, the shape of product that survives the economics TechCrunch is describing.

Where to Go From Here

If you'd rather not build this yourself, bizflowai.io is where I publish my work on practical AI automation for solopreneurs and small teams. Whatever route you take, the test above is the same: price the completed, checked job, not the prompt.


Want more like this?

I'm Lazar Milićević, a senior engineer who builds AI automations that actually ship, and I publish practical AI automation for small teams.

Planning an AI automation project or need a second opinion on your architecture?

Visit bizflowai.io for our services, case studies, and AI consulting.

Frequently asked questions

Why does the economics of consumer AI matter?

A TechCrunch analysis, "The ugly economics of consumer AI," argues that frontier labs' caution around consumer AI isn't just a verdict on whether the technology works. The cost of serving consumers matters too. For founders, the lesson is to ask whether you can afford every step between a request and a correct outcome, not just whether the AI gives good answers.

How do I test whether AI is affordable for a recurring business workflow?

Pick one recurring workflow and log the next twenty real cases. For each, record the trigger, minutes of human work, tools touched, final deliverable, and every judgment point. Then prototype only the repeatable path and measure actual model and tool charges, review time, retries, and corrections against the original manual time. This gives you real numbers from your own operation instead of a demo-based savings estimate.

What costs should I count when budgeting for AI in business workflows?

The model call is only one cost. You should also count integration fees, retries, human review, support, and mistakes. A wrong invoice number, for example, can create an audit problem even if the AI call was cheap. Judge the cost of the completed job, not the price of a single prompt. A pricier model call can be worth it if it reliably prevents several manual handoffs.

How should I design an AI workflow so it finishes the job instead of just chatting?

Don't pay a model to hold open-ended conversations about every email. Give the workflow a specific trigger and a specific deliverable. For an invoice email, the system extracts fields, checks them against the customer record and purchase order, prepares an accounting entry, and asks a human to review missing or conflicting information. Success is a checked entry or a clearly assigned exception.

What should an AI workflow do with ambiguous cases?

Ambiguous cases should go into a review queue for a human rather than letting the system silently guess. In an invoice workflow, missing or conflicting information is flagged for human review. A successful run ends with either a checked entry or a clearly assigned exception, which avoids silent errors such as a wrong invoice number causing an audit problem later.