Harness Tuning Took an Agent From 43.5% to 93%

Your agent works in the demo, then fails on the third real task you hand it, and the obvious fix is a bigger model or a fine-tune you can't afford to run. A new Salesforce preprint argues you may be fixing the wrong layer. They froze the model, rewrote everything around it, and watched a pass rate go from 43.5% to 93.0% on a browser benchmark.
This post walks through what they did, which numbers hold up, what they don't prove, and how to borrow the method for your own agents without a research lab.
What Salesforce actually measured (and what the headline hides)
Salesforce AI Research and Salesforce Agentforce built a framework called DarwinX that evolves an agent's harness while the model weights stay frozen. The headline result is real but narrower than "browser tasks" suggests: it is WebArena-Infinity pass@1 on 1,260 real tasks, in the paper's "audit-clean" form.
Here is the nuance, because it matters for how much to trust the number. The authors audited every WebArena-Infinity trajectory for validity and removed invalid ones. Under that audit:
| Agent | Raw score | Audit-clean score |
|---|---|---|
| Base agent | 53.0 | 43.5 |
| DarwinX-evolved agent | 94.4 | 93.0 |
The base agent dropped nearly ten points when invalid runs were removed. The evolved agent barely moved. So the gap widened under scrutiny, from roughly 41 points raw to 49.5 points audit-clean. That is a good sign: an evolved harness that was merely gaming the benchmark would more likely lose ground under an audit, not hold it. According to the paper's audit, invalid trajectories fell from 293 to 17.
Caveats I'd keep in front of you:
- This is a research preprint (arXiv:2608.07545, submitted 31 July 2026) with self-reported results. I haven't seen independent replication.
- The numbers come from benchmarks, not business workflows. A 43.5% to 93% jump is not a typical outcome you should expect on your invoicing agent.
- I'm not going to name the model behind the browser result, because I haven't confirmed which one the paper used for that specific run.
Primary sources: the paper abstract on arXiv, the full HTML version, and the VentureBeat write-up.
The harness is everything except the weights
The harness is the prompts, tools, skills, and workflow code wrapped around a model call. When people say "the agent got better," they often mean the model. In practice most of what you control as an application developer lives in the harness, and DarwinX's whole premise is that this layer is where hosted-model users have leverage.
VentureBeat frames this as relevant to developers who build on hosted models and have no fine-tuning pipeline of their own. That describes almost every solopreneur and small team. You can't touch the weights of a hosted frontier model. You can touch:
- System prompts and instructions: what the agent is told about its job and constraints.
- Tool definitions: which tools exist, how they're described, what they return, how errors surface.
- Skills: reusable procedures the agent can pull in ("verify the output before finishing").
- Workflow structure: planning steps, retries, checkpoints, when to stop.
- Context handling: what gets fed back after each step.
A concrete way to think about it: the model is the engine, and the harness is the drivetrain, steering, and dashboard. Engine swaps are expensive and out of your hands. A lot of reliability problems are really steering and dashboard problems, like the agent not knowing what "done" looks like or not being able to see that its last action silently failed.
Why self-improving harnesses usually fall apart
Letting an agent inspect its own failures and edit its own harness sounds simple, but it tends to stall. Two failure modes show up, and they're the reason DarwinX exists.
Regression. An edit that helps one task makes the agent worse at another. You add a rule to fix a form-filling failure, and now the agent over-checks on simple navigation and times out. If you only measure the task you just fixed, you never notice.
Plateau. If you repeatedly improve a single version of the harness, you walk down one path. Each patch is locally sensible, and the whole thing converges on a local optimum. You've polished one lineage and never explored a different one that might have been better.
Anyone who has maintained a long prompt knows both. The prompt grows by accretion, every incident adds a paragraph, nobody dares delete anything, and quality flatlines while token cost climbs.
How DarwinX avoids both traps
DarwinX borrows from evolution, and the mechanism is more practical than the metaphor sounds. Two ideas do the work.
Preserve-and-extend. According to the paper, DarwinX admits a new harness variant only if it extends coverage without regressing. A change has to solve something new and keep everything the previous version already solved. That directly attacks the regression problem. It is a regression gate, the same discipline you'd apply to code with a test suite.
An archive of lineages. Instead of one harness being patched forever, the system keeps alternative lineages and can recombine them. That attacks the plateau problem, because you're not committed to a single path.
Real verifiers as fitness. Fitness comes from each benchmark's own verifier. There are no gold solutions and no hand-picked winners. The score is whatever the task's checker says, not what the agent claims about itself. This is the piece most home-grown "self-improving" setups skip, and it's the one I'd copy first.
In pseudocode, the shape is:
archive = [baseline_harness] # keep multiple lineages, not one
for generation in range(N):
parent = select(archive)
failures = run_and_collect_failures(parent, task_set)
candidate = propose_edit(parent, failures) # LLM edits prompts/tools/skills
results = run_all(candidate, task_set) # verifier-scored, not self-scored
# preserve-and-extend: must keep old wins AND add something new
if keeps_all_prior_passes(results, parent) and adds_new_passes(results, parent):
archive.append(candidate)
That's a simplification of the idea, not the authors' code. The real implementation is in the open-source Beagle repo.
What the evolved harness actually learned
The most useful detail for builders is what changed. VentureBeat reports the evolved harness added seven skills. They tell the agent to:
- define what a correct result looks like before starting,
- check generated files and values before finishing,
- ground outputs in real tool execution rather than asserting them.
None of that is exotic. It's "know your acceptance criteria, verify before you declare victory, and don't make things up." The gain didn't come from clever new capabilities. It came from encoding basic engineering hygiene into the agent's procedure, and from finding which hygiene steps mattered by letting failures drive the edits.
The compute story is also worth reading carefully. The evolved agent did not just burn more turns everywhere:
| Task group | Median turns, before | Median turns, after |
|---|---|---|
| Tasks both versions already solved | 12 | 13 |
| Six newly solved tasks | 11 | 22 |
On tasks the agent could already handle, behavior barely changed. On the newly solved tasks, it spent about double the turns. That's what you want: extra effort concentrated where it converts failures into passes, not a blanket "think harder" setting. It also suggests a cost-control angle. A harness that spends effort selectively is cheaper than one that over-verifies everything.
Does it generalize? What the other benchmarks say
One benchmark proves little, so the spread of results matters. The paper tested four setups:
| Benchmark | What it tests |
|---|---|
| Terminal-Bench 2.1 | In-domain |
| TerminalWorld | Held-out tasks |
| WebArena-Infinity | Synthetic-to-real transfer |
| Terminal-Bench 2.1 → SWE-bench Verified | Cross-benchmark transfer |
The harness improved on all four, with gains ranging from 3.4 points on SWE-bench Verified to 49.5 points on WebArena-Infinity. The authors summarize a single evolution loop as adding about 17 points on average. I'm quoting that as the paper's own claim, not an independent finding.
Three details I find more informative than the average:
- Synthetic-to-real. The WebArena-Infinity harness was evolved only on synthetic browser intents, then tested on real, unseen tasks. Improvements that survive a distribution shift are more believable than ones tuned on the test set.
- Cross-benchmark transfer. A harness evolved on Terminal-Bench 2.1 transferred unchanged to SWE-bench Verified, going from 80.8 to 84.2 (+3.4), per the Beagle repo. A smaller gain than the browser result, but it came with no re-tuning.
- Strong baselines still improve. On Terminal-Bench 2.1, the Monet agent on a frozen GPT-5.5 went from 75.5% to 83.2%, and the paper reports 84.7% with a stronger base model. Gains shrink as the baseline gets stronger, which is the sober part of this story: the less headroom, the smaller the win.
The pattern across all of it: big gains where the baseline was failing for fixable procedural reasons, modest gains where it was already strong.
How to run a small version of this on your own agent
You don't need an evolutionary framework to get most of the benefit. You need the discipline underneath it. Here's the stripped-down version I'd run on a business agent (an inbox triage agent, a lead-qualification agent, an invoice-extraction agent).
1. Build a task set with a real checker. Collect 30 to 100 real inputs your agent has handled, each with a pass/fail check that doesn't rely on the agent's opinion. For invoice extraction, that's "total matches the PDF's total." For triage, "label matches what a human assigned." No checker, no improvement loop; you're just guessing.
2. Baseline it and read the failures. Run the current agent, score it, and cluster the failures by cause rather than by task. Expect a handful of recurring patterns: skipped verification, misread tool output, premature "done."
3. Make one change per candidate. Add a skill, tighten a tool description, add a verification step. One change, so you know what caused any movement.
4. Gate on no-regression. Rerun the entire set. Accept a change only if it fixes something new and breaks nothing that previously passed. Here's a minimal gate:
def accept(candidate_results: dict, baseline_results: dict) -> bool:
prior_wins = {t for t, ok in baseline_results.items() if ok}
still_wins = {t for t, ok in candidate_results.items() if ok}
new_wins = still_wins - prior_wins
regressions = prior_wins - still_wins
return len(new_wins) > 0 and len(regressions) == 0
5. Keep the old version. Don't overwrite. Tag harness versions so you can branch from an earlier one when a path dead-ends, which is the poor man's archive.
6. Watch cost per task. Log turns and tokens before and after. If median effort jumps on tasks that were already passing, your new skill is over-triggering.
Two honest limits. Strict no-regression gating on a small, noisy task set will sometimes reject good changes because of variance, so rerun borderline cases a few times before judging. And a task set is only as good as its coverage; an agent can pass all 50 of your tests and still fail on input 51.
What this does and doesn't prove
Read this as strong evidence for a direction, not a forecast for your numbers. It shows that on these benchmarks, with a verifier-driven loop and a regression gate, harness changes alone moved pass rates by large amounts, and that the gains held up under an audit and under transfer. It does not show that the same method yields comparable lifts on messy small-business workflows. The paper, per a secondary report from Crypto Briefing, says nothing about production deployment.
My takeaway for practitioners: before you reach for a bigger model, check whether your agent has defined success criteria, verifies its own outputs against real tool results, and is regression-tested when you edit it. Those three are cheap, and this work suggests they can be worth far more than a model upgrade when the baseline is failing for procedural reasons.
How BizFlowAI approaches this
The part of this research that matches how I work is the order of operations. When a client agent underperforms, I don't start by swapping models. I start with a task set drawn from their real inputs, a checker that doesn't trust the agent's self-report, and a baseline number. Then changes to prompts, tool descriptions, and verification steps get accepted only if they don't break what already worked. That's unglamorous, and it's where reliability tends to come from.
I'd be careful not to oversell the parallel: the Salesforce results are benchmark results from a research team, and I'm not claiming client workflows see anything like a 43% to 93% swing. If you want to see what a harness-first reliability pass would look like on an agent you already run, book a BizFlowAI discovery call and we'll look at your failure cases together.
Work with BizFlowAI
If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.
Request a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.
More guides like this on the BizFlowAI blog.
Frequently asked questions
What is an agent harness and why does it matter more than the model?
An agent harness is everything around the model weights: system prompts, tool definitions, skills, workflow code, and how context is fed back after each step. If you build on a hosted model, you cannot change the weights, but you can change the harness. Salesforce's DarwinX research showed that evolving only the harness raised a browser-agent pass rate from 43.5% to 93.0% on 1,260 real tasks with the model frozen.
How can I improve my AI agent's reliability without fine-tuning the model?
Work on the harness instead of the weights. Add explicit acceptance criteria so the agent knows what a correct result looks like, make it verify generated files and values before finishing, and require outputs to be grounded in real tool execution. Score changes with a real verifier rather than the agent's own claims, and only keep a change if it fixes something new without breaking tasks that already passed.
What is DarwinX and how does it evolve an agent harness?
DarwinX is a Salesforce AI Research framework that automatically evolves an agent's prompts, tools, skills, and workflow while the model stays frozen. It proposes edits based on observed failures and admits a new variant only if it preserves all previous passes and adds new ones. It also keeps an archive of alternative harness lineages so it does not plateau on a single path, and it uses each benchmark's own verifier as the fitness signal.
Why do self-improving AI agents regress or plateau?
Regression happens when an edit that fixes one task quietly breaks others, because only the newly fixed task was measured. Plateau happens when you keep patching a single version of the harness, so it converges on a local optimum and never explores better alternatives. Fixes are a regression gate that requires all prior wins to be kept, plus maintaining multiple lineages that can be recombined.
What kinds of skills did the evolved harness add to the agent?
According to reporting on the Salesforce paper, the evolved harness added seven skills that encode basic engineering hygiene. They tell the agent to define what a correct result looks like before starting, check generated files and values before declaring completion, and ground outputs in actual tool execution instead of asserting them. The extra effort was concentrated on newly solved tasks, where median turns roughly doubled from 11 to 22, while already-solved tasks barely changed.