Skip to main content

Choosing metrics to evaluate B2B agents before the pilot

Rafael Torres
Rafael TorresAugust 14, 202614 min. read
Choosing metrics to evaluate B2B agents before the pilot

A company announces it will pilot AI agents, and in the same week four people open a discussion about which model to use, which orchestration framework, which runtime. Nobody asks the one question even once: what does "working" mean, for this operation, for an agent? That question decides the pilot. And it is almost never asked.

The first pilot decision is not technical

The first decision in a B2B agent pilot is choosing what to measure, not which framework or which model. Whoever jumps straight to the technology choice is answering a question that has not been formulated yet, and discovers it at the end of the quarter, with an agent running and no ruler to say whether it was worth it. The cost of the pilot is not the token. It is the test that ends without a number to argue with or against.

Three people feel that cost first: the operations leader, who has to justify the head added or the head that was never added; the RevOps owner, who inherits the agent without knowing which target it runs against; and the technology leader, who becomes the owner of a system whose "success" nobody defined. Those three together are the reason half of pilots die quietly. Not because the agent failed. Because there was never a prior agreement on what would count as success.

There is a number that shows the asymmetry: public agent benchmarks are abundant, but the report that matters to the CFO, the one that connects the agent to the business result, is a ruler the operation itself has to produce. It does not arrive ready-made from any vendor. That gap is the entire subject of this piece, and the reader who ignores it pilots blind.

The most elegant trap is this one. The technical team picks metrics it knows how to measure well: latency, cost per token, task completion rate. The business team picks metrics that sound good in a meeting: "more efficiency," "less rework." The two sets never speak to each other, and the pilot ends with a list of numbers everyone recognizes as true and no one recognizes as decisive. That is the point where the project starts to rot without ever throwing an error.

What a domain benchmark actually proves

A domain benchmark proves that an agent can turn a real, private task into a correct deliverable inside a controlled environment, graded criterion by criterion by an LLM judge. That is an important proof, and it is all it is. It says nothing about what happens when that agent touches the real operation, with dirty data, undocumented exceptions, and a real deadline.

The evaluation class has a specific shape. The agent reads case documents in a sandboxed environment, executes a multi-step task, and produces a real deliverable: a memo, a disclosure schedule, a deposition summary. Not a trivia answer, not a single-turn chat. It is compound work, assessed end to end.

Cite the concrete example. Harvey LAB-AA, Artificial Analysis's implementation of Harvey's Legal Agent Benchmark, runs agents against 120 private legal-practice tasks spread across 24 areas. Each task is graded criterion by criterion by a single LLM judge, a rubric that scores each criterion separately instead of handing out one forgiving overall grade. What is measured there is capability: the distance between the deliverable produced and the expected standard for that specific criterion.

What the Artificial Analysis evaluations list brings together is exactly that family: LAB-AA, AA-Briefcase, AA-Analyst, GDPval-AA v2, Terminal-Bench. Domain benchmarks that prove agents produce real, multi-step deliverables in a sandbox. And they stop there, at the edge of the sandbox.

Here is the confusion that costs real money. An organization reads "the agent won a domain benchmark" and translates it as "the agent will deliver value in my workflow." The translation is wrong. The benchmark measures whether the task was completed within an agreed criterion. It deliberately does not measure whether completing that task that way is worth money to anyone. Two different things, and the second is the one that pays for the project.

Why completed task is not delivered value

A completed task and delivered value are not the same thing, and the distance between them has a name: the benchmark-to-production gap. A benchmark keeps the data clean, the environment controlled, and the grading criterion fixed by a judge. Production throws dirty data, partially integrated systems, and a success criterion that shifts week to week. The gap is not an implementation detail. It is the hole pilots fall into.

The benchmark assumes three conditions that the real operation breaks one by one. First, clean input: the document the agent reads in the benchmark is complete and well labeled. In production, the same document arrives half-missing, in twenty formats, with an attachment nobody asked for. Second, known scope: every benchmark task has a defined start and end. In production, the work arrives tangled, with dependencies nobody listed. Third, stable criteria: the LLM judge uses the same rubric on every task. In production, what counts as "good" changes when the quarter's target changes.

The example makes it concrete. A legal agent evaluated and validated to draft memos can score high on LAB-AA because the benchmark hands it a clean document set and a well-bounded request. In a real firm, the same agent receives an email with ten attachments, a client who changed direction midstream, and a partner who wants the answer "with more context." The capability is there, proven by the benchmark. The value is not. It depends on factors the benchmark never promised to cover.

There is a second hidden cost in this confusion, and it is accounting in nature. When the pilot's ruler inherits the benchmark's ruler, you measure the average task accuracy and call it return. What gets left out is the part of the operation the agent did not touch, the part that needed human repair, the part that generated silent rework. Exactly the part that decides whether the agent saves a head or merely moves one around.

That is why the first ruler alone, the capability one, produces a number everyone praises and no one uses to decide anything. It answers "does the agent know how?" The question the CFO asks is another one: "how much does the company gain per week with the agent doing it?" The two demand different instruments, architected at different moments, and almost no one separates the two.

How to assemble the ruler before the pilot

Assembling the ruler before the pilot means stacking two different layers of measurement and then measuring the distance between them. The first layer is domain capability, gauged by a benchmark of the kind Artificial Analysis documents. The second is the business value of your own operation, defined by the metric that matters for that specific routine. The pilot starts only when both exist and the gap between them is known.

The order matters more than the content of each layer. Whoever assembles the value ruler after running the agent is measuring what already happened, without having established beforehand against what. It is like choosing the sales target at the end of the month: every conclusion will be retrospective and no decision will have become easier. The ruler has to exist before, and be written somewhere the team can dispute.

The table below draws the separation between the two metric types, because it is the boundary that decides almost everything:

DimensionCapability metric (benchmark)Value metric (business)
What it measuresWhether the agent completed the task against an agreed criterion, in sandbox, with clean dataWhether having that work done produced measurable results in the real operation
What it does NOT sayNothing about the business result, the residual rework, or the cost of integrating the agent into the flowNothing about the isolated technical quality of the work, outside the effect it produced
When to use itTo select and calibrate the agent's capability before it touches the operationTo decide whether the agent stays, scales, or dies, at the end of the pilot

The assembly sequence has four steps, and none of them is optional. First, establish the capability baseline using a domain benchmark relevant to the area. Second, write the business-value ruler for your own operation: cycle time, rework, cost per delivery, time to response. Third, measure the gap between the two layers: where the agent is capable but the real flow knocks it down, that is where the integration project lives. Fourth, fix the pilot's acceptance criteria on top of that gap, with a number, an owner, and a date.

The third step is what separates a mature pilot from a curiosity test. High capability and low value, in the same agent, point to the problem being in the integration, not the technology: the agent knows how, it just does not reach the place where it knows. Low capability and high value point to the problem being in the agent, and no process tweak will save it. Without measuring both layers, the team confuses the two diagnoses and fixes the wrong side.

It is in this drawing, and not in any leaderboard number, that Nexforce Agents enters. The pilot's ruler is defined before any execution, and the execution happens in a sandboxed, traceable way, via Nexforce Work and Nexforce Code. What the benchmark documents as capability, Nexforce Agents lets run inside the operation with a record of what happened at each step, so the gap between the two layers shows up as data and not as suspicion.

inline-01.png

The strongest counterargument

The strongest counterpoint is simple: benchmarks suffice, because business value is too complicated to measure rigorously, so the best use of time is to run the agent and see what happens. It is an honest argument, defended by serious people, and it has real appeal when the quarter is tight and nobody wants to stop the project to draw a ruler. It collapses on a bill the proposal itself hides.

The reasoning behind it is: measuring capability is cheap, measuring value is expensive, so measuring only capability is the rational path. The flaw is in the premise that failing to measure value costs zero. Failing to measure value postpones the cost to the end of the pilot, when the decision to scale, keep, or kill the agent has to be made on the basis of a number that does not exist. The cost did not disappear. It was pushed to the point where it hurts most.

There is a second layer to the counterargument, and it is the more seductive one. "Big companies decided on the basis of benchmarks and it worked." That confuses correlation with capital burn. An organization with a comfortable budget can pilot blind and survive, because the error is absorbed by the rest of the operation. The average company running B2B agents has no such cushion, and it is for them the ruler matters. For whoever has money to spare, any evaluation methodology looks like bureaucracy; for whoever does not, it is what keeps the test from turning into a loss.

The definitive answer to the counterargument is empirical, and it has already been said here: the gap between capability and value does not describe a measurement difficulty, it describes the very work of integrating the agent into the operation. Whoever refuses to measure the gap is not saving the effort of evaluating. They are choosing not to know where the agent will fail, and to find out the worst way, with the system in production and the outcome in the quarter's result.

What changes for those adopting agents today

The belief that needs to change is a single one: stop piloting blind. Whoever adopts agents today, or plans to, should stop starting the conversation with technology and start with the ruler. The framework and the model come after, and choosing them becomes obviously easier when there is an already-written success criterion that both must satisfy. An agent is only a product decision when there is a value metric beside it; without one, it is an experiment.

The change has a concrete, uncomfortable organizational consequence. Someone from the business side needs to sit with the technical side before the pilot's first token. Not after. And they need to come out of that meeting with the value ruler written, with an owner and a deadline. Most pilots skip that meeting because it is the boring part, the one that requires saying, in numbers, what "it worked" means. It is exactly the part the benchmark cannot do for anyone.

Looking at what is already published here, the boundary gets sharper. The model gateway at scale and the operational cost of an AI gateway in production deal with the infrastructure agents run on; the economics of model routing deals with cost per token. None of them is the quality-and-value ruler of a work agent. That ruler is another layer, and it is the territory this piece occupies.

The landing point is honest. The domain benchmark answers whether the agent knows how; the business ruler answers whether that is worth money in your operation. Both need to exist before the pilot, and the distance between them is the map of the integration work. A team that assembles both layers and measures the gap is piloting with its eyes open. The rest is betting, and calling the bet strategy.

FAQ

What is the difference between a domain benchmark and a business-value metric?

The domain benchmark measures whether the agent completed a real task against an agreed criterion, in a sandbox environment and with clean data. The business-value metric measures whether having that work done produced measurable results in the real operation, such as less cycle time or less rework.

Do I need to measure the benchmark-to-production gap in my pilot?

Yes. The gap between the capability proven by the benchmark and the value measured in the operation is the map of the integration work itself. Without it, the team cannot distinguish an integration problem from an agent problem.

Do agent benchmarks like LAB-AA serve to choose my agent?

They serve to calibrate technical capability before the agent touches the operation. What they do not do is say how much value that agent delivers in your specific flow, which is the question that decides the pilot.

When should I define the ruler, before or after running the agent?

Before. Defining the value ruler after running the agent turns every conclusion into hindsight, without a prior criterion to decide against. The ruler written with an owner and a deadline needs to exist before the pilot's first token.

Does Nexforce Agents help measure value, or just execute?

Nexforce Agents is where the ruler defined before the pilot is executed in a sandboxed, traceable way, via Nexforce Work and Nexforce Code. Measuring value remains a ruler the operation itself has to produce.

Referências e Leitura Complementar

The ruler is the pilot

The question that opens this piece, "what does a working agent mean?", has no technical answer. It has a business answer, and it needs to be written before any framework. The decision to measure capability and value in two layers, and to look at the distance between them, is the first decision of any serious pilot. Whoever makes it early wins a pilot that ends with a number to argue with or against. Whoever does not ends with an agent running and a question without an answer, and then the project dies slowly, without ever throwing an error.

Nexforce

Deploy Work and Code Agentswith zero software licensing costs

Automate operational tasks and code writing autonomously with dedicated agents integrated into your systems

Free Trial

Related articles