Teverant AI · Insights

2026-06-24

Enterprise AI agent project evaluation: 6 engineering signals for deciding whether it's worth building

Before launching an enterprise AI agent project, how do you make a rational call from an engineering standpoint? This article breaks the vague question of "can it be done?" into 6 quantifiable engineering signals: whether the task can be asserted, process stability, the determinism boundary (pass^k), the ROI ceiling, whether risk can be governed, and whether evaluation can be built after the fact. Each signal comes with a clear way to test it, helping technical teams complete a rigorous AI agent project evaluation at the approval stage and avoid continued investment in the wrong direction.

First ask "can it be done?": breaking a vague judgment into 6 engineering signals

In most agent project approval meetings, the final decision rests on two things: a demo that runs smoothly, and a vendor saying "we've done this use case before." Neither is false, but neither answers the question that really needs asking: if this task is handed to an agent, can it converge reliably from an engineering standpoint? Treating "is it worth doing?" purely as a business question often means writing the next three months of rework into the contract in advance.

The root of the problem is that evaluation itself has not yet converged. Benchmarks for measuring agents are everywhere today, from academic reasoning problems to industry-oriented end-to-end tasks, with more than a dozen established categories alone, each with its own definitions and scenario coverage. As a result, the same agent can score very differently on different benchmarks, and decision-makers have no single yardstick for comparing them side by side. Having many benchmarks is not the same as having a standard; precisely because there is no accepted frame of reference, selection is all the more easily swayed by a single attractive number.

More importantly, there is a very steep step between demo conditions and production conditions. In controlled examples with only a few steps, today's strong models really can deliver convincing performance. But once you move to complex workflows that require multi-turn planning, cross-tool calls, and accumulated state, even top-tier models see the share of runs that complete the entire chain on the first try drop sharply, far from the level at which you could confidently entrust it to end users. This is not a weakness of any particular model; it is the shared ceiling of the current generation of technology on long-horizon tasks. The message for decision-makers is direct: do not infer the stability of ten thousand executions from one successful demo.

So we break "can it be done?" apart into six engineering signals that can be verified one by one. They are not a scorecard but a series of gates up front. If any one of them won't close, however many resources you put in afterward, you are gambling:

  • Data availability: Can the correctness of the task be clearly "asserted," and are there enough real samples to define success and failure? This gate determines whether you will have anything to test later.
  • Process stability: Beyond the final success_rate, is every intermediate step observable and reproducible? A high overall success rate does not necessarily mean the process isn't relying on luck.
  • Determinism boundary: When the same input is run repeatedly, how widely do the results scatter? This directly determines whether it can face users directly, or can only sit behind a human safety net.
  • ROI boundary: Where is the capability ceiling, and will evaluation saturate too early? Return on investment is not about the peak value, but about whether the capability curve still has room to keep rising.
  • Risk governability: When errors happen, is there a guardrail that can stop them? Without guardrails that can actually be deployed, no matter how high the average performance, it should not go live.
  • Evaluation can be built later: Not having a ready-made evaluation set today does not mean the project should be rejected. It is the last box to tick before the Go/No-Go decision, not a hard condition that blocks approval.

The order of these six signals is deliberate. The first two answer "can the task itself be engineered?", the middle two answer "how good is good enough?", and the last two answer "can we contain things when problems occur?" An agent project worth doing is not one that scores full marks on all six, but one that knows clearly which signal it is stuck on, how big the gap is, and what it will cost to close it. Below we break them down one by one, starting with the one most often skipped yet most fatal: whether the success or failure of the task can actually be clearly asserted.

Signal 1: can the task be "asserted"? Data and sample availability

When evaluating agent projects, many teams start with the wrong question: "Is the current model strong enough?" That question comes later. What really holds you up at the approval stage is a more basic fact: can you write this task as a set of testable "assertions"? In other words, given a certain input, can you clearly state "the correct result should look like this," and find a few dozen real scenarios to fill in that statement? If you can't even do this, no matter how strong the model is, you have nothing to grab onto.

In engineering terms, an "assertion" is a sample: one input, one expected behavior, one evaluation criterion. The process of breaking a task into assertions is essentially interrogating yourself about whether the task has a clear right and wrong. If you can break it down, the business logic has converged; if you can't, it is usually not a technical problem but a sign that the requirements themselves are still in a gray zone. What this gate often stops are the projects that "sound very worth doing, but where no one can say what counts as success."

The sample size doesn't have to be daunting. A common misconception is that an evaluation set needs thousands of items to be meaningful, so teams hesitate to start. In fact, in its 2024 guidance on agent evaluation practices, Anthropic suggested that 30 to 50 high-value samples are enough to get started; other engineering guidance for agent development (2024) also noted that effective evaluation can begin with 20 to 50 real failure cases, without waiting to accumulate hundreds of tasks. The reasoning is simple: in the early stages, each change to an agent system usually produces behavioral differences large enough to see with the naked eye, and changes that coarse can be distinguished with a small sample. When later on you need to tell apart a subtle difference between 92% and 93%, that is the time to expand the sample. In other words, sample size should follow the precision of the question you need to answer, rather than chasing statistical polish from the outset.

What really determines sample quality is coverage, not quantity. A set containing only happy-path samples amounts to demonstrating that "it runs when everything goes smoothly," which is of almost no value for judging feasibility. Samples worth collecting must span several fundamentally different paths:

  • Normal path: The input is well-formed, the process flows smoothly, and the task is completed the way it should be. This is the baseline, but only the baseline.
  • Edge path: Incomplete information, missing parameters, ambiguous context. Does the agent ask follow-up questions and fall back sensibly, or does it push ahead and make up an answer?
  • Failure path: A tool call it depends on fails, an interface times out, or an unexpected structure is returned. This category directly tests its robustness in an uncertain environment.
  • Risk path: Scenarios involving irreversible operations, sensitive actions, or high-cost decisions. Should it stop and ask for approval, or will it charge ahead on its own initiative?

This four-way split is in line with some more granular breakdowns used in the industry. One evaluation method (2024) further divides tasks into five categories: normal completion, missing information, tool failure, high-risk actions, and distracting noise, and recommends building 50 to 100 tasks at the start. The granularity differs, but the core is the same: you need to deliberately construct the situations where things "don't go smoothly," because an agent's real value shows precisely in how it handles the unexpected, not in how it performs under ideal conditions.

Hidden here is a signal even more worth watching than "not being able to gather enough samples." When you try to construct samples for a particular path and find you simply can't write them (for example, you can't say "what the correct behavior is when information is missing," or can't define "which actions count as high-risk"), this usually isn't because samples are hard to find. It is because the business rules haven't been thought through. In that case, what's missing is not data but an understanding of the problem itself. Pushing ahead anyway only outsources vague requirements to a system that can't be tested, and in the end no one can say whether it is doing the right thing.

So the verdict of this section is clear-cut: not being able to gather samples is a stronger reason to stop than a model not being strong enough. Model capabilities will improve over time, interfaces will be upgraded, and costs will fall; these are all variables you can wait on. But if a task cannot be asserted and cannot be turned into testable samples, then in engineering terms it is a black box: you can neither judge whether it can be done now, nor know whether it is being done correctly once it is. Spending a day or two before approval to carefully gather 30 to 50 samples covering all four paths is itself the cheapest feasibility test there is. If you can gather them, the remaining signals are worth examining; if you can't, you may have just saved several months of rework.

Signal 2: process stability, the middle layer beyond success_rate

Many projects focus on a single number at approval: how many tasks were ultimately done correctly. If the number is high, the team feels confident; if it's low, they think it can't launch. The problem is that this number compresses the entire execution chain into a binary result, swallowing every incident that happens in between. An agent can put its steps in the wrong order during planning yet stumble onto the right final answer because later steps happen to correct it; it can pass parameters that don't match the schema when calling a tool and only get through by luck after three retries; it can truncate half of the retrieved evidence yet happen to land its final output in the correct range. In all these cases the final success rate looks fine, but once you deploy to production and the input distribution shifts, they are exposed immediately.

So the second signal for judging whether an agent task is worth doing is: can you see "how it was completed," not just "whether it was completed"? If your current evaluation can only produce an overall success rate, then the project's real stability is a black box to you, and any stability commitment made at approval is a guess.

To open the black box, observation must be split into four independent layers, each answering a different question, and all four must be examined together rather than picking the one that looks best:

  • Task layer: Is the final result correct? This is the familiar success rate; it is a conclusion, not a process.
  • Evidence layer: Does the conclusion hold up? Is the agent's answer backed by corresponding retrieval citations, and is the evidence preserved in full? A "correct answer" without supporting evidence is as good as no answer in compliance and traceability scenarios.
  • Execution layer: Were the tools driven correctly? The tool-call success rate and schema-validation failure rate directly reflect whether the interfaces between the agent and external systems are stable. If this layer collapses, success at the task layer is often piled up through retries and cannot be reproduced.
  • Risk layer: Were any lines crossed? The rate at which unauthorized operations are blocked, and the rate of budget or quota overruns. This layer is the bottom line: even if the first three layers look great, a leak here means the system cannot face real users.

The value of these four layers lies in how they corroborate one another. The task layer tells you the result; the other three tell you whether that result is trustworthy, reproducible, and controllable. Looking only at the task layer is like inspecting only a building's facade, with no idea whether corners were cut on the rebar inside the walls. I've seen many teams with an impressive task layer in the demo stage whose schema failure rate at the execution layer shot up as soon as they entered a phased rollout, because the tool return formats in the test set were too tidy and broke the moment the real environment changed. Problems like this will never be exposed by looking at success rate alone.

The four-layer framework solves the problem of "seeing faults," but there is a more hidden dimension: efficiency. Two agents can both get a task right, one in five steps and the other in twenty steps plus seven repeated calls. At the task layer they look identical, but in engineering terms they are worlds apart. More steps mean higher latency, higher token costs, and a larger surface for errors, and over the long run stability will inevitably be worse. So process metrics need to be observed separately: average number of steps, the ineffective-call rate (calls that were made but contributed nothing to the task), and the repeated-call rate (the same tool hit again and again with nearly identical parameters). These numbers help you distinguish "barely getting it right" from "getting it right cleanly." The former is a ticking time bomb once it goes live.

Beyond efficiency, the quality of the tool calls themselves also needs to be broken down, because "the call failed" and "the call succeeded but was wrong" are two completely different kinds of problems, and lumping them into one success rate makes the metric useless. Break it into four dimensions:

  • Right choice: Using the tool that should be used, with no mismatches such as using a query interface to perform a write operation.
  • Right parameters: The fields, types, and values passed in are valid, without relying on downstream fault tolerance to cover for them.
  • Right timing: Calling at the step where the call belongs, neither too early (before the information is complete) nor too late (still querying when it could already answer).
  • Right interpretation of results: Correctly parsing what the tool returns, without treating an error as data or an empty result as valid content and carrying on.

Of these four dimensions, timing and interpretation of results are the easiest to overlook and the most likely to cause incidents in production. Wrong parameters usually produce an immediate error and leave a trail; but with wrong timing or wrong interpretation, the agent often doesn't raise an error. It carries the wrong premise all the way forward until it surfaces in the final output, by which point it is very hard to pinpoint which step went wrong. So these two dimensions deserve their own statistics rather than being folded into a generic tool success rate. In addition, for high-risk tools (those that can modify data, spend money, or send external requests), violating calls must be counted separately, not mixed in with ordinary calls and expressed as a ratio. An ordinary query tool being called incorrectly now and then is harmless; a payment or deletion tool called incorrectly even once is an incident. Their tolerances differ by orders of magnitude, and they shouldn't be measured the same way.

Putting all of this together, the criterion for signal 2 is clear: if all you can get right now is a final success rate, then the project's process stability is unknowable, and at approval you should list "building four-layer observability + process metrics + per-dimension tool-call statistics" as prerequisite work, rather than adding it after launch. Only once you have this observability in place are you entitled to say the agent's process is "stable"; without it, so-called stability is just an optimistic judgment with no data behind it.

Signal 3: the determinism boundary, where pass^k decides whether it can face users

When validating agents, many teams look at just one number: run it once, did it succeed? If it did, they conclude "this can be done," and then schedule it, approve the project, and make external commitments. The problem is that vouching for a single result and vouching for "right every time" are two entirely different engineering commitments. The former needs luck to cooperate only once; the latter requires the system not to break down across consecutive calls. Conflating the two is a common starting point for an agent project's reputation to collapse after launch.

To tell them apart, the industry uses two metrics: one answers "succeeded at least once in k attempts," and the other answers "succeeded every single time across k consecutive attempts." The former measures the capability ceiling and is meaningful for exploratory, retryable tasks; the latter measures the delivery floor and determines whether an agent can face real users directly. The key to the decision is which one your business actually depends on. If end users expect a correct result in every interaction, only the latter counts, and no matter how good the former looks, relying on it is self-deception.

There is a counterintuitive piece of arithmetic here that every decision-maker should work through. Suppose the success rate for a single task is 75%. That sounds quite good: only one failure in four. But if a complete business process requires three consecutive steps to succeed, then, estimating them as independent events, the probability of getting through the whole chain is 0.75 cubed, leaving only just over 40%. In other words, a step you consider "basically reliable" on its own, once chained together, gives users less than even odds of getting a correct result. This collapse is not linear; it worsens exponentially with the number of steps. The longer the chain and the more nodes in series, the more severely small flaws at individual points are amplified.

Looking at this pattern in reverse is even more useful: when a user-facing agent has an absurdly high complaint rate after launch, yet spot checks of individual steps "look okay," it's most likely not that one stage is broken, but that from the very beginning you vouched for a multi-step chain using a single-attempt success rate. Exponential decay won't let you off just because every step is "about right"; it simply multiplies each step's "about right" together into what users experience as "often wrong."

Even more critical is the ceiling of the underlying capability itself. Industry evaluations generally show that in sufficiently complex real-world scenarios, even the strongest current models can have single-task success rates as low as around 30%. Plug 30% into the chained multiplication above: two steps in series leaves less than 10%, and three steps drops to nearly zero. This means that in scenarios that demand strong consistency, face users, and tolerate no errors, a solution with a base success rate of only about 30% is not something that "just needs a bit more optimization before it can launch"; it is structurally unqualified for delivery. No amount of prompt tuning can close a gap of that magnitude.

Does that mean anything with an insufficiently high success rate should be killed outright? No. How tight this boundary is depends on whether the business can tolerate a safety net. Here is how my decision framework divides it:

  • Scenarios where humans can provide a safety net: The agent's output first goes through a human review or correction step, so someone catches it when it's wrong. These scenarios can relax the "right every time" requirement; a single-attempt success rate high enough to significantly reduce manual work makes it worth doing, and the focus shifts to measuring how much manual effort is saved rather than pursuing zero errors.
  • Scenarios that allow asynchronous retries: The task doesn't require a real-time response, and failures can be automatically rerun or retried via a different path. Here, what really matters is the probability of "succeeding at least once within a few attempts," not single-attempt determinism; retrying is itself an engineering technique for raising the success rate.
  • Scenarios with neither a safety net nor retries: Users receive results in real time, results take effect immediately, and an error means an incident. Payments, compliance determinations, and automated external replies all fall into this category. This line must be marked in red in the project approval document: what it demands is a deterministic floor of consecutive success, not the capability ceiling of one successful run.

So the action for decision-makers in this section is very concrete: first confirm which category the agent you plan to build falls into, then decide which metric to use for acceptance. Break the task into the number of steps it actually has to go through, estimate a conservative success rate for each step, multiply them together, and see what's left at the end of the chain. If that number can't support the business's consistency requirements, and the scenario happens to allow neither a safety net nor retries, then however stunning the single-step demo, the conclusion should be to postpone or not do it. This is not pessimism; it is working out in advance the costs that will inevitably surface after launch. Doing this one extra multiplication at the approval stage saves an entire team's time in incident post-mortems.

Signal 4: the ROI boundary, capability ceilings and the evaluation saturation trap

What gets believed most easily in approval meetings is an upward curve. Someone will cite the pace of progress on public benchmarks: on coding evaluations like SWE-bench Verified, frontier models pushed their scores from around 40% to 80% within a single year. The slope really is steep, steep enough to convince the room that "it'll be deployable if we just wait another six months." But taking that curve directly as the ROI expectation for your own project is the most common misjudgment I've seen.

The problem is that scores and capabilities do not correspond linearly. As a benchmark approaches saturation, the tasks still left unsolved on the leaderboard are precisely the hardest, most counterintuitive ones that depend most on long-chain reasoning. Even a genuine leap in model capability shows up in the score as a shift of only a few percentage points. In other words, the closer you get to the ceiling, the more each point is worth and the more steeply the cost of gaining it rises, while the "rate of progress" you extrapolate from the slope of the curve systematically overestimates future gains. What decision-makers see is a smooth extrapolation; what it corresponds to in engineering is the stretch where marginal returns diminish sharply. Treat the optimistic slope from the demo stage as your growth assumption after deployment, and the ROI model is off from the very first step.

A more hidden distortion happens when you focus on a single metric. I have a real example from an enterprise knowledge base agent that shows how a single-metric view can deceive. The team revised the prompt, and on an offline run of 20 samples, the task success rate rose from 71% to 83%. Looking at that number alone, a twelve-point improvement, writing "significantly optimized" in the approval materials would be no stretch at all.

But lay out the other dimensions for the same batch of samples and the picture changes completely:

  • Citation accuracy (citation_rate) collapsed from 88% to 61%: the price of the higher success rate was that the model began fabricating or misattributing citation sources, which is an almost fatal flaw in knowledge base scenarios;
  • The average number of tool calls rose from 2.1 to 3.8: more of the successes were piled up by "trying a few more times," and every call is real money spent on tokens and API costs;
  • p95 latency doubled outright: the waiting time experienced by the slowest-served users doubled, and in interactive scenarios that is often the tipping point for churn.

This is the key point: in isolation, the success rate went up, but taken as a whole, the net benefit may be negative. The extra two tool calls, multiplied by daily active users and per-call cost, are an ongoing expense; doubled latency means either scaling out or sacrificing the experience; and the collapse in citation accuracy puts the manual review workload right back. The work you thought the agent was saving people gets sent back to manual verification because the citations can't be trusted. Stack these three together and they could easily eat up the entire twelve-point success rate dividend, or even leave you worse off.

So ROI judgments cannot be anchored on the optimistic reading of any single metric; they must be anchored on two things. The first is the capability ceiling: is this type of task in a steep-growth phase on public benchmarks, or already close to saturation? If your core use case happens to correspond to the hardest tasks on the leaderboard, don't expect free gains from model iteration in the short term; you can't skimp on any of the engineering safeguards you need to invest in. The second is the total cost picture: convert latency, number of calls, token spend, and the most easily overlooked cost of all, manual review, into the same table, then net them against the improvement in success rate. A solution that improves on one metric while every other dimension deteriorates is not, in engineering terms, an optimization; it is a cost transfer. You have simply moved the price from one visible place into several invisible ones.

The actionable guidance for decision-makers: require that any "success rate improvement" claim come with three supporting data points on the same batch of samples: citation accuracy, number of calls, and p95 latency. If any one of them is missing, the ROI conclusion does not hold and cannot be used as grounds for a Go. It is easy to pick a flattering number at the demo stage; only a solution that stays positive when the full cost table is laid out is worth pursuing.

Signal 5: risk governability. No guardrails, no launch

The first four signals answer "can it be built?"; this section answers a different question: once it's built, can you stop it when it goes out of control? Once an agent can call tools and write data, each of its decisions is no longer just text generation but an action with side effects. The most realistic criterion for whether a project is worth launching is not how well it performs on average, but whether you can absorb the cost when it makes mistakes. If the answer is "no," the project's priority should be pushed back, and the governance layer should be completed before launch is even discussed.

I tend to look at governance in three layers: can it be blocked, can it be recovered, and can it be seen. If any one of the three is missing, the risk is uncontrollable.

Layer 1: can it be blocked before the action happens?

In practice, a minimum viable production safeguard comes down to a few hard checkpoints, each corresponding to a class of pitfall that has already been hit. If the structured output from the model fails schema validation, it should not enter downstream execution; blocking it outright is far safer than "best-effort parsing," since JSON with half its fields missing is often harder for downstream systems to debug than an outright error. Every write operation must first pass a policy check, taking "whether it can write, where it writes, and who authorized it" out of the model's free improvisation and handing it to deterministic rules. Once a single task's tokens or tool calls approach the budget limit, degrade immediately rather than pushing through; otherwise a run stuck in a loop can blow up both cost and latency at the same time. And there is one easily overlooked category: when the model still reaches a conclusion without retrieving any evidence, that answer must be flagged as high-risk rather than returned to the user alongside normal answers.

What these checkpoints have in common is that none of them depends on the model "thinking it through" on its own; instead, they take decision authority back to the engineering side. That is precisely the point of guardrails: the model may make mistakes, but you draw the boundaries within which it makes them.

Layer 2: can it recover after an error?

Blocking is only the first step. What better distinguishes a project's maturity is its ability to recover from errors. Here I don't recommend using a blanket "recovery success rate," because the correct recovery behavior differs completely by error type, and lumping them together actually hides the problems.

  • Missing parameters: The correct response is to stop and ask the user, not to guess with a default value. Whether the agent can recognize "insufficient information" and proactively ask is a hard capability dividing line.
  • Tool timeout: What's needed is a limited number of retries with backoff, not mindless retrying until the other service goes down, nor giving up after a single failure.
  • Insufficient permissions: The only correct response is to stop and report, and the agent must never learn to "work around permissions." An agent that looks for ways around permissions is far more dangerous than one that throws an error.

Only by tracking recovery behavior separately for each error type can you see clearly whether the model really handles exceptions, or merely performs well when things go smoothly.

Layer 3: can it be seen in production?

No matter how well the first two layers are done, if you can't see what's happening, you have no governance. Keep traces for all key runs, not a sample but every single one, because the run that goes wrong is often exactly the one you didn't sample. At the metrics level, task completion rate alone is far from enough: it only tells you "how many succeeded," not "how they succeeded and what nearly went wrong."

Production monitoring that can truly support judgment needs roughly a dozen dimensions working together. Tool-call count, task completion rate, average execution steps, user follow-up rate, tool failure rate, timeout rate, and repeated-call rate characterize "whether it runs smoothly"; high-risk action confirmation rate, user cancellation rate, human takeover rate, negative user rating rate, and number of safety interceptions are what really characterize "how big the risk is." In the second group, the human takeover rate and the number of safety interceptions deserve particular attention: the former shows in what proportion of cases the system still needs a human safety net, and the latter shows whether the protective layer is actually doing its job. If these metrics simply can't be collected in your existing observability stack, then risk governability should be judged as not met.

Looking at the three layers together, the conclusion is direct: only when it can block, recover, and be seen does a project meet the conditions for launch. If any layer is missing, the agent should not yet face real users or real data. Guardrails are not an optimization to add after launch; they are a precondition that determines whether it can launch at all.

Signal 6: evaluation can be built later. It is not a blocker for approval, but the final box to tick before Go/No-Go

The first five signals are about "whether the thing itself can be done." With the sixth signal, the focus shifts to "whether we are capable of knowing if we've done it right." Many teams get stuck here: they have no ready-made evaluation system at all, so their judgment stalls. Either they reject a project that should have gone ahead because "it can't be measured," or they launch an unreviewable black box on the principle of "let's get it running first." Both are misjudgments. The value of an evaluation system is beyond doubt, but when to build it and whether to approve the project are two separate matters.

A common path in real engineering is to have the product first and, once there is real traffic, go back and fill in evaluation. Agent teams that had already built a sizable user base set up evaluation systematically only after the product was running, combining three layers: static analysis scoring, real-task testing in the browser, and a model acting as judge. The whole process took months, not years. Another approach is to let evaluation grow alongside the product: in the early stage, humans score items one by one; once enough judgment samples have accumulated, human judgment is distilled into an LLM grader, with scoring logic built around dimensions such as "does not break existing content, actually completes the instruction, and produces output of acceptable quality," while keeping a step for periodic human calibration. It eventually evolves into two independent suites: one watching the quality baseline and one guarding against regressions. Both paths show the same thing: evaluation can be added later, and the cost of adding it is controllable and predictable. So at the approval stage, it should not be a hard, single-vote veto, but the final item that must be ticked off before launch.

Conversely, what deserves more caution is not "having no evaluation" but "having a flawed evaluation." A poorly designed evaluation will systematically underestimate your agent and lead you to kill, with your own hands, a project that is actually running well. On one benchmark for code and scientific research, a highly capable model initially scored only just over 40%, which looked absurdly bad. When researchers later went through it item by item, they found the problem was on the evaluation side: the grader treated numerical approximations as errors (an answer that should have been judged correct was marked wrong because of trailing decimal digits), some task descriptions were ambiguous, and some randomized tasks could not be reproduced reliably at all. After these holes were filled and a more reasonable judging standard was adopted, the same model's score came close to perfect. The gap was not in the agent but in the ruler.

Another example is subtler. In a conversational benchmark for flight booking, an agent dug into the rules and found the user a better-value option than the reference answer, only to be judged a failure because that option was not on the "correct path" hard-coded in advance in the evaluation script. It did better than the evaluation expected and lost points for it. What makes this kind of error dangerous is that it penalizes precisely the agents that are genuinely creative and capable. If you look only at the score, you'll conclude that "this path doesn't work," and then turn around and invest in a solution that scores nicely but is actually mediocre.

So the real meaning of the sixth signal is: evaluation must be built, and more importantly, it must be built right. Bringing the six signals together: you don't need the evaluation system in place at approval, but before launch you must satisfy a release checklist. At least 30 offline samples with explicit assertions, so that "right or wrong" has an objective basis. Metrics broken out across four levels (data, single step, process, and business), so that a blanket success rate doesn't mask the layer where the real fault lies. Support for replaying key failure samples exactly as they occurred; otherwise, even when you fix a bug, you have no way to verify it. All high-risk write operations wired into production guardrails, keeping irreversible actions behind the gate. Before-and-after comparison results kept for every version change, so that each change can clearly be shown to have made things better or worse. Passing all six signals does not guarantee project success, but if any one of them stays red for long, it's worth stopping to redo the math before you invest.

Do these six signals have a priority order? Which ones should be an outright veto if they fail?

There is an order, but it's not a simple ranking. The first three signals (whether data and samples can be obtained, whether the process is stable, and whether the determinism boundary is adequate) are "foundation-level." If any of them stays unmet for long, the project basically doesn't hold up and should be rejected outright or postponed. The last three are "governance-level": the ROI boundary helps you judge whether it's worth doing, risk governability determines whether it can face real users, and building evaluation later is the final box to tick before launch. Failing at the foundation level is a hard veto; failing at the governance level usually means "fix it before launch," not "give up."

If success rates are only 30%, does that mean enterprise agents aren't worth building yet?

No. Reaching a veto from a single blanket success rate is exactly the mistake this framework is meant to avoid. A 30% end-to-end success rate may mean that one specific stage in the process is dragging things down. After you break the metrics into four layers, you'll often find the data layer and single-step layer are fine, and the problems are concentrated in process handoffs or a certain class of edge inputs, which are fixable. What's more, as the two examples of evaluation flaws above show, a low score may itself be a problem with the ruler. Pinpointing which layer the failures occur in before deciding whether it's worth doing is far more reliable than making the call based on a single overall score.

If we don't have an evaluation system at the approval stage, does that mean we can't make a judgment?

Quite the opposite. An evaluation system is an engineering asset that can be built once the product matures, taking shape within months, and should not be a precondition for approval. What you really need to confirm at the approval stage is "whether the task can be asserted," that is, whether the results have an objective standard of right and wrong and whether samples can be obtained. As long as that holds, evaluation can be added sooner or later; if even that doesn't hold, that is the real signal to stop, and building evaluation later won't save it.

How do we avoid being misled by a polished demo or benchmark scores?

With a demo, what matters is whether it can be replayed consistently: how many times out of ten runs of the same scenario it succeeds, not one edited success. With benchmark scores, first read the scoring logic itself: how does it judge right and wrong, does it count numerical approximations as errors, are task descriptions ambiguous, and could a better non-standard solution be wrongly penalized? Highly capable agents really do get buried under low scores. Look at "the score" and "how the score was produced" separately, then add before-and-after comparisons and failure replays, and you won't be led around by a single number.