Teverant AI · Insights

2026-08-24

Deploying AI agents in the enterprise: a 5-step implementation plan

From assessing the need and screening use cases to breaking down workflows, designing permissions, integrating systems, and evaluating results, this article lays out a 5-step method for deploying enterprise AI agents, helping enterprises go from a minimal closed loop to a controlled launch within 4–6 weeks.

Decide first: does the enterprise need an AI agent, or an ordinary automation tool?

When an enterprise introduces AI capabilities, the first task is not choosing a model but determining the type of task. Knowledge retrieval, workflow automation, and autonomous execution solve different problems. If every requirement is lumped under "agent," the project takes on extra integration, permission, and audit costs, and becomes harder to pass acceptance reliably.

Requirement characteristicsBetter-suited solutionBasis for the judgment
The main task is looking up policies, product materials, or internal documentsRAG knowledge assistantThe system only needs to retrieve relevant content and generate an answer; it does not need to modify business data or drive downstream processes
Clear operating steps, fixed decision conditions, and a relatively stable target interfaceRule engine or RPAThe task path can be orchestrated in advance, input and output boundaries are clear, and there is no need to change the plan on the fly based on semantics
Requires contextual judgment, pulls data from multiple systems, and chooses next actions based on intermediate resultsAI agentThe execution path cannot be fully predefined; the system must call tools, adjust steps, and handle exceptions within defined constraints

So "can answer questions" should not be the criterion for approving an agent project. An agent aimed at business execution must at least form a closed loop that can be inspected: receive the task, gather the required information, make a judgment, call systems to perform operations, return the result, and hand off to a human when it exceeds its permissions, lacks information, or encounters a tool failure. If the system ultimately just offers a paragraph of advice, and employees then log into multiple systems themselves to complete the operations, it is still an assistant, not a complete execution agent.

To judge whether you need an agent, first check whether the workflow contains genuine dynamic decision-making. For example, a customer request may need to be handled based on the customer's history, the current contract, and inventory status, and different outcomes may trigger different approvals, notifications, or data writes. A workflow like this may suit an agent. Conversely, if every input maps to exactly one action, deterministic automation is usually easier to test and makes risk easier to control.

The basis for approval should also shift from model performance to business outcomes. The team first needs to be clear about what it wants to change: shorter handling time per task, less rework and fewer errors, lower cost per task, higher effective conversion, or broader service coverage outside business hours. Once the goal is set, define the current baseline, the measurement definitions, the observation period, and how the manual comparison will be done. Model scores can help diagnose problems but cannot replace business acceptance.

  • If the benefit can only be described as "more natural answers" or "a smarter-looking demo," the project is not yet ready for business approval.
  • If the task's value is clear but a wrong operation could cause significant losses, first restrict the actions it can execute and keep a human confirmation step.
  • If the investment in system integration, data authorization, and exception handling clearly exceeds the expected benefits, prioritize an assistant or partial automation.
  • Only when the task occurs frequently, the cost of manual handling can be identified, and execution results can be tracked are the conditions in place for a further ROI calculation.

Market growth, changing deployment timelines, and flagship case studies only show that the technology is entering enterprise workflows; they cannot replace your own enterprise's judgment. The final decision should still come down to three questions: does the use case genuinely require dynamic judgment, can the risks be constrained through permissions and human intervention, and can the expected benefits cover the cost of building and continuously running it? When all three are not true at the same time, not building an agent is often the safer engineering choice.

Step 1: screen use cases, using five criteria to find the workflows most worth deploying

The goal of screening is not to find the "most advanced" business area, but to determine whether the first pilot can form a closed loop that is verifiable and can be rolled back. Customer service, sales follow-up, IT operations, and invoice processing can all go into the candidate pool, but should not be approved simply because they are common. The enterprise still needs to check whether it has a stable task volume, relatively consistent ways of working, and data and system capabilities the agent can call.

We recommend building a use-case scoring sheet. It is not a universal industry standard, but a way to align judgment across business, technology, and risk control teams. Evaluate each candidate workflow on the following five dimensions:

Evaluation dimensionQuestions to confirmCharacteristics suited to a first pilot
Task frequencyDoes the task occur continuously? Can manual effort be recorded?Stable business volume, making before-and-after comparison easy
Rule clarityCan the inputs, decision conditions, and output requirements be stated? Is there a path for handling exceptions?Main steps are clear, and the few exceptions can be handed off to a human
Data availabilityDo the required documents, records, and business fields exist? Is calling them permitted?Data sources are well defined, and quality issues can be pinpointed
System operabilityDoes the agent need to read or write? Do the target systems offer stable interfaces?Limited call scope, with operation results that can be tracked and undone
Error toleranceWhat impact would an error have? Can it be reviewed, intercepted, or remedied?Errors do not directly trigger irreversible, high-risk consequences

Don't simply add up the scores. Task frequency and rule clarity determine whether the value is easy to verify, data and system conditions determine whether the solution can be implemented, and error tolerance determines whether it is appropriate for the agent to take actions. A workflow with very high business volume is still not suitable as a first project if the data it needs has not been organized, or if a single mistake could have serious consequences.

The first use case should also be deliberately narrow in scope. A safer approach is to serve only one clearly defined user group, handle one type of task with relatively fixed inputs and outputs, and connect only a few necessary systems. For example, rather than overhauling the entire customer service chain from the start, you could limit it to retrieving materials, drafting replies, and human confirmation for one type of inquiry. This lets you observe the agent's decision process and also makes it easier to tell which problems come from the model, the knowledge data, or the system interfaces.

Setting exclusion criteria matters more than raising candidates' scores. The following situations should be kept out of the first pilot for now:

  • The task is occasional, making it hard to build a stable sample and impossible to measure return on investment continuously;
  • Ownership of the workflow is unclear, and when exceptions occur no one decides how to handle them;
  • Key data is missing, its quality cannot be controlled, or access rights cannot be obtained lawfully;
  • The business rules rely mainly on individual experience and have not yet been turned into describable processing boundaries;
  • The agent's actions could directly trigger irreversible payments, authorizations, or critical configuration changes.

Once screening is complete, the output should not be just a list of use-case names. For each selected item, spell out at least the target users, the starting point of the task, the expected output, the data required, the systems involved, the human review points, and the actions it is prohibited from executing. Only when these boundaries can be confirmed by the business owner is the use case ready for the next step, workflow breakdown.

Step 2: break down the workflow, turning "let the agent handle it" into an executable task chain

Once the use case is chosen, don't jump straight into writing prompts or configuring tools. First map the existing business process into a task chain that can be inspected, interrupted, and rolled back. "Let the agent handle refund requests" is still far too broad: it doesn't say where requests come in, which records need to be read, under what conditions it can proceed, who is authorized to submit the refund, or where things go after a failure. Even if a requirement like this can produce a demo, it will have a hard time making it into production.

We recommend drawing the flowchart with the following structure and assigning an owner to every node:

Workflow elementQuestions to clarifyTypical outputs
Trigger conditionsIs it started by a user request, a system event, a scheduled task, or a manual assignment?Event definitions, trigger scope
Input dataWhich fields, documents, and historical records are needed? Is the data complete and readable?Input list, rules for missing items
Decision rulesWhich conditions have fixed definitions, and which require contextual judgment?Rule tables, decision branches
Tool callsWhich business systems need to be queried or written to? How are failed calls retried?List of interfaces and permissions
Business actionsDoes the agent only make suggestions, or can it create tickets, change statuses, or submit approvals?Action definitions, approval requirements
Output resultsWho receives the result, in what format, and does the supporting evidence need to be saved?Structured results, operation records
Exception handlingWho does it hand off to when it encounters conflicting data, tool timeouts, or uncertain judgments?Paths for termination, retry, and human takeover

The most important judgment during breakdown is separating deterministic processing from open-ended judgment. Tasks such as field format checks, amount reconciliation, status matching, and threshold checks should go first to programs, rule engines, or RPA. Their inputs and outputs are well defined, and they are easier to test. Models are suited to work that requires understanding context, such as intent recognition, summarizing long texts, synthesizing materials, and generating candidate solutions. Don't have a model calculate results that code can compute precisely, and don't force large numbers of fixed rules to cover semantic judgments.

Every node should be labeled with who executes it, using four tags: agent, rules or program, RPA, and human. The choice is not about which technology is newer, but whether the task requires semantic reasoning, whether a stable interface already exists, whether the action can be reversed, and whether the consequences of an error are acceptable. For example, the agent can read the refund explanation and organize the reasons, a rules program verifies the amount and order status, RPA can enter information into a legacy system that lacks an interface, and the final refund submission is confirmed by an authorized person.

The flowchart should also highlight key decision points. Actions such as approving refunds, issuing external quotes, changing contract terms, and handling production incidents typically have financial, legal-commitment, or business-continuity implications. For these, the agent should not be allowed to run all the way from understanding the request to execution. A safer approach is to have it prepare materials, provide the basis for its judgment and a recommended action, and pause before anything is actually written to a system or takes external effect, so the responsible person can confirm. Human confirmation can't be a mere formality button: the interface should present the original input, the cited evidence, the proposed action, and the scope of impact together, so the reviewer can make a judgment rather than re-investigating the entire process.

Exception paths should be designed at the same time as the normal path. At a minimum, handle missing input, documents that can't be parsed, conflicts among multiple data sources, tool call failures, insufficient model confidence, and human response timeouts. If PDFs or scanned materials can't be read directly, run them through text recognition and structuring first, then pass them to retrieval or judgment. Any task that cannot meet its preconditions should stop at a clearly defined node, preserve its context, and hand off to a human, rather than having the agent guess the missing information and carry on.

The first pilot should first verify whether a single agent can complete a minimal closed loop: receive a task, obtain the necessary information, form a judgment, call a limited set of tools, output a result, and wait for confirmation at risk points. Only once responsibility boundaries are clear, and the logging, permissions, and exception mechanisms of the single-agent flow run reliably, is it worth splitting into multiple agents for finance, compliance, business, and so on. Multiple agents can isolate specialized responsibilities, but they also introduce problems with task handoffs, state synchronization, permission propagation, and fault localization. If a single agent hasn't closed the loop yet, adding collaborating roles usually just spreads what were clear errors into the orchestration layer.

The acceptance deliverable at this stage is not a conceptual flowchart, but a set of executable definitions: the inputs and outputs of each step, who executes it, the tools it can call, permission boundaries, human confirmation points, how failures are handled, and the audit trail requirements throughout. Only once these are clear do subsequent system integration and results evaluation have a verifiable foundation.

Step 3: design permissions, defining what the agent can see and what it can do

Agent permissions cannot simply reuse the way employee accounts are configured. Employees decide whether to proceed based on policies, experience, and the boundaries of their responsibilities, whereas an agent may treat a one-time authorization as a capability that remains available indefinitely. Permission design therefore cannot only answer "is it allowed to access the system?"; it must separately constrain the data it can read, the tools it can call, and the actions it can submit or execute.

Permission layerQuestions to clarifyTypical controls
Data permissionsWhich fields, customers, and time ranges can it view?Field masking, tenant isolation, restricted query conditions, no bulk downloads
Tool permissionsWhich interfaces can it call, and what are the call frequency and parameter ranges?Interface allowlists, parameter validation, quota limits, timeouts and circuit breakers
Action permissionsCan it make suggestions, create drafts, or directly complete business operations?Tiered actions, approval thresholds, two-person review, confirmation before execution

These three layers cannot substitute for one another. Allowing an agent to read customer profiles does not mean allowing it to export them in bulk; allowing it to query orders does not mean allowing it to change order status; allowing it to calculate a refund plan does not mean it has the authority to initiate a payment. In engineering terms, "read, write, export, submit, approve, execute" should be split into separate capabilities, rather than covering the whole workflow with one broad role.

Every agent should use its own machine identity, never reusing an employee account, let alone borrowing an administrator account. The identity configuration should at least be bound to a specific purpose, access scope, call quota, and expiration time. When a pilot ends, a task is withdrawn, or the identity goes unused for a long time, it should be deactivated promptly and its credentials revoked. Keys should not be written directly into prompts, scripts, or configuration files; they should be issued on demand and rotated by a unified credential management mechanism.

The scope of automatic execution should be determined by business risk, not by model confidence alone. Operations that are reversible, limited in impact, and governed by clear rules can be completed automatically once parameter validation passes; those involving large sums of money, sensitive personal information, external contractual commitments, or irreversible changes should stop at an approval node. Model scores can only serve as an input to approval; they cannot replace authorization rules.

  • Low-risk tasks: information classification, record completion, internal draft generation, and similar tasks can be processed automatically, but results must still be recorded.
  • Medium-risk tasks: ticket routing, drafting customer replies, modifying business fields, and similar tasks can be submitted by the agent and take effect after a designated person confirms.
  • High-risk tasks: financial operations, account deactivation, price commitments, sending sensitive data externally, and similar tasks may only produce recommendations and evidence, never direct execution.

Human-in-the-loop should not be just a "confirm" button in the interface. The approval page needs to show the input information, cited evidence, proposed action, key parameters, and affected objects together, so the approver can judge why the agent is doing what it does. If the approver sees only a one-line conclusion, human review easily degenerates into a rubber stamp.

Permission controls also need to be built alongside the audit chain. For every task, you should be able to reconstruct the user's request, the materials retrieved, the judgment formed by the model, the interfaces actually called, the parameters used, the approver, and the final execution result. Sensitive fields can be masked in audit records, but the links between events must not be lost as a result. The goal of auditing is not to store every conversation, but to be able to answer "who, under what identity, based on what information, performed what action on which object."

Finally, define exception-handling paths in advance: reject immediately when an unauthorized request is detected; pause the task after consecutive call failures; block subsequent steps when an execution result deviates from expectations; invoke compensation or rollback first when side effects have already occurred; and hand off to a human when automatic recovery isn't possible. Only when pausing, reversing, and taking over all actually work has the agent been given controlled permissions, rather than being plugged into a super account with no boundaries.

Step 4: integrate systems, bringing the agent into real business operations

Whether an agent can go into production depends not on whether the model can answer questions, but on whether it can obtain trustworthy context and complete business actions through controlled interfaces. Integration work should start from one concrete workflow: which materials need to be read, which systems queried, what results written back, and how to recover after a failure. Don't build a company-wide "big platform" first and then go looking for use cases.

Prepare usable knowledge first, rather than handing files straight to the model. Policy documents, product materials, historical cases, and similar content can be provided to the agent through RAG. Before ingestion, complete deduplication, chunking, metadata enrichment, and permission tagging, and retain the file version, publication date, scope of applicability, and location in the original text. Retrieval results must be traceable back to their source passages; otherwise business staff will struggle to verify them, and outdated policies may continue to influence decisions.

PDFs, scanned contracts, invoices, and image-based tables cannot simply be stored as full text. Use OCR first to extract text, tables, and key fields, then perform layout correction and quality checks. For fields that affect downstream actions, such as amounts, dates, and customer names, set a confidence threshold; when recognition results are uncertain, route them to a human for confirmation rather than letting the agent guess from context. Structured content must also inherit the access permissions of the original file, so the knowledge base doesn't become a new way around business authorization.

Prefer stable APIs for system integration. APIs have clear input and output boundaries, making permission control, error handling, and call auditing easier. For legacy systems that have no usable interface but whose page structure has been stable for a long time, RPA can perform interface operations such as lookups, downloads, and data entry. In that case the agent decides the next task, and RPA only executes actions that have already been constrained. If pages are redesigned frequently, CAPTCHAs are common, or operation results are hard to confirm, the maintenance cost of RPA may exceed that of building an interface, and this should be verified before the pilot.

CRM, ERP, order, ticketing, and messaging systems should not each be exposed to the agent separately. A safer approach is to wrap business capabilities into purpose-specific tools, such as "look up customer status," "create after-sales ticket," and "read order details," rather than opening up generic database queries or arbitrary script execution. The unified entry point should at least handle the following:

  • Identity propagation: use the real permissions of the current user or service account, without expanding the scope of authorization because an agent is involved.
  • Parameter constraints: validate field types, value ranges, and required fields, and reject anomalous parameters generated by the model.
  • Call protection: set concurrency limits, timeouts, idempotency keys, and limited retries to prevent duplicate orders or duplicate notifications.
  • Audit records: store the requester, tool called, input parameters, execution result, and failure reason to support issue tracing.
  • High-risk interception: operations such as payments, deletions, bulk modifications, and external sends go to human approval before execution.

Decide on the deployment model last, and don't default to private deployment. Selection should consider data sensitivity, response latency, call frequency, and long-term maintenance capacity together. Public model APIs are suitable for validating the workflow first, but you need to confirm data retention, data transfer, and the vendor's compliance boundaries; private deployment strengthens control over the environment but brings work such as compute provisioning, model upgrades, inference optimization, monitoring and on-call coverage, and ongoing adaptation.

Decision itemQuestions to confirm
Data boundariesAre original documents, retrieved passages, and call logs allowed to leave the environment the enterprise controls?
Response requirementsHow long a wait can the business process tolerate, and are peak call volumes stable?
Cost structureCompare actual API spend with the total cost of compute, operations, and model adaptation
Operational capabilityAre version management, failover, security patching, and performance regression testing in place?

The acceptance criterion for this step is not "the systems are connected," but that the agent can obtain the correct materials within its permissions, call the designated tools, recognize execution failures, and stop at high-risk nodes to wait for confirmation. Only when this controlled chain runs end to end has the agent truly entered real business operations.

Step 5: evaluate results, using business metrics rather than demo impressions for acceptance

An agent that can get through a demo is not necessarily one that can work reliably in a real workflow. Acceptance should answer three questions: have business results improved, is the new risk controllable, and is it still worth continuing once all costs are counted? To avoid having nothing to compare against after launch, the project team should fix the measurement scope, sample conditions, and calculation methods before the pilot, and record a comparable baseline for the manual process.

Freeze the baseline first, then discuss improvements

The baseline should cover at least average handling time, first-contact resolution rate, share of human takeovers, error rate, cost per task, and final business outcomes. The last item needs to be defined per use case: for example, effective lead conversion in sales workflows, issue resolution and subsequent complaints in customer service workflows, and order completion or exception recovery in operations workflows.

Definitions must be written into the acceptance document. For example, does "handling time" start when the request enters the queue, or when the agent first calls the model? Does "first-contact resolution" allow follow-up questions? Does a human merely confirming count as an intervention? If different definitions are used before and after the pilot, the results cannot be compared.

DimensionRecommended metricsConfounding factors to exclude at acceptance
Business resultsRevenue change, cost savings, conversion outcomesPromotions, seasonal fluctuations, changes in customer mix
Workflow operationEnd-to-end time, throughput per unit of time, human takeover rateDifferences in task difficulty and intake channel
Models and toolsCorrectness of results, tool call success, retries and timeoutsSystem failures, API rate limiting, missing data
Business riskUnauthorized access, erroneous execution, customer complaints, and human reversalsDuplicate events and unconfirmed alerts

These metrics cannot substitute for one another. Faster responses with more complaints should not be judged a success; correct model answers combined with frequent failures when writing to business systems also cannot go into full production. The project team should set hard stop conditions in advance, especially for unauthorized access, irreversible operations, and errors that affect customers' rights and interests.

Validate in four stages instead of granting execution rights outright

  • Offline testing: Use anonymized historical tasks and specially constructed edge-case samples to check task understanding, tool selection, parameter generation, and refusal strategy. The test set should include normal workflows, insufficient information, system exceptions, and malicious input.
  • Shadow mode: Receive real requests and generate recommendations, but do not call any interfaces that change business state. Compare its output with the results of manual handling and downstream business outcomes, focusing on stability across different task types.
  • Low-traffic phased rollout: Open it only to limited users, limited workflows, or low-risk tasks. High-impact actions continue to require human confirmation, and the reasons for every takeover, modification, and reversal are recorded.
  • Full launch: Expand the scope gradually once the agreed thresholds are met, keeping monitoring, auditing, degradation, and shutdown mechanisms in place. After any change to the model, prompts, knowledge base, or interfaces, revalidate the affected task chains.

Shadow mode cannot rely on comparing text similarity alone. What really needs checking is: if this recommendation were adopted, would the order status, customer issue, or operational event have been handled correctly? For tasks without a clear reference answer, business staff can do blind reviews and record the verdict ("usable as is," "usable after edits," "not usable") along with the reasons.

Standardize how ROI is calculated, and don't count only model costs

Annual net benefit can be defined as annual attributable benefits minus all annual costs, with total investment as the denominator for calculating ROI. Benefits include verifiable revenue gains, freed-up labor hours, and reduced error losses; investment should include model calls, business system modifications, data preparation, human review, monitoring and auditing, and ongoing operations. Freed-up hours should count as a benefit only when they actually reduce outsourcing spend, avoid new hiring, or are redirected to measurable work.

External case studies can help round out your list of metrics, but their efficiency gains or benefit ratios cannot be applied directly. Enterprises differ in task complexity, labor costs, historical data quality, and depth of system integration. Whether to expand the rollout should ultimately rest on a baseline with consistent definitions, real business samples, records of risk events, and a complete accounting of costs.

Completing the first pilot in 4–6 weeks: from a minimal closed loop to a controlled launch

4–6 weeks is a suitable planning window for a first pilot; it does not mean every project can go into production within that period. The goal is not to deliver a complete platform, but to answer one question: under clearly defined boundaries and controlled permissions, can an agent reliably complete a real business workflow and generate benefits sufficient to cover its costs and risks? If system interfaces, data quality, or compliance approvals are not yet ready, narrow the scope rather than compressing the testing phase.

PhaseMain workDeliverablesExit criteria
Week 1Evaluate candidate use cases, map the existing workflow, collect the manual handling baseline, and log data, permission, and execution risksPilot boundaries, flowchart, baseline records, risk list, stop conditionsInputs and outputs can be defined, owners are clear, and results can be accepted
Weeks 2–3Prepare the necessary knowledge, connect the systems the pilot needs, and validate the task chain in a sandboxMinimal knowledge set, interface list, test cases, exception-handling pathsAt least one business loop runs end to end
Weeks 4–5Run in shadow mode first, then open to limited traffic; authorize only reversible, low-impact actionsTakeover records, failure logs, error attribution, permission audit recordsExceptions can be detected, intercepted, traced, and taken over by a human
Week 6Review quality, business value, risk events, and total costReview decision to expand, adjust, or stopThe acceptance thresholds agreed before the pilot are met

Weeks 2–3 are where scope creep is most likely. In engineering terms, prioritize getting the minimal chain working: "receive the task, read the necessary information, form a judgment, call tools, return the result, human confirmation." Reports, complex orchestration, and non-essential interfaces can come later. Adding features before the loop is stable only widens the scope of fault localization.

Once in shadow mode, the agent can generate recommendations and proposed actions but does not directly affect production data. The team should compare against manual results and record whether each deviation stems from missing knowledge, model judgment, tool calls, or the workflow definition. During the phased rollout, open up action permissions gradually, while keeping mechanisms for human takeover, audit trails, fast reversal, and emergency shutdown.

The conclusion in week 6 should not be based on answer accuracy alone. The review needs to examine business benefits, execution quality, risk records, and all costs together. All costs should cover model calls, system integration, human review, operations and maintenance, and exception handling. If the direction of benefits holds and the main problems are fixable, adjust and continue; if key risks cannot be contained, or the workflow itself lacks stable inputs, stop expanding. A full launch is not the end of the project either; the model, knowledge content, and business rules need ongoing maintenance as workflows change.

Which use case should an enterprise prioritize for its first AI agent?

Prioritize workflows with clear boundaries, high frequency, recordable labor costs, verifiable results, and controllable error impact. The first pilot should not be a task with unclear cross-departmental ownership, unstable data sources, or where a single error could cause major losses. Being able to form a short closed loop matters more than how complex the use case is.

Does an AI agent have to use private deployment?

Not necessarily. The deployment model should be determined by data sensitivity, regulatory requirements, network boundaries, model capabilities, and operational capacity together. Sensitive data doesn't have to go into the model by default either; you can start with field masking, minimal data transfer, access isolation, and log auditing. If private deployment is truly required, you also need to confirm whether the enterprise has the capability to handle model updates, capacity management, and incident response.

When connecting an agent to ERP or CRM, how do you choose between API and RPA?

When a system offers a stable API, use the API first, because its parameters, return values, and permission boundaries are easier to verify. RPA suits legacy systems that lack interfaces and can't be modified in the short term, but page changes can break the workflow. When writes, approvals, payments, or status changes are involved, whichever approach you use, add identity verification, idempotency controls, operation auditing, and human confirmation.

How do you decide whether an AI agent pilot should be expanded or stopped?

The prerequisites for expanding a pilot are that business benefits are reproducible, the causes of failures can be explained, the burden of human takeover is acceptable, and risk events stay within preset boundaries. If critical errors keep recurring, system dependencies remain unstable over the long term, total costs have no room for improvement, or the system can only run with extensive manual remediation, narrow the scope or stop. Stop conditions should be written into the pilot plan in week 1, not changed on the fly when the results disappoint.