Teverant AI · Insights

2026-08-10

Enterprise AI agents: use cases, process, and a deployment plan

A systematic guide to enterprise AI agent applications: use case screening, goal definition, system integration, access control, performance evaluation, and the complete method for going from POC to production.

1. Decide first: which business scenarios deserve an agent first

When a company starts an agent project, the first step is not choosing a model but confirming whether the business actually needs "answers" or "execution." If, after a user asks a question, all that is needed is a knowledge explanation, a document summary, or search results, an ordinary Q&A system is usually enough. An agent becomes necessary only when the task requires the system to break down steps, select and call enterprise tools, read execution results, handle exceptions, and then carry on with subsequent actions.

The engineering difference between the two directly affects cost. The main path of a Q&A system is "input → generate → return"; an agent must manage task state, tool interfaces, identity and permissions, failure retries, and operation records. Wrapping a simple knowledge query in an agent does not naturally create more business value; it adds latency, points of failure, and governance burden. Use case screening should therefore start from the process itself, not from "which work can be connected to a large language model (LLM)."

Processes suited to a first batch of projects usually have all of the following characteristics:

  • High frequency: they occur continuously every day or week, so the time saved on each instance accumulates into a steady return.
  • Repetitive manual actions: the main work is looking up data, filling in fields, checking against rules, generating materials, or moving information between systems.
  • Describable steps: most branches have clear conditions, and experienced employees can write down an operating checklist and exception-handling procedures.
  • Already-digitized inputs: the required content exists in databases, business systems, document libraries, or accessible external data sources.
  • Verifiable results: you can determine whether a ticket was closed, a discrepancy was found, or a report was generated, rather than judging by subjective impressions alone.

Screened by these conditions, classifying and routing customer service requests, verifying accounting discrepancies, compiling business reports, triaging security or operations alerts, and collecting and analyzing online information are usually better starting points than strategic judgment or complex negotiations. The former have stable inputs, clear actions, and checkable outputs; the latter tend to depend on tacit experience, and accountability for them is hard to hand to an automated system.

Evaluation dimensionQuestion to answerSignals to prioritizeSignals to defer
Business valueHow many labor hours does the process consume, and what losses do errors or delays cause?Stable volume, long wait times, calculable labor costsSporadic demand; benefits can only be described as "a better experience"
Degree of standardizationCan the process be mapped with its main branches listed?Clear rules, a limited number of exception typesEvery case relies on on-the-spot judgment; methods vary from person to person
Data and system readinessWhere are the inputs, are the interfaces usable, and what is the data quality?Data is already structured; key systems support controlled callsLarge amounts of material are missing and must first be entered or identified by hand
Risk controllabilityCan execution errors be detected, reversed, and attributed?Actions can be previewed before confirmation, rolled back, and logged end to endErrors are irreversible and involve major financial, compliance, or personal-safety liability

In an actual review, set hard thresholds first and then do a composite score. If key data is unavailable, operations cannot be audited, or the consequences of errors cannot be recovered from, the project should not be launched just because its business value looks high. Scoring is for ranking, not for covering up missing fundamentals. The first project in particular should avoid spanning multiple departments, multiple core systems, and multiple sets of approval relationships at once; otherwise the team will struggle to tell whether problems come from the model, the interfaces, or the division of organizational responsibility.

Benefit assessment should also focus on the "process compression ratio," not just on answer speed. In publicly reported company cases, after information on trending products was collected automatically, organizing work that used to be measured in hours was compressed to minutes; after attendance data aggregation, rule checks, and payroll calculation were chained into an automated process, the payroll cycle for a team of many people dropped from days to minutes. Because the materials for these cases do not provide verifiable report names and years, these numbers should not be used directly as a budgeting basis, but their benefit structure is worth referencing: an agent's value comes from reducing handoffs, waiting, and repetitive operations, not merely from generating text faster.

So an agent use case worth validating first should be expressible as a clear engineering proposition: given which inputs, allowed to call which systems, performing which actions under which constraints, and accepted against which metric in the end. If these questions still cannot be answered, sort out the process and data first rather than rushing to build an agent.

2. Start with one quantifiable task: define the agent's goals and boundaries

Agent projects most easily lose focus at the first step. "Improve customer service efficiency" and "help employees find information faster" can serve as directions, but not directly as acceptance criteria. In engineering terms, first pick a task that can be observed and reviewed, then break the business goal down into process metrics and outcome metrics. Otherwise, it may look like it answers questions in a demo, yet once in production there is no way to tell how much time it actually saved or how much manual work it reduced—or whether it just added new review work.

Rewrite the goal as task metrics first

Take customer service: you should not focus only on whether answers are correct. At a minimum, track time to first response, first-contact resolution rate, share of handoffs to human agents, answer usability, and the fully loaded cost per service interaction together. Accurate answers with frequent handoffs to human agents point to problems in knowledge coverage or process orchestration; fast responses that mislead customers can amplify after-sales risk. Metrics also need a clearly defined measurement basis—for example, whether "resolved on first contact" means the customer did not follow up again, or the ticket was closed within the specified time—rather than letting each team interpret it its own way.

Goal layerSuggested questionsAcceptance method
EfficiencyIs customer wait time shorter, and has human handling time dropped?Compare against similar tickets before launch
QualityAre answers grounded, and can they be used directly to handle business?Sample reviews with error types recorded
Business outcomeIs the problem resolved in the current session, and are repeat reassignments reduced?Trace the chain of sessions, tickets, and handoffs to human agents
Cost and riskHow do the costs of each call, human review, and error handling change?Calculate per closed-loop task, not just model fees

Accuracy can serve as a threshold, but it cannot be the only conclusion. In enterprise customer service practice, the business usually requires replies to reach a high level of reliability while also being professional, concise, and directly usable by employees. Set up an error taxonomy: factual errors, outdated citations, missing conditions, unauthorized access, unclear wording, and failure to hand off when required. Only then can you tell whether a problem arises in retrieval, the prompt, tool calls, or the business rules themselves.

Write the agent up as a reviewable task contract

Before development, pin down five things in a document: what input it receives, what result it should produce, which systems it can call, which actions it is allowed to perform, and in what situations it must hand off to a human. Inputs should specify source and format—for example, whether the customer's question, order number, vehicle model, or contract number must be present. Outputs should specify fields, cited evidence, and status, rather than vaguely requiring "the best answer."

  • Callable tools: expose only the query, retrieval, or submission interfaces needed for the current task.
  • Permitted actions: distinguish between reading, drafting, submitting changes, and final confirmation, and start with low-risk actions by default.
  • Handoff conditions: stop automated processing when refunds, complaint escalations, unverifiable identity, conflicting knowledge, insufficient confidence, or system anomalies are involved.
  • Audit requirements: retain inputs, retrieved content, tool parameters, execution results, and records of human takeover.

"Automation" does not mean handing an employee's account over to the model. An agent should be like a constrained digital employee: it has a job scope, operating limits, approval nodes, and clear conditions for stepping away. Guard especially against prompt injection, erroneous calls, and overly broad permissions stacking up into irreversible operations. Any action that modifies master data, makes external commitments, incurs costs, or affects customer rights should have human confirmation or a rollback mechanism.

Plan the timeline by complexity, and don't underestimate knowledge preparation

A simple FAQ task can usually be planned in weeks; once the scenario expands to multiple processes such as orders, after-sales, and inventory, the timeline often moves to a monthly scale; when multiple specialized agents, task dispatching, and state coordination are involved, it should be managed as a quarter-scale project. The truly time-consuming part is often not connecting the model but the knowledge cold start and RAG build-out: inventorying materials, deduplicating versions, partitioning permissions, and testing chunking and retrieval all require participation from business staff. In industry practice, this part often accounts for a large share of total implementation effort.

The POC must resemble production, not just test ideal questions

A POC should first be limited to one business queue, using anonymized real historical questions and real process data, and keeping interference such as typos, missing context, repeated follow-ups, and outdated materials. Public experience at automakers has already shown that a Q&A system that performs well in controlled testing may, once it reaches the customer environment, face multiple vehicle models, thousands of pages of manuals, and delayed updates to new product knowledge all at once; what really separates good systems from bad ones is often a small number of high-risk edge-case questions.

So at the end of a POC, you need to answer at least three things: which tasks can be completed automatically and reliably, which must be handed off to a human, and whether the main bottleneck to further investment is data, interfaces, or rules. Only after this round of validation against real business results does the agent have the foundation to move on to system integration, permission design, and operations at scale.

3. System integration: getting the agent into real enterprise workflows

Whether an agent is truly in production cannot be judged by whether it can generate a decent paragraph of text, but by whether it can obtain business data, make judgments, invoke system actions, and write the results back into the original process. An application that can only answer questions in a standalone chat window is essentially still an information assistant; in production, an agent usually needs to connect to the knowledge base, ERP, CRM, ticketing platform, file storage, browser, code execution environment, and internal APIs.

Integration design should start by mapping out the task chain, not by piling on tools. Take "processing a customer refund request": you need to specify where the order and communication records are read from, which rules the judgment is based on, which interface is called to create the approval, how failures are retried, and which system the final status is written back to. Every node should define its input fields, output structure, timeout policy, and where exceptions go, so the model does not guess system parameters in natural language.

Integration targetEngineering focusCommon failures
Enterprise knowledge and filesParsing, chunking, indexing, citation location, and version updatesDocuments are retrievable, but table relationships, section hierarchy, or image information is lost
ERP, CRM, and other structured systemsField mapping, query constraints, consistent definitions, and result validationThe same metric has multiple definitions, and the agent returns numbers whose source it cannot explain
Tickets and approval workflowsState machine, callbacks, idempotency control, and human takeoverInterface retries create duplicate tickets, or the process cannot recover after an interruption
Browsers and code executorsRuntime isolation, access scope, and artifact retentionWeb page changes break scraping; generated scripts lack execution limits
Internal APIsParameter validation, error-code handling, call logs, and compatibility strategyThe model assembles parameters directly, and after an interface upgrade it silently produces wrong results

Knowledge integration cannot rely on demoing with a handful of Word or PDF files. Materials in production often include scanned documents, spreadsheets, slide decks, archived web pages, images, and large composite files. During acceptance, check whether the relationship between headings and body text is preserved, whether tables spanning pages can be reconstructed, whether image content is searchable, whether attachments are linked to the main document, and whether old indexes are invalidated promptly after updates. Otherwise, even when retrieval hits the right document, it may return fragments missing their context.

Structured data should preferably be connected through a controlled query layer, not by letting the model access the production database directly. Managers can use natural language to query order changes, business metrics, or financial details, but the system still needs to convert the question into a restricted query, validate the time range, organizational scope, and metric definitions, and then return the data along with its source. For process-type tasks, you also need to add data collection, computation, report generation, exception flagging, and writing results back, upgrading "looking up information" to "getting the work done."

For tool integration, adopt a uniform contract: declare each tool's purpose, parameter types, return structure, permission requirements, and side effects. Read calls and write calls should be separated; actions such as creating, deleting, paying, and sending messages must have an idempotency key and a clear confirmation mechanism. The entire call process should leave behind a task ID, the model's decision, tool parameters, the system response, and the final status, so you can pinpoint whether a problem occurred in retrieval, reasoning, or interface execution.

Architectural complexity should be determined by task dependencies, not by the number of agents:

  • For short-chain tasks such as knowledge Q&A or single-system queries, a single agent with retrieval and tool calling is enough.
  • When high-frequency simple requests coexist with low-frequency complex judgments, route between small and large models: the small model handles classification, extraction, and fixed replies, while the large model handles reasoning and disambiguation.
  • Adopt multi-agent collaboration only when requirements analysis, task execution, and result verification can each have their inputs and outputs defined independently, and need to be scaled or audited separately. Protocols such as A2A are suited to carrying communication between roles, but they cannot replace state management and failure recovery.

In practice, build out in this order: "read-only queries → controlled generation → execution after human confirmation → automatic execution when conditions are met." First verify that data can be read reliably and results can be traced, then gradually open up write capabilities. The completion standard for system integration is not that the interfaces are connected, but that a task can be executed as a closed loop within its permissions, can be paused, rolled back, or handed off to a human when an exception occurs, and leaves a record at every step.

4. Access control: manage the agent as an "executing digital employee"

The risk of an enterprise agent comes not only from wrong model answers but from wrong answers being turned directly into system actions. An agent that can call databases, tickets, email, or finance interfaces already has execution privileges in substance. So permission design cannot stop at "who can use it"; it also has to answer three questions: what it can read, what it can execute, and how to stop and trace it when something goes wrong.

1. Derive least privilege from the task boundary

Do not open the entire database, the complete file directory, or every API to the agent just because it is convenient to connect. First list the minimum data and tools needed to achieve the goal, then restrict systems, objects, fields, actions, and time ranges separately.

  • Data scope: expose only the necessary business domains, records, and fields. A customer service agent can read order status but usually does not need to see payment vouchers, full ID numbers, or employee information.
  • Tool scope: query, create, modify, and delete should be split into separate permissions; being allowed to read tickets must not imply being allowed to close or bulk-delete them.
  • Identity scope: the agent should use its own service identity, not reuse an administrator account or employees' long-lived credentials. Different agents, environments, and business departments should not share the same key either.
  • Runtime scope: set limits on call volume per invocation, batch size, operation frequency, active hours, and spending, so that an error loop cannot turn into large-scale writes or continuous resource consumption.

Permission checks must be enforced by the tool gateway or the business system, not by relying on "do not access sensitive data" in the prompt. A prompt is behavioral guidance, not a security boundary.

2. Let the risk of each action determine whether human review is needed

Handling levelApplicable actionsControls
Automatic executionLow-risk, reversible queries or updates with clear rulesLimit scope and frequency; log after execution and spot-check
Human confirmationPayment requests, contract changes, payroll processing, customer commitments, data deletion, external publishingGenerate a draft of the operation first, showing the target, rationale, and scope of impact; execute only after an authorized person approves
ProhibitedBypassing approvals, escalating its own privileges, disabling audit, exporting large volumes of sensitive data, and similar behaviorBlock hard at the policy layer; never expose the corresponding tools or credentials to the model

Human confirmation cannot be just an "approve" button. The approval page should clearly show the tool the agent is about to call, the key parameters, the before-and-after difference in the data, and the potential impact. For high-risk actions, also use two-person review, tiered limits, or step-up authentication.

3. Audit records must be able to reconstruct a task

Production audit should at least link together the user request, retrieved materials, prompt and model version, tool call parameters, model inputs and outputs, policy blocks, approvers, execution results, and exception information. Only with a unified task identifier across every step can you replay the full chain after an incident, rather than seeing only the final answer.

For write operations, also save the pre-change state, the idempotency identifier, and the compensation plan. Reversible actions should support automatic rollback; for actions that cannot be undone directly, define freezing, reversal, or manual handling procedures in advance. The logs themselves need access control, masking, tamper protection, and retention periods, so the audit system does not become a new outlet for sensitive information.

4. Keep model risk out of the execution path

Prompt injection can hide in web pages, emails, attachments, or knowledge base documents, so external content must be treated as untrusted input. Retrieval results must not modify system instructions or expand permissions on their own. Tool parameters should go through type validation, business rule checks, and sensitive-field filtering; model-generated content must never be concatenated directly into database statements, scripts, or API requests.

The knowledge base needs controls over write sources, review processes, and versions, to prevent incorrect content from influencing decisions over the long term. For hallucination, require key actions to cite verifiable data records; when evidence is insufficient, data conflicts, or confidence conditions are not met, the agent should stop and hand off to a human rather than guess to fill the gaps.

5. Permission governance must also cover production operations

After launch, the control plane should not be about security alone. The team also needs to monitor call costs, execution success rates, approval backlogs, abnormal retries, and the availability of external systems, and configure budget alerts, concurrency limits, circuit breaking, degradation, and failover. SLAs should also be defined at the task level—for example, handling time limits, failure recovery time, and conditions for human takeover.

Finally, set up a regular review mechanism: clean up idle accounts and keys, check for attempts at unauthorized access, revoke tool permissions that are no longer needed, and update policies based on incidents and operator errors. One simple standard tells you whether the permission design is adequate: even if the model's output is completely wrong, the system can still contain the impact within a range that is detectable, stoppable, and recoverable.

5. Performance evaluation: look beyond answer accuracy to whether the task gets done

The object of agent evaluation is not a single answer but an end-to-end task. If the content of an answer is correct but the customer's intent was not recognized, the result was not written back to the business system, or employees still have to recheck and re-enter it, the task cannot be counted as complete. Companies should split evaluation into three layers—model, process, and business—answering "is the output trustworthy," "is the process a closed loop," and "does the investment produce business returns" respectively.

Evaluation layerSuggested metricsQuestions to answer
Model layerAnswer accuracy, share of verifiable citations, factual error rateIs the content correct, is the evidence real, is anything fabricated?
Process layerEnd-to-end completion rate, share of human intervention, exception rate, average handling timeCan the agent complete the process independently, and at which step do failures occur?
Business layerHours saved, cost per task, business conversion, incremental revenue, customer satisfactionCompared with the previous approach, is there an attributable improvement in business results?

The three layers of metrics cannot substitute for one another. Higher accuracy in knowledge Q&A does not necessarily mean shorter handling times for customer service tickets; growth in call volume does not mean labor was saved. If employees frequently rewrite answers or have to re-enter data manually across multiple systems, process gains remain limited no matter how good the model metrics are.

Go-live thresholds should be set according to scenario risk, not with a single company-wide standard. For customer service, internal knowledge queries, and similar scenarios, correct answers, professional wording, concise content, and no need for employees to rewrite can be the baseline conditions. Where payments, credit, claims, contracts, or regulatory reporting are involved, amounts, parties, dates, rule versions, and approval status must also be checked separately, with human review nodes retained. High-risk actions should not be cleared on a composite average score alone, because a small number of serious errors can be masked by a large number of easy samples.

The evaluation set should come from real business, not only from ideal questions written by the project team. It can be grouped into routine tasks, long-tail questions, missing information, rule conflicts, system anomalies, and malicious input, and should cover different departments, channels, and time periods. Important business needs a combination of real ticket replays and continuous human spot checks: check not only the final result but also whether tool calls, cited materials, parameter values, and system writes are correct.

Business value should be confirmed through before-and-after comparison or group experiments. A safer approach is to keep a human baseline and have a test group use the agent, then compare completion time, rework rate, unit cost, and final output under the same task types, workload, and service hours. In industry cases, financial reconciliation being cut from days to hours, and daily throughput for insurance claims pre-review rising by an order of magnitude, are results that can be checked; but because the available materials do not provide verifiable report names and years, the precise figures should not be taken directly as a universal benchmark.

  • Do not use the number of logins, calls, or conversation turns to calculate ROI directly; they only describe usage.
  • Task completion rate must have a clearly defined denominator and distinguish fully automatic completion, completion after human confirmation, and human takeover.
  • Cost accounting should include model calls, retrieval, system integration, human review, failure retries, and day-to-day operations.
  • Improvements in revenue or conversion require a control group, to avoid attributing external factors such as marketing campaigns and seasonal changes to the agent.

Go-live evaluation is not a one-time acceptance. Companies should continuously collect employee edits, reasons for handoffs to human agents, user feedback, tool call failures, and abnormal tasks, building a replayable library of failure samples. The operations team can, on a fixed cycle, clean out invalid knowledge, standardize business terminology, adjust retrieval scope, refine prompts, and rerun the full evaluation set after changes to policies, products, or processes.

Finally, establish clear handling rules: when metrics are stable, gradually expand the scope of automatic execution; when a certain type of error keeps rising, downgrade to suggestion mode; when a key system fails or a high-risk field cannot be confirmed, hand off to a human immediately. Only then do evaluation results truly govern the agent's permissions and rollout pace, rather than just producing an accuracy report.

6. From POC to production: organization, platform, and operations for long-term reliability

A POC validates whether a model can complete tasks on controlled samples; a production system faces continuously changing knowledge, real interfaces, concurrent requests, and accountability tracing. Between the two is not a single deployment but the process of filling in a whole set of engineering capabilities. If knowledge is still uploaded manually on an ad hoc basis, interface failures have no fallback, and call costs cannot be broken down by business, the system should not be opened to all users, however good the demo looks.

Production capabilityMinimum requirementSignals to watch continuously
Knowledge maintenanceClear data sources, update cycles, expiration rules, and release approvalsRetrieval hit rate, share of outdated content, completeness of cited sources
System integrationFixed interface contracts, with timeout, retry, idempotency, and rate-limiting mechanismsCall success rate, response latency, number of dependency system failures
Cost managementRecord model and tool call consumption by department, scenario, and taskCost per task, budget burn rate, spikes in abnormal calls
Capacity and resilienceSupport horizontal scaling of instances, with backup paths for key dependenciesConcurrency capacity, queue backlog, recovery time, number of degradation triggers
Version governanceKeep separate versions of models, prompts, knowledge bases, and tool configurationsScope of change impact, rollback success rate, drift in live performance
Accountability tracingRetain inputs, decision process, tool operations, and human confirmation recordsAudit coverage, unauthorized access incidents, problem ownership, and time to resolution

Organizationally, the agent should not be driven by the technology department alone. The business owner sets the value target and is accountable for launch results; process experts define the normal path, exception branches, and human takeover points; knowledge and data administrators maintain content quality; the IT integration team secures identity, interfaces, and the runtime environment; security and compliance staff review permissions and the boundaries of data use; and the operations and evaluation team handles spot checks, metric reviews, and closing out issues. Major changes should be reviewed jointly by these roles rather than decided by developers on their own.

At the platform layer, observability comes first. In high-frequency call scenarios, you should be able to see the consumption of each type of task and set tiered budget thresholds; when load rises, the platform must automatically add compute instances while preventing downstream systems from being overwhelmed by bursts of requests. For key processes, design model switching, read-only mode, rules-engine fallback, and human takeover in advance. Service levels cannot just say "the system is available"; they should also cover end-to-end success rate, latency ceilings, recovery time, and audit completeness for key operations.

Pay-per-use pricing and serverless architecture suit early projects with unstable demand, reducing capacity planning and infrastructure maintenance work. But they solve the problem of resource supply; they cannot replace permission governance, cost allocation, version control, or failure drills. Especially as traffic grows, the team still needs to account for constraints such as cold starts, peak pricing, vendor quotas, and cross-region disaster recovery.

Rollout phaseSuitable scenariosGo-live threshold and exit conditions
Phase 1Internal Q&A, FAQ, a single approval or query processKnowledge is traceable and errors can be corrected by humans; if high-frequency questions cannot be answered reliably over time, pause expansion
Phase 2Customer service assistance, meeting summaries, business data interpretationMulti-system integration testing completed, with spot checks and human review in place; if the time saved does not cover operating costs, narrow the scope
Phase 3Specialized, high-risk tasks such as equipment inspection and production schedulingMust go through shadow runs, permission isolation, failure drills, and business sign-off; if unauthorized execution or unexplained anomalies occur, roll back immediately

The core judgment in moving from POC to production is not "the answers seem smarter" but whether the system keeps running reliably through knowledge changes, traffic fluctuations, dependency failures, and staff handovers. Every phase should have clear entry criteria, an observation period, and stop conditions; only when someone owns accountability, changes can be tracked, and failures can be contained does the agent truly become part of the enterprise's production systems.

7. FAQ: common questions about deploying AI agents in the enterprise

Should a company start with an ordinary AI assistant or go straight to an AI agent?

The deciding factor is not model capability but whether the task requires operating systems. If the need is mainly to search materials, summarize documents, generate drafts, or support judgment, starting with an ordinary AI assistant is more appropriate: the integration scope is small, and wrong results are easy for people to spot and correct.

An agent is needed only when the task includes explicit actions—for example, creating a ticket, updating customer status, initiating an approval, or generating a replenishment suggestion after checking inventory. Even then, it is not advisable to aim for full automation from the start; instead, open up capabilities in the order "read-only queries → generating suggested operations → execution after human confirmation → automatic execution under defined conditions."

A practical test: if a wrong model output would only produce a substandard paragraph of text, it is more like an assistant; if an error could change business data, trigger a flow of funds, or affect customer rights, it must be built with agent engineering, including identity, permissions, approvals, audit, and rollback.

Which enterprise processes are best suited to deploying AI agents first?

Prioritize processes with relatively standard inputs, stable execution steps, and verifiable results. Typical candidates include classifying and routing service tickets, enriching sales leads, checking contract clauses, creating requests after internal knowledge lookups, and collecting information across systems to generate business records.

When screening use cases, check five conditions:

  • The task occurs frequently, and manual handling carries an observable time cost;
  • The process boundaries are clear, and exceptions can be listed and handed off to a human;
  • The required data can be obtained through official interfaces or controlled tools;
  • Completion can be confirmed by system fields, approval results, or follow-up actions;
  • The impact of failure is contained, with paths for reversal, compensation, or human takeover.

Scenarios that should not be in the first batch usually fall into two categories: those that depend on large amounts of tacit experience, where even the business owner cannot articulate the decision rules; and those that involve irreversible high-risk actions, such as making payments directly, bulk-deleting data, or independently making major compliance decisions. For such processes, let the agent organize information and flag risks first, without granting it final execution authority.

How should permissions be designed when an agent connects to ERP, CRM, or internal systems?

Do not let the agent reuse an administrator account, and do not pass all of a user's permissions directly to it. A safer approach is to give the agent its own identity and split permissions into two parts: "the data scope it can access" and "the types of actions it can perform." Even if an employee can view complete customer records, the agent may not need to read every field; even if it can write to the CRM, that does not mean it can delete customers or bulk-change account ownership.

Access control should cover four layers:

  • Identity layer: distinguish the initiating user, the agent's identity, and the service that actually executes, and keep the full delegation chain;
  • Tool layer: authorize query, create, modify, and delete separately, and avoid unbounded general-purpose interfaces;
  • Policy layer: decide dynamically whether execution is allowed based on department, data classification, amount, time, and business status;
  • Execution layer: require human approval for sensitive operations, and record inputs, decision rationale, call parameters, returned results, and the final operator.

The production environment should also have call quotas, short-lived credentials, idempotency control, and circuit breakers for anomalies. When unauthorized access, duplicate submissions, or consecutive failures occur, the system should stop subsequent actions rather than let the model keep retrying on its own.

How do you judge whether an agent project is worth further investment?

Do not just count whether answers "look correct." What a company is paying for is task outcomes, so the object of evaluation should be the full chain: whether the agent understood the request, got the right data, chose the appropriate tool, executed successfully, and brought the business status to the expected state.

When the project is approved, define a human baseline first, then compare task completion rate, average handling time, share of human intervention, rework, the cost of recovering from exceptions, and resource consumption per task. Where customer or compliance risk is involved, also record erroneous actions, attempts at unauthorized access, exposure of sensitive information, and approval blocks separately; overall averages must not be allowed to mask high-risk failures.

Whether to keep investing comes down to three conclusions. First, do successful tasks bring confirmable business benefits? Second, are failures concentrated in fixable interface, knowledge, or process problems? Third, as usage grows, do human review and operations burdens remain manageable? If the agent can only run on demo examples and needs frequent human fallback as soon as the data changes, the problem usually lies not in model parameters but in task boundaries, system interfaces, or business rules that have not yet been engineered. In that case, narrow the scope and redesign rather than expanding deployment directly.