Teverant AI · Insights

2026-08-19

How to build an AI customer service bot: an enterprise integration and evaluation plan

A systematic guide to building an AI customer service bot, covering service boundaries, channel integration, knowledge base construction, dialogue flows, human handoff, phased rollout and evaluation metrics, to help enterprises improve customer service efficiency without taking on undue risk.

1. Define the service boundary first: which questions go to the bot and which must be handled by humans

Before AI customer service goes live, the first thing to settle is not model size but how far the bot is responsible for taking a case. If the goal is written simply as "automatically answer user questions," two kinds of deviation easily follow. In one, the bot gives overly confident judgments on high-risk matters; in the other, many questions that seem answerable actually still require looking up orders, verifying identity or initiating an after-sales process. If the boundary isn't clearly defined, the larger the knowledge base, the larger the scope for mishandling.

Divide the scope of automation by inquiry type

The first batch suited to the bot is usually high-frequency matters with stable rules, clear answer sources and a manageable cost of error: product specifications and usage limits, the current stage of an order, shipping status, eligibility conditions for promotions, and return and exchange policies the company has already fixed. What these scenarios have in common is that the user expects confirmation of information, not free judgment by a customer service agent.

Complex complaints, order problems involving large amounts, refund disputes and matters that require determining liability should be handed to humans by default. For instance, whether damage to a product occurred during warehousing, shipping or use often can't be determined from a single conversation. And if a user asks for an exception to existing policy and the bot makes promises on its own to keep the conversation smooth, it creates compensation, compliance and complaint risks down the line. Industry practice generally has automated systems take on repetitive, lower-risk requests and leaves complex disputes to staff with the authority and judgment to handle them.

Define a "handling outcome" for every scenario

Scenario design can't stop at "prepare a standard answer." The same question may correspond to different system actions. We recommend spelling out the following outcome types in the scenario table:

OutcomeWhen it appliesCompletion criteria
InformationParameters, rules, policy explanationsCites valid content; the answer stays within policy
Data collectionAfter-sales requests, problem reports, supplementary proofAll required fields collected, with privacy boundaries explained
System lookupOrder, shipping and account statusReturns the corresponding record after identity verification
Process initiationRefunds, exchanges, ticket registration and similar mattersTask successfully created and the user informed of next steps
Human handoffDisputes, risks, questions beyond the bot's authorityContext preserved; reason for handoff and expected wait explained

The value of this is that the system won't mistake "generated a paragraph of text" for "the problem is solved." For example, the completion condition for a shipping inquiry is not describing how to look it up, but confirming identity, calling order data and returning a verifiable result. Likewise, the completion condition for a refund dispute is not restating policy, but identifying the dispute type, collecting the necessary materials and routing it to the right queue.

Use question tiers to decide the initial scope

Before launch, build a question tiering table that records at least inquiry frequency, risk level, whether business data is needed, average human handling time and the expected outcome after automation. A tiered structure can be used:

  • High frequency, low risk: automate first; suited to standard policies, basic specifications and common status lookups.
  • Medium frequency, data required: handle through identity verification and business workflows; relying on knowledge base text alone is not recommended.
  • Low frequency, high risk: provide a direct human entry point; the bot only identifies the type and collects context and materials.

Choosing the first scenarios means weighing historical inquiry volume, human handling time and business value together. A matter with modest inquiry volume that takes a lot of time to handle each time may be more worth automating than a simple high-frequency Q&A; conversely, a very frequent scenario that involves liability for compensation shouldn't be opened to automated decisions just because users ask about it a lot. Getting the boundaries, data permissions and handoff conditions of a few scenarios working first, then expanding coverage, avoids an unmanageable number of knowledge base items and runaway process branches. It also makes it easier, after launch, to tell whether problems stem from insufficient answer quality or from incomplete process design.

2. Channel integration is not "dropping in a chat box": design the message, identity and business system chain first

When connecting a customer service bot to channels, first look at where customers come from, then decide what to connect first. The website or a standalone web page suits the first round of validation: traffic is controllable, changes are cheap, and it is easy to observe real questions. When existing customers mainly communicate on mobile or in enterprise collaboration tools, consider connecting entry points such as an app, WeChat Official Accounts, WeChat Customer Service, WeCom and DingTalk. Don't launch multiple entry points at once for the sake of "omnichannel coverage"; otherwise the same question may get different answers, and the human team will struggle to tell which chain a complaint came from.

The core of channel planning is not embedding a reply box but confirming a complete message path. For each entry point, verify at least the following capabilities item by item:

Link to confirmEngineering questions to settle
MessagingCan text, images and other messages be received reliably, and are replies subject to length, frequency or format limits?
Conversation recordsIs context between the user, the bot and human agents stored in one place, and can it be searched by conversation?
IdentityCan the customer identifier of a logged-in user be obtained, and how are anonymous users linked to follow-up service?
Human takeoverAfter handoff, can the conversation history, user profile and actions already taken be passed to the agent together?

The identity chain should be established at the entry point wherever possible. Bind logged-in users directly to an internal customer ID; when the bot looks up orders, membership benefits or after-sales records, don't rely only on a name the user types into the conversation. Anonymous visitors can first be linked via phone number, order number or a one-time session identifier, but avoid exposing full identity information in chat content for long periods. Data such as phone numbers, addresses and order details should also be masked as the business requires, with limits on which fields agents and the bot can each see.

Once a question involves real-time business status, the knowledge base is no longer the main data source. Information such as order progress, shipment tracking, refund requests and discount eligibility should be retrieved from order systems, CRM, ERP and logistics systems through controlled APIs or workflows. Before a call, verify the customer's identity and ownership of the resource; after the call, return only the results needed to complete the current task. For example, when a user checks an order, the system can return the status of orders under that user's name, but a single query should not open up all orders or internal notes. Write operations also need confirmation steps, idempotency controls and operation records, to avoid duplicate refund submissions or duplicate tickets.

Every channel needs designed failure paths. On an API timeout, first tell the user that real-time results can't be retrieved right now, and offer a retry or a handoff to a human. When the same message arrives twice, deduplicate by message ID to avoid duplicate replies and duplicate business actions. When a user goes unresponsive for a long time, save the context so the conversation can resume when they return rather than starting over. When identity can't be confirmed, stop calling sensitive APIs and switch to collecting the necessary information or handing off to a human. A bot that can't understand a question also shouldn't keep rewording its answer in a loop; after repeated failures it should state clearly what it can handle and keep a human entry point available.

Before launch, you can run acceptance item by item against a "channel → identity → system → human" chain diagram: can messages come in, can replies go out, are records traceable, is identity trustworthy, are business calls restricted, does someone take over after failures? Only when these basic chains are stable is it worth expanding channels and the scope of automation.

3. Building the knowledge base: from uploading documents to retrievable, maintainable answer assets

A knowledge base is not a matter of uploading a few PDFs, web pages and internal documents in one place and leaving the model to make sense of them. What customer service really needs is a set of answer units that can be retrieved accurately, have a clear scope of application and can be withdrawn at any time. Raw materials can serve as input but can't be the final form of knowledge; tables of contents, repeated statements, historical versions and missing context in long documents all amplify retrieval bias.

Split materials by business scenario first, then decide chunk granularity

The first step is to inventory material sources, including product documentation, FAQs, pricing and promotion rules, order and delivery information, return and exchange policies, service commitments and historical tickets. Then, rather than organizing content by file name, split it by user task, such as pre-purchase inquiries, product usage, order lookups, promotion verification and after-sales handling. Ideally, one user question maps to one clearly bounded knowledge unit, rather than leaving the retriever to guess the answer from an entire manual.

Keep the business conditions when splitting. For example, "returns are supported" is not a complete answer; you also need to state the applicable products, application deadline, product condition, proof requirements and exclusions. We recommend organizing content into the following fields:

FieldPurpose
User question and standard answerDescribes how customers might ask, and how the bot is allowed to answer
Applicable conditionsLimits by region, product, membership status, channel or time range
Exceptions and prohibitionsPrevents the bot from applying general rules to special cases
Effective date and versionDetermines whether the content is still valid
Source and ownerSupports review, accountability and future updates

Version management matters more than "the more content the better"

Pricing, refunds, compensation, contract terms and after-sales commitments are high-risk knowledge. Such answers must be traceable to a specific source and version and, where necessary, should retain the publication date, expiry date and approval records. Once a new policy goes live, old entries can't stay in the retrieval pool competing alongside new ones; they should be explicitly marked as expired, archived or removed from the online index. Otherwise, however fluent the model's language, it may cite a promotion that has ended or an outdated handling standard.

The basic chain of retrieval-augmented generation is to find relevant chunks in the knowledge content first and then compose a reply based on them. Retrieval results should therefore carry the applicable conditions and limits with them, rather than returning only a conclusion stripped of context. When the retrieved content isn't enough to determine a user's eligibility, the bot shouldn't fill in the gaps itself. It should first ask for key variables such as the order number, purchase channel, product condition and when the issue occurred; if the conditions still can't be confirmed, it should hand off to a human.

Run acceptance on real conversations, not just standard phrasings

Before the knowledge base goes live, draw an acceptance set from historical tickets and chat logs, covering different phrasings of the same question, typos, colloquial abbreviations, successive follow-ups, negative sentences and the edges of rules. For example, don't just test "how do I return an item"; also test "can I still return it if I've opened it," "can I apply if I didn't place the order myself" and "can I return something I bought at a promotional price." Each sample should record at least the expected answer, the conditions that must be cited, the fields the bot may ask about and the situations that should be handed to a human.

After launch, unresolved conversations, user corrections, repeated follow-ups and answers rewritten by human agents are all signals of knowledge gaps. Summarize these records weekly and sort them into three types of problem: "no matching content," "content exists but wasn't retrieved" and "retrieved correctly but the answer overstepped." Then add entries, adjust chunking and revise answer constraints accordingly. A knowledge base built this way is an answer asset that can be retrieved, audited and continuously maintained as the business changes.

4. Dialogue flows and human handoff: teaching the bot when to continue and when to stop

Where customer service bots most often go wrong is not in failing to generate answers, but in continuing the conversation when they lack what they need to handle it. Companies should first break each type of business into explicit handling flows, then specify for each step the information required, the actions allowed and the exit conditions. The goal of a dialogue flow is not to have the bot say as much as possible, but to have it move forward when it has enough information and hand over to a human promptly when risk rises.

1. Drive business flows with a "minimum information set"

Each flow should define the minimum fields needed to make one decision, avoiding both collecting too much data up front and reaching conclusions when key conditions are missing. Taking refunds as an example, the flow can usually follow the sequence "confirm the order → understand the reason → verify the status → decide the path":

  • Confirm the object: obtain the order number and, where necessary, confirm order ownership using the logged-in identity, phone number or other verification information.
  • Understand the request: record the refund reason and distinguish cases such as not yet shipped, shipped, delivered and duplicate purchase.
  • Verify conditions: check the order status, payment status, after-sales window and whether a request is already in progress.
  • Execute the branch: if the rules are met, provide the entry point or initiate the next step; if not, explain the specific limitation; if the system can't decide or the user disputes the outcome, hand off to a human.

The key here is not to make the flow complex but to give every branch a basis. LLMs are good at phrasing and asking follow-up questions, the knowledge base supplies policies, product rules and similar content, and the dialogue flow constrains sequence and action boundaries. In practice, define "can look up," "can explain," "can submit a request" and "must be approved by a human" separately.

2. Write down stop conditions in advance

Handing off to a human shouldn't be a button hidden at the bottom of the page; it should be a formal branch in the flow. We recommend configuring at least the following triggers:

  • Retrieval results are insufficient to support a clear answer, multiple policy clauses conflict, or answer confidence falls below the threshold set by the business;
  • The user directly says "get me a human" or "complaint," or asks for a supervisor;
  • Clearly negative expressions appear, such as anger, threats or repeated rejection, especially in scenarios involving compensation, service failures or harm to the user's interests;
  • Sensitive information is involved, such as ID documents, account security, payment disputes or medical and health matters, or the matter requires human authorization or exception approval;
  • The content is something the rules explicitly forbid the bot to promise, such as setting a compensation amount, committing to a handling deadline, changing contract terms or guaranteeing that a problem will be solved.

These conditions should be recordable and reviewable, not left solely to the model's own judgment. For complex disputes, problems with high-value orders and heated complaints, industry practice generally still has humans bear final responsibility, with the bot mainly covering simple, high-frequency requests with stable rules.

3. When handing off, move the context, not the customer

If customers have to explain their problem all over again after being handed to a human, the automation before that point merely delayed handling. A handoff event should generate a structured summary that includes at least the following:

ContextWhat should be recorded
User requestThe customer's original question, the outcome they currently want, and complaint or sentiment tags
Conversation progressConclusions the bot has already given, rules it has cited, and options the customer has explicitly rejected
Business dataConfirmed fields such as order number, order status, refund reason and identity verification result
System callsAPIs queried, results returned, reasons for failure and the time of the last update
Recommended actionItems the human agent should verify, policies that may apply, and next steps

The summary should both preserve the basis for judgments and distinguish what was "provided by the user," "retrieved from systems" and "inferred by the model," so agents don't mistake speculation for fact. Where sensitive data is involved, it should be masked according to permissions; a handoff must not widen the circle of people who can see the data.

4. Human takeover needs service levels too

A handoff to a human is not the end of the flow; companies also need to define service commitments after takeover. Once users enter the queue, show the current status and estimated wait time. Outside working hours, provide a message, callback or ticket option and state the expected response window. For urgent issues such as payment risk, account theft or major security incidents, set up a higher-priority escalation channel. If the wait clearly exceeds expectations, let users choose to keep waiting, leave contact details or switch to another channel.

In the early stage after launch, focus on checking three kinds of records: whether the bot stops when it should, whether handoff summaries are sufficient for a human to take over, and whether human agents frequently correct the bot's earlier judgments. A lower handoff rate is not better in itself; if handoffs decrease while complaint escalations, repeat inquiries or human rework increase, the system is merely hiding problems. A truly usable flow lets the bot handle the steps with higher certainty and clearly hands uncertainty and questions of responsibility over to humans.

5. Rollout approach: validate with a phased rollout and human quality checks, then gradually expand automation

AI customer service isn't suited to covering all inquiries at once. The core of going live is not flipping a switch but controlling the risk boundary at each stage: first confirm that the bot can reliably handle low-risk questions, then let it take on scenarios that require reading business data and performing operations.

1. Advance by risk tier, not by number of features

In the internal testing phase, use historical conversations, masked samples and manually constructed questions, focusing on high-frequency content with relatively fixed answers, such as business hours, delivery areas, return and exchange rules and common usage instructions. During the low-traffic phased rollout, let through only a small share of real conversations and keep a human takeover entry point. If answers to standard questions are stable and misleading answers are under control, then choose a channel with clear business boundaries for the official launch.

Scenarios such as order status lookups, refund progress and after-sales requests should be opened only after a single channel is running stably. These requests typically involve identity verification, API calls and permission checks and can't be completed with model-generated text alone. Expanding to multiple channels also shouldn't mean simply copying configurations; re-check each channel's user identity, message formats, session timeouts and how human agents take over.

2. Before launch, use a checklist to verify both "can answer" and "knows when to stop"

Testing can't just check whether the bot answers FAQs correctly; it must also verify whether it pulls back when uncertain. Cover at least the following scenarios:

  • Standard questions: whether answers are accurate and key conditions and limitations are complete;
  • Vague phrasing: when user information is incomplete, whether the bot asks follow-up questions first rather than guessing;
  • Unauthorized requests: before identity is confirmed, whether it refuses to look up other people's orders or modify account details;
  • Policy validity: whether old rules, discontinued promotions and expired commitments are identified and blocked;
  • Prompt attacks: when users ask it to ignore established rules or reveal internal instructions, whether it holds to business boundaries;
  • Sensitive information: whether ID numbers, contact details, payment information and the like are kept from being echoed back when unnecessary;
  • API failures: when the order or after-sales system times out or returns empty values, whether it says it can't confirm right now rather than fabricating results;
  • Human handoff: whether it hands off promptly in cases of low confidence, repeated misunderstanding, an explicit request for a human, or complaints.

3. Keep a human bypass during the phased rollout

During the phased rollout, it is best to have people monitoring conversations in real time, at least spot-checking high-risk conversations and random samples. Quality reviewers shouldn't record only "correct" or "incorrect," but classify by what happened next: resolved, only partially completed, wrong conclusion given, replied without actually helping, and kept answering automatically when it should have handed off to a human. The last category deserves separate tracking, because it tends to generate complaints more readily than an isolated case of poor wording.

The human bypass should also support quick corrections: agents can add the correct answer, flag the trigger conditions and record why the user was dissatisfied. For operations involving orders, refunds and accounts, context should be preserved after human takeover so users don't have to describe the problem again. If a certain type of error keeps recurring, first narrow the scope of automation, then fix the knowledge or the flow; don't use more templates to paper over the root cause.

4. Close the iteration loop with weekly reviews

Each week, pull unresolved questions, negative ratings, content rewritten by humans and recent policy changes from the conversation logs, and determine for each whether the cause is a knowledge gap, a retrieval failure, incomplete flow design or insufficient prompt constraints. Once the cause is confirmed, update knowledge content, add examples, adjust handoff conditions, and compare different answer strategies through A/B tests. In industry practice, ongoing log reviews, regular knowledge additions and validated prompt adjustments usually produce steadier improvement than piling in large volumes of material all at once.

Every release should keep a version record and a rollback plan, noting at least what changed, the channels it applies to, the scenarios affected and the quality review conclusions. Only when the error rate, handoff anomalies and performance in sensitive scenarios are all within acceptable ranges is it appropriate to increase traffic; otherwise, treat the phased rollout as an observation period rather than forcing the launch forward.

6. Evaluation: build a unified metric system around first response, resolution rate and handoff rate

After AI customer service goes live, the most common misjudgment is taking "replies quickly" to mean "serves effectively." A bot can send a welcome message instantly or hand out template answers in bulk, but that doesn't mean the user's problem was understood, let alone solved.

1. First response time: count only the first reply that carries real information

First response time is the time from when a user initiates an inquiry to when they receive the first valid reply. When measuring it, separate the bot's automatic first response, the first response after human takeover, and exceptions such as API timeouts and message delivery failures. Exception records can't be counted as normal responses, and welcome messages with no business information, such as "Hello, how can I help you?", shouldn't count as valid first responses.

We recommend recording the mean, the median and the slower percentiles together. The mean is useful for observing overall trends, the median better reflects the typical experience, and the slower percentiles can expose queuing at peak times, channel callback delays or congestion in human takeover. Measure each channel separately: web, WeCom, phone-to-text and other entry points have different messaging mechanisms, and mixing them masks real problems.

2. Resolution rate: define the denominator first, then what "resolved" means

Resolution rate can't be replaced by the auto-reply rate. The former concerns whether the user's problem was closed out; the latter only shows how many automatic messages the system sent. Companies should first define the denominator, for example by including conversations with a clear request that completed at least one valid round of interaction and excluding ads, duplicate messages and clearly invalid inquiries, and then set the rules for judging resolution.

A safer approach combines several signals: no human intervened before the conversation ended, the user explicitly confirmed resolution, there was no follow-up about the same problem within an observation window afterward, or the related ticket was closed normally. Different businesses can use different weights, but the rules must be fixed, with a sampling review mechanism in place. For high-risk scenarios such as refunds, account security and compliance complaints, even if the bot gave an answer, the end of the conversation alone shouldn't count as resolution.

MetricRecommended definitionWhat to check
First response timeFrom inquiry initiation to the first valid replyWhether welcome messages or exception messages are wrongly counted
Resolution rateShare of in-scope conversations that were closed outDenominator scope, repeated follow-ups, ticket status
Handoff rateShare of valid conversations with a human takeoverWhether a handoff was warranted, whether it was resolved afterward, how long the wait was

3. Handoff rate: high isn't necessarily bad

The handoff rate needs to be interpreted in light of question types. Missing product information, intent recognition errors and incomplete flow configuration can cause the bot to hand off frequently; but in complaints, complex after-sales cases, identity verification and high-risk transactions, promptly stopping automated answers and passing the case to a human is actually the right behavior. So don't set a "the lower the better" target divorced from business scenarios.

In your analysis, add at least three metrics: whether questions that should be handed off were successfully transferred, whether they were resolved after transfer, and how long users waited in the handoff queue. Also break results down by intent, customer tier, channel and time period, to identify two types of error: "should have been resolved automatically but was handed off" and "should have been handled by a human but the bot kept asking questions."

4. Establish supporting metrics and a baseline for comparison

First response, resolution rate and handoff rate are the primary metrics, but they can't replace a complete operational view. We recommend also tracking answer accuracy, first-contact resolution rate, average time from question to closure, user satisfaction, human ticket volume, overnight coverage and cost per conversation. Accuracy can be assessed through human spot checks and hit rates on key intents; first-contact resolution specifically tracks whether users need to repeat themselves or add information over multiple rounds.

For every metric, keep a pre-launch baseline first, then compare post-launch data for the same period; also distinguish new and returning customers, working hours and nighttime, pre-sales and after-sales, and different entry points. Public project case studies commonly report noticeably faster responses, some common questions handled independently by the bot, and reduced pressure on human staff at night, but these results can only serve as a reference for target ranges and can't be promised directly to another company. Ultimately, they should be revalidated with several consecutive weeks of the company's own real conversations, human review and cost accounting.

7. FAQ: the 4 questions enterprises ask most before adopting AI customer service

Can a company without a complete FAQ launch AI customer service directly?

Yes, but automated answers shouldn't be opened to all customers right away. The number of historical Q&A pairs is not a hard threshold for launch; the real prerequisites are whether the bot can answer by citing controlled materials, and whether it can refuse or hand off to a human when materials are missing.

Companies short on materials can start from three kinds of content. The first is product documentation, including feature descriptions, usage limits, billing rules and troubleshooting steps; the second is the website FAQ, help center and service policies; the third is human tickets, from which you can extract questions that are asked often, have stable answers and follow clear handling paths. When organizing this content, don't just upload whole files; split it into retrievable knowledge units and tag each with the applicable product, customer type, effective date and responsible department.

The first batch of knowledge doesn't need to cover every question. We recommend starting with highly repetitive, lower-risk scenarios, such as finding features, operating instructions and order status explanations. Matters whose answers depend on specific conditions, such as refund disputes and contract interpretation, should go to humans first. Before launch, also test with real user phrasing, because customers rarely phrase questions the way document titles are written.

How should the resolution rate of AI customer service be calculated?

The resolution rate usually refers to the share of bot-handled conversations that needed no further human handling and in which the user's problem was closed out. When calculating it, you must first define "resolved." Merely sending an answer doesn't count; the user confirming the issue was handled, completing the target operation, or not contacting again about the same problem within an agreed observation period comes closer to genuine resolution.

The auto-reply rate indicates how many of the conversations entering the customer service system received a reply generated by the bot. It reflects the reach of automation, but a reply is not the same as a correct answer. The handoff rate indicates the share of bot conversations that ultimately entered the human queue, and is used to observe the limits of the bot's capabilities and gaps in its knowledge.

The three metrics can't be judged in isolation. A rising auto-reply rate alongside a falling resolution rate may just mean the bot is intercepting more questions; a very low handoff rate with rising repeat inquiries may mean the handoff conditions are too strict; and a high resolution rate that covers only a small number of simple inquiries doesn't show an improvement in overall service efficiency either. Evaluation should look at first response time, the number of user follow-ups, the repeat contact rate and handling time after human takeover together, grouped by question type and channel, so averages don't mask high-risk scenarios.

When must a conversation be handed off to a human?

Scenarios involving complex after-sales disputes, strong complaints, problems with high-value orders, and anything that requires identity verification, liability determination, compensation approval or exception authorization should remain with humans. The bot can collect the order number, a description of the problem and the necessary proof, but it shouldn't commit on its own to refund amounts, liability or handling deadlines.

Handoff rules should have at least two layers. The first is deterministic rules: hand off directly when intents such as complaints, legal risk, personal safety or account funds are detected. The second is a confidence fallback: stop generating and hand off to a human when no reliable grounding can be retrieved, multiple answers conflict, the user repeatedly rejects answers, or the bot still can't confirm intent after several rounds.

A handoff can't just display "please contact a human agent." The system should pass the conversation summary, identified intent, knowledge cited, user identity and business documents to the agent together, so the customer doesn't have to tell the story again. When no human is available, the queue status, expected response method and follow-up contact channel should also be made clear.

How often should AI customer service performance be evaluated after launch?

When product launches, price adjustments, service policy changes or clusters of complaints occur, don't wait for the regular cycle; run a dedicated review immediately.

Logs to keep include the user's original question, the bot's reply, the retrieved knowledge chunks and their versions, confidence, the number of dialogue turns, the reason a handoff was triggered, the agent's subsequent replies, user feedback and the final handling status. Where personal information is involved, set up masking, access permissions and retention periods; don't keep raw data indefinitely for the convenience of training.

Human corrections can't stay buried in chat logs. Each change should retain its version, owner, test cases and go-live time, so that human experience turns into knowledge base and process improvements that can be verified in the next round.