Teverant AI · Insights

2026-08-03

How to build a Telegram bot: a guide to enterprise automation integration

Want to know how to build a Telegram bot? This article walks through bot creation, token security, permission design, webhooks, session state, and monitoring and alerting, laying out how to integrate a bot into enterprise automation and take it into production.

1. Define the business boundaries first: what a Telegram bot should handle

When an enterprise builds a Telegram bot, the first step is not choosing a programming language but deciding which interactions deserve to go through the bot channel. Bots are best suited to tasks with a clear entry point, rules that can be described, and results that can be verified. If the requirement is still "understand every question and complete every operation," the project can easily turn into a general-purpose support bot with runaway permissions, chaotic state, and no way to audit it.

Start by dividing use cases into three categories, and define the response time, data sources, and failure handling for each.

Use case typeTypical triggersSuitable tasksMain boundaries
User-initiated inquiries and lookupsCommands, text, images, button actionsChecking order status, retrieving documents, submitting leads, viewing account statusMainly read-only lookups and structured data collection; hand complex judgments off to a human agent
Group management and notificationsMembers joining, command calls, rule matches, scheduled operationsWelcome guidance, FAQ responses, event reminders, handling rule-violating contentFirst define which group messages the bot can read and who is authorized to run admin commands
Business event messagesEvents generated by order, ticketing, inventory, or approval systemsShipping notices, ticket updates, anomaly alerts, to-do notificationsMessages convey status and should not replace the system of record in the business system

Telegram's platform constraints directly shape the design. A bot cannot create a direct message relationship with a user out of thin air; the user must first open a conversation with the bot and start it, usually by clicking a link and then sending /start. So even if an enterprise already has a user's account, it cannot assume it can send them a direct message. When proactive push is needed, complete the conversation binding in advance during registration, subscription, or a service flow, and store the usable chat ID, authorization status, and unsubscribe status.

Reading messages in groups is not open by default either. Under the mechanism described in Telegram's official bot documentation, when privacy mode is enabled, a bot mainly receives commands, mentions, and replies relevant to itself, not every ordinary message in the group. If the business genuinely depends on all group messages, the privacy setting can be changed, but you must also assess data minimization, notifying group members, log retention, and the risk of false triggers. If the bot only needs to post announcements or respond to commands, there is no need to widen its visibility.

During requirements review, describe every request consistently in four parts, "trigger source, processing rule, business system, reply action," rather than listing feature names. For example: the user taps the "Track shipment" button; the service verifies the binding between the user and the order; the order system returns the latest milestone; the bot displays the result and offers a path to a human agent. A group alert can be described as: the monitoring system generates an anomaly event; the rule engine deduplicates it and determines its severity; the on-call system creates an incident record; the bot sends a notification with an acknowledgment button to the designated group.

This way of describing requirements surfaces four kinds of problems early: whether the trigger is trustworthy, whether the rules can be executed deterministically, who owns the business data, and how a failed reply is compensated. Any requirement that cannot fill in all four parts of the chain should not go straight into development.

High-risk actions also need to be separated from ordinary conversation. At a minimum, these include:

  • Money-related operations, such as refunds, compensation, or balance adjustments;
  • Group governance operations, such as removing members, muting them, or changing admin permissions;
  • Core business changes, such as canceling orders, changing shipping details, or closing tickets;
  • Data export operations, such as bulk exports of customer, order, or conversation records.

These actions should not be executed immediately just because a sentence of natural language has been recognized. The minimum control is to show the target, the scope of impact, and the key parameters, and have the operator confirm again. When money, bulk data, or irreversible changes are involved, the request should also go through human approval, recording the initiator, approver, original request, execution result, and time. The resulting business boundary should be explicit: the bot receives intent, displays status, and initiates processes, while key decisions remain with the permission system and the business systems.

2. Create the bot and manage the token: treat the credential as a production secret

Creating the bot takes only a few minutes. What really affects production security is account verification, credential storage, and the ability to respond after a leak. Don't treat the bot token as an ordinary configuration value: whoever holds it can call the Bot API directly, sending and receiving messages as the bot or changing some of its behavior, so its security level should match that of a database password or a cloud platform access key.

Confirm the creation entry point first, so you don't hand your key to an impostor account

After searching for BotFather in Telegram, don't judge by the avatar or display name alone. Check two conditions: the account username must be exactly @BotFather, and it must carry Telegram's official verification badge. Once confirmed, send /start, then create the bot with /newbot.

During creation you will submit two kinds of names in turn:

  • Display name: Shown to users. It can use the business name, but should not be mixed up with internal environment names.
  • Bot username: Must be unique across all of Telegram, may contain only English letters, digits, and underscores, must be 5 to 32 characters long, and must end in bot.

If the enterprise maintains development, testing, and production environments, create a separate bot for each, for example by including dev, staging, or prod in the username. Don't let multiple environments share one token; otherwise test scripts may send messages to real users, and auditing and fault isolation become harder.

Fill in the basic profile to make the bot easier for users to understand

Having the token doesn't mean the bot is ready to open to users. At a minimum, complete the following configuration in BotFather:

CommandWhat it configuresEngineering recommendation
/setuserpicBot avatarUse a recognizable, stable icon so it isn't confused with a personal account
/setdescriptionDescription on the bot's profile pageClearly state its purpose, scope, and where to get support
/setabouttextShort introduction shown before a conversation startsSay in one sentence what the bot can do; don't list internal capabilities
/setcommandsCommand menuExpose only commands that are live and will stay available, and keep descriptions short

The command menu should be released in sync with the backend's actual capabilities. If the menu still lists commands that have been retired, users will read the lack of response as a system failure. If the bot handles business functions such as approvals or lookups, the introduction should also state the data scope and how to reach human support.

Inject the token only at runtime, and keep it out of your code assets

The API token returned by BotFather should go straight into a controlled secret-storage pipeline. Small deployments can inject it through environment variables; production systems are better served by a cloud key management service or an internal enterprise credential system, restricting read access to the bot's runtime identity and a small number of ops staff.

  • Never write the token into source code, configuration templates, image build files, or test examples.
  • Never commit it to a Git repository; deleting the current file does not remove the secret from commit history.
  • Logs, exception stack traces, and HTTP debug output must be redacted and should not record the full Bot API URL.
  • Don't pass the token around in group chats, ticket comments, or screenshots; if a handover is truly necessary, use a controlled secret-sharing channel.
  • Use separate bots and separate credentials for development, testing, and production, and configure permissions and alerts separately for each.

Establish a leak response process in advance

If you find the token in a public repository, logs, chat history, or on an unknown host, don't investigate the impact first and then decide whether to rotate. The correct order is to stop the damage immediately: use BotFather's /revoke, or go into /mybots and select the bot, to revoke the current credential and generate a new token.

Then write the new value into the production secret system, restart or perform a rolling update of the affected instances, and check each item: whether the new token can call the Bot API, whether the webhook is still receiving messages normally, whether background jobs have resumed, and whether the old token no longer works. Finally, clean up repository history, build caches, and copies of logs, review the call records during the exposure window, and record the cause of the incident and the remediation steps. A rotation is only truly complete once the old credential is invalid and every running instance has switched over.

3. Permission design: separate platform, system, and business permissions

In an enterprise integration, what a bot "can see" and what it "can do" should not be governed by the same set of switches. A safer approach is to split permissions into three layers: the platform layer controls the data and group management capabilities Telegram provides, the system layer constrains credentials and backend resources, and the business layer determines whether the current user is authorized to perform a specific operation. Loosening any one layer never substitutes for authorization checks in the others.

Platform layer: decide message visibility based on how the bot is triggered

If the bot only handles slash commands, messages that reply to it, or mentions containing its username, you should usually keep the group privacy setting on. That way the bot doesn't continuously receive ordinary group conversation, which keeps irrelevant data out of logs, queues, and models, and reduces the risk of false triggers and the spread of sensitive information.

Only when the business explicitly depends on ordinary group messages, such as detecting tickets, alerts, or compliance keywords in natural language, should you change Group Privacy through BotFather. After changing the setting, remove the bot from the target group first and then invite it back, so the old group membership state doesn't keep applying the previous configuration. Before launch, verify separately with plain text, commands, mentions, and replies; don't rely on the setting shown in BotFather to judge whether it has taken effect.

Group admin permissions should also be approved one at a time, rather than all granted at once for convenience. We recommend building a permission table based on actual actions:

Business actionPlatform permission that may be neededEngineering judgment
Automatically removing rule-violating contentDelete messagesLimit the applicable groups, rule scope, and message types, and keep an audit record of deletions
Temporarily restricting rule-violating membersRestrict membersSet clear trigger conditions and a recovery mechanism; misjudgments should be reversible by a human
Generating group invite linksInvite users or manage invite linksLinks should have a defined purpose, expiration, and scope so they don't become permanently open entry points
Notifications, lookups, process submissionsUsually no admin role neededRun as an ordinary group member by preference, so permissions don't expand for deployment convenience

System layer: isolate bot identities and production credentials

Development, testing, and production environments should use different bot identities and tokens. Splitting only the database while reusing the same token can still let test code read production messages, let a wrong webhook overwrite the production URL, or let development logs leak production credentials. Environment isolation should also cover callback URLs, message queues, caches, audit logs, and downstream API accounts.

The production token should be managed as a server-side secret, readable only by the deployment process and a small number of authorized ops staff. Don't write it into the code repository, image build arguments, front-end configuration, or searchable logs, and error messages should not output the full request URL or request headers. For rotation, you need an operational process covering updating the secret, redeploying, verifying callbacks, and revoking the old credential. Reads, changes, and rotations should all leave audit records.

Business layer: a Telegram identity is only an entry point for mapping

Usernames can be changed and may be empty, so a username must never be used to grant enterprise permissions directly. The reliable approach is to record the Telegram user_id and bind it to the enterprise account, job role, and organizational unit.

The binding flow should be initiated from the enterprise side; for example, after logging into an internal system, the user generates a short-lived binding code and sends it to the bot to complete the link. When the bot receives an operation request, it first verifies the Telegram identity mapping, then checks account status, role permissions, data scope, and whether the current group is allowed to perform the action. When employees leave, change departments, or group membership changes, authorization should be revoked in sync rather than persisting forever in the bot's database.

High-risk actions such as payments, bulk changes, exporting sensitive data, or deactivating accounts should not be executed on the strength of a single Telegram message. Depending on the risk, add a one-time verification code, secondary confirmation in the enterprise system, an approval workflow, or human review, and return to the user the object awaiting confirmation, the scope of impact, and the expiration time. The final audit record must at least link the requester, the conversation, the basis for authorization, the operation parameters, the approval result, and the execution status.

The acceptance criterion for permission design is not "the bot can complete the task," but being able to answer three questions: why it can see this message, which service can use this credential, and on what grounds the current user is performing this operation. Only when all three permission layers have clear boundaries will you avoid having to keep adding admin permissions just to keep things running as you connect more groups and business processes.

4. Taking webhooks to production: verification, idempotency, retries, and fast responses

In development you can start with polling: the process actively pulls updates, needs no public entry point, and is well suited to local breakpoint debugging and quick validation. In production, especially as message volume grows or multiple instances are needed, switch to webhooks. Telegram pushes updates to a public HTTPS URL provided by the enterprise, so the service no longer has to poll continuously and can more easily plug into gateways, queues, and centralized monitoring.

The public entry point must use HTTPS and one of the external ports allowed by the official Telegram Bot API documentation. The actual bot service doesn't have to listen on that port directly; a reverse proxy, API gateway, or load balancer can terminate TLS and forward traffic to the internal application. Before deploying, confirm DNS resolution, the certificate chain, gateway timeouts, request body size limits, and firewall policies. Don't assume that "it opens in a browser" means Telegram can deliver to it reliably.

The responsibilities of the webhook receiving layer should be kept as narrow as possible: verify the source, parse the update, register the idempotency key, write to the queue, and immediately return success. AI inference, CRM lookups, ticket creation, file processing, and bulk sending should never block the incoming request. Otherwise a single slow downstream query can lengthen the response time, trigger platform retries, and amplify the load on the service even further.

Processing stageSynchronousAsynchronous
Entry validationHTTPS, webhook secret, data structure, and required-field checksArchiving anomalous samples and security analysis
Event registrationGenerate the idempotency key and atomically write the receipt recordEnrich with user, group, and business context
Business processingOnly enqueue and returnModel calls, system lookups, rule evaluation, and message sending

Verification cannot rely solely on the URL being hard to guess. If the enterprise gateway supports access control, you can add network-level restrictions as well, but an IP allowlist should not be the sole basis, because upstream address policies may change.

Duplicate delivery must be treated as a normal case, not an edge-case anomaly. For example, syncing the same lead to different sales reps can be executed separately for each, but the same recipient must never have a duplicate record created or receive the same message repeatedly because of a retry.

The idempotency record should be written in the same transaction as the task enqueue wherever possible, or use transactional messages, the Outbox pattern, or similar approaches to avoid "the record succeeded but the task was never enqueued." Processing states should distinguish at least received, in progress, completed, retryable failure, and terminal failure. Retries need backoff and a cap, and should be decided by error type: network jitter can be retried, while parameter errors and permission denials should usually go straight to manual review.

  • Keep the original update, idempotency key, processing version, retry count, and related business object in the queue to support tracing.
  • Tasks that exceed automatic retry capacity go to a dead-letter queue; they should never be silently dropped.
  • Provide a manual compensation entry point that shows the failure reason, lets operators replay after correcting data, and runs the idempotency check again.
  • Record requests and results on the sending side as well, so you are never left unable to explain whether a follow-up message was actually sent even though receipt succeeded.

The criterion for production readiness is not that the webhook receives messages, but that the entry point can acknowledge quickly, duplicate events cause no duplicate side effects, downstream failures don't bring down the receiving layer, and every failed task has a path that can be queried, retried, and taken over by a human.

5. Session state: upgrading from keyword replies to recoverable business processes

Keyword replies only need to handle the current message, but multi-turn business flows must answer three questions: what the user is doing, which step they have reached, and whether the next message still belongs to that flow. Whenever ordering, ticket submission, identity verification, approvals, or information collection are involved, the conversation should be implemented as an explicit state machine, not as an ever-growing pile of conditionals in code.

Don't keep process state only in the bot process's memory. When the process restarts, a container migrates, the service scales horizontally, or a request lands on another instance, in-memory state is lost. What users see is usually not a clear error, but a bot that suddenly forgets the context, asks the same question again, or even writes an answer into the wrong field. Problems like this are hard to fix with retries.

A minimal session record should include the following fields:

FieldPurposeEngineering notes
Subject identifierIdentifies the user, group, or member within a groupIn group chats, don't rely on the chat ID alone; decide by business need whether to combine it with the user ID
Flow and current stepDetermines which business process is running and which next steps are allowedStep names should be stable and not depend on display text
Key parametersStores collected options, reference numbers, and temporary inputsStore sensitive information minimally and apply access controls
Update time and expiration timeIdentifies stalled sessions and cleans up temporary dataOn expiration, move to an explicit terminal state rather than deleting the entire history
Flow versionDistinguishes between old and new flow definitionsWhen releasing a new version, decide whether to continue the old flow, migrate state, or require a restart

Storage should be layered by the nature of the data. Data that affects business accountability, such as order results, ticket numbers, authorization decisions, and approval status, must be written to a persistent database. Session storage only guides the flow; records in the business system are the ultimate source of truth. When resuming a session, re-query the business state; never rely on the cache alone to decide whether an order has been created or a permission has taken effect.

State transitions should follow the order "validate the current state, then execute side effects, then commit the new state," and must handle concurrent messages. A user may send two messages in a row, or tap the same button several times. Each update should carry a session version number or use an atomic conditional update to prevent two requests from advancing the flow at the same time. Operations such as creating orders or submitting tickets also need a business idempotency key, so that consuming the same update twice or retrying a request doesn't produce two records.

Every flow should have exception exits designed in advance, rather than implementing only the ideal path:

  • Cancel: Terminate the current flow, release temporary resources, and state which actions have been completed and which have not.
  • Back: Allow returning only to defined nodes; once an irreversible business result has occurred, you cannot simply roll back the interface state.
  • Timeout: After a session expires, refuse to reuse old parameters, and prompt the user to confirm the current business status and re-enter the flow.
  • Start over: Create a new session instance while keeping the termination reason of the old one, to support auditing and troubleshooting.

Out-of-sequence input also needs deterministic behavior. When text or a button callback arrives that doesn't match the current step, don't guess the user's intent and force the flow forward. A safer approach is to return the operations currently available; if a change in the business record is detected, sync the factual state first, then decide whether to continue, end, or rebuild the session. Button callbacks can carry the flow instance ID, step identifier, and version information, but business parameters sent back by the client must not be trusted; the server must validate them again.

Human takeover should be a separate state, not just an ordinary tag. Once takeover begins, auto-replies and automatic flow progression must pause, though inbound messages can still be recorded. The history must distinguish bot messages, human replies, system actions, and business state changes, and store the takeover time, the handler, and the reason it ended. After the issue is resolved, automation should be explicitly resumed by a person or a controlled rule, and then, based on the current state of the business system, the flow should either continue or start over. Takeover must never be lifted automatically just because the next user message arrives.

Don't test only normal conversations during acceptance. At a minimum, cover service restarts, cache expiration, repeated button taps, out-of-order messages, concurrent operations by the same user, flow version upgrades, and recovery after human takeover. A multi-turn bot ready for production is not one that can remember a few sentences, but one that can explain its current state at any point of interruption and safely continue or end the business process.

6. A minimal architecture for automated operations: closing the loop across events, rules, tags, and outreach

An enterprise bot should not cram all its business logic into the webhook handler. That may get you live quickly, but as rules multiply, you run into hard-to-maintain code branches, duplicate messages, and user state you can't trace. A more robust minimal architecture separates message intake, event delivery, rule evaluation, state storage, and message sending, decoupling operational configuration from the underlying communication.

ModuleMain responsibilitiesEngineering notes
Telegram intake layerReceives updates such as commands, plain text, images, and button callbacksParses everything into internal events; does not execute complex business logic directly
Event queueBuffers inbound messages and business system eventsRetains event IDs; supports peak shaving, retries, and duplicate-consumption control
Rule or workflow engineEvaluates trigger conditions, selects audiences, and advances flowsRule versions must be traceable, and changes must be able to roll back
Tag and session storeStores user attributes, subscriptions, flow nodes, and recent interactionsTags should record their source, update time, and validity period to avoid endless accumulation
Business connectorsConnect to product, order, support ticket, membership, and other systemsLimit read scope; write operations require separate authorization and auditing
Sending service and operations consoleCreates send jobs, controls rate, configures templates, and reviews resultsDistinguish between generated, submitted, actually failed, and user unsubscribed

The focus of the inbound path is "understanding what the user did." After the intake layer receives an update, it should first normalize the event type and then hand it to the rule engine. For example, a command can start a subscription flow, text can go to intent detection, an image can be passed to recognition or human review, and a button callback confirms a selection. Tags, form results, and flow state produced by rule execution should be written to the storage layer rather than kept only in process memory.

The outbound path starts from internal enterprise events, not from scheduled mass messages. Events such as a product launch, an order entering delivery, or a change in service request status are converted by connectors into a unified format and enter the queue; the rule engine selects recipients accordingly, and the sending service then calls the Bot API to deliver the content to the appropriate users or groups. Business systems should only publish facts and should not assemble Telegram messages directly; otherwise templates, frequency controls, and unsubscribe policies end up scattered across multiple systems.

An operable rule can't just be "send when the condition is met." At a minimum it should include the following:

  • Trigger conditions: Specify whether it is started by user behavior, a business event, or a scheduled task, and define how events are deduplicated.
  • Target scope: Filter by subscription status, region, customer stage, or past behavior, while excluding people you are not authorized to reach.
  • Content templates: Manage missing variables, language versions, button links, and template versions, rather than assembling messages ad hoc at runtime.
  • Sending constraints: Set a cap on outreach per unit of time, message priority, and quiet hours, to prevent multiple flows from piling up on one another.
  • Exit mechanism: Provide a clear way to unsubscribe, and make unsubscribes immediately become an exclusion condition for subsequent rules.
  • Results tracking: Store the rule version, match reason, send result, button interactions, and downstream business conversions, forming a chain that can be audited.

Tags can only support decisions; they cannot replace business facts. For example, "has placed an order" should come from the order system, while "interested in new products" can come from subscriptions or interaction behavior. The two differ in data reliability, update frequency, and usage permissions. The operations console needs to show where each tag came from, and expired tags should become invalid automatically. Where sensitive attributes are involved, visibility and export capabilities should also be restricted.

Without a dedicated development team, you can use visual automation tools to validate auto-answers, conversational forms, and topic subscriptions first. The goal is to confirm whether anyone actually uses the flows, not to carry core business right away. Once you need to read customer data, modify orders, perform complex authorization, or have clear requirements for continuous availability, focus on examining where the platform stores data and how it handles access auditing, credential custody, backup and recovery, failure replay, and migration. A tool that can't answer these questions is fine for prototyping but not suitable as a long-term production hub.

Whether the loop is truly closed can be judged by a simple criterion: for any message, you can trace back the triggering event, the rule used, the basis for target selection, the template version, the send result, and the user's subsequent actions. If any one link is missing, the operational results are hard to explain and failures are hard to pinpoint.

7. Monitoring, alerting, and a launch checklist: making messages traceable and failures recoverable

The acceptance criterion for launching a bot should not be merely "it can reply to messages," but that for any update you can locate its processing path, determine the cause when sending fails, and restore the business through retries, compensation, or human takeover. We recommend generating a unified trace ID for every inbound update and carrying it through the webhook, the queue, the business systems, and the message-sending side.

What should logs and metrics record?

The chat_id can be stored as an irreversible hash, or with only its last few characters retained. Logs must never contain the token, full user profiles, the raw text of sensitive conversations, or downstream system credentials. When troubleshooting genuinely requires viewing messages, use controlled sampling, field masking, and time-limited access.

Monitoring should cover three layers: at the entry layer, watch webhook request volume, failure rate, and response latency; at the execution layer, watch queue depth, the wait time of the oldest task, and retry volume; at the business layer, watch send success rate, the share of 403 blocks, the number of 429 rate-limit responses, session completion rate, and human takeover rate. Looking only at interface availability won't reveal problems where "the system is healthy but the flow never completes."

How should alerts be tiered?

LevelTypical situationsHow to handle
Immediate alertConsecutive webhook failures, severe queue backlog, production token invalid or revokedNotify the on-call engineer, pause non-critical jobs, and prioritize restoring the entry point and sending capability
Business alertSend errors for a single template, falling session completion rate, spike in human takeover rateJoint investigation by operations and engineering, checking templates, rules, and downstream interfaces
Watch eventBrief latency increases, a small number of recoverable retriesRecord the trend; escalate once it exceeds a duration or cumulative threshold

Alert conditions should combine a ratio, an absolute count, and a duration, to avoid false alarms at low traffic and also to avoid missing cases at high traffic where a small percentage corresponds to a large number of failures.

What failure drills should be run before launch?

  • Simulate the 403 that occurs when a user blocks the bot, and confirm the system stops pointless retries and marks the conversation as unreachable.
  • Deliberately trigger the send rate limit, and verify that the system backs off as instructed by the server rather than immediately resending in a loop.
  • Following the constraints in Telegram's Bot API, test the 4096-character limit for a single text message; overlong content should be split before sending, with the order of the parts preserved.
  • Deliver the same update_id repeatedly, and confirm the idempotency key prevents duplicate charges, duplicate record creation, or duplicate outreach.
  • Restart the service while tasks are being processed, and check whether unacknowledged messages are re-queued and whether execution state can be restored from persistent records.
  • Simulate a business system timeout, and distinguish between "the request was not executed" and "it executed successfully but the response was lost"; in the latter case, query the result first before deciding whether to resend.
  • Rehearse a token rotation, and confirm the old credential is invalidated promptly, the new credential is injected through the secret management system, and no remnants remain in logs or configuration repositories.

Failed messages should be handled according to the nature of the error: network jitter, timeouts, and rate limiting can use exponential backoff with jitter; parameter errors, user blocks, and insufficient permissions should usually not be retried automatically; when the business state is uncertain, the message goes into a compensation queue, and a reconciliation job queries the final result. Every retry must have a cap on attempts, a deadline, and an idempotency key.

Should Telegram bot development use polling or webhooks?

Polling is fine for local development and short-term validation, since it is simple to deploy and easy to debug. Production environments usually choose webhooks to reduce polling overhead and shorten the time for events to arrive, but this requires public HTTPS, request verification, fast acknowledgment, asynchronous processing, and failure monitoring.

Why doesn't the bot receive ordinary messages in a group?

First check the bot's group privacy mode, admin permissions, and the message type. When privacy mode is enabled, the bot usually receives only commands, mentions, or messages related to it; after changing the setting, re-verify the bot's permissions in the target group. Don't substitute broader permissions for business judgment; request only the message scope the use case actually needs.

What should we do if the API token leaks?

Immediately revoke and regenerate the token in BotFather, update the production secret, then restart or roll out the affected instances, while clearing the old value from caches, build artifacts, and automation jobs. Then audit the webhook configuration, send records, and anomalous source addresses to determine whether any unauthorized operations occurred during the exposure. Merely deleting the token from the code repository is not enough to eliminate the risk.

Why do bot messages fail to send or get sent twice?

Send failures are commonly caused by users blocking the bot, permission changes, rate limiting, overlong content, invalid parameters, or network timeouts. Duplicate sends usually come from webhook redelivery, a consumer executing twice, or a call that succeeded but whose response never reached the client, which then sent the request again. The fix is to build idempotency records keyed on the business event ID, store the send status along with the message ID returned by Telegram, and have retry jobs check the local execution result first instead of sending again directly.