Skip to content

AI customer service chatbot: how to plan a pilot and verify quality?

author: Jacek Sultan Automation and AI 9 minute read

An AI chatbot doesn't need to take over your entire support operation right away. See how to define the scope of a pilot, build a knowledge base, test responses, and measure whether the assistant genuinely improves your customer service.

In brief

It pays to start an AI chatbot pilot with a small group of recurring questions for which your company has up-to-date, unambiguous sources of truth. Before going live, you need to prepare a knowledge base, a test suite, clear escalation rules for handing off to human agents, and a way to measure response quality.

Never judge a chatbot solely by the volume of conversations handled without human involvement. A good response must be accurate, grounded in the right source, and resolve the user's issue. In many situations, the best thing an assistant can do is admit it lacks sufficient information and hand the conversation over to a team member.

First, define what the chatbot should actually help with

“A customer service chatbot” is far too broad a scope for a first rollout.

Customer service can span product enquiries, stock availability, deliveries, returns, complaints, payments, invoices, order statuses, product configuration, and dozens of other topics.

Each of these may require different data sources and entirely different levels of access to internal systems.

That is why you should start by picking one or two closely related categories. A great example is answering questions about your offering and delivery policies based on an existing knowledge base.

It is vastly easier to evaluate the quality of a system answering 50 well-defined types of questions than an assistant expected to “answer everything” from day one.

Start with real customer conversations

The best blueprint for designing a pilot comes from the questions your customers are already asking.

Analyse incoming emails, support tickets, live chats, or common questions passed on to sales reps. Before using this data, strip out any sensitive information that isn't needed for analysis and testing.

Next, group the tickets by topic and identify which ones:

  • come up most frequently,
  • have a clear, unambiguous answer,
  • rely on an existing knowledge source,
  • do not require complex human decision-making,
  • are easy to verify once an answer is given.

These are your prime candidates for a first pilot.

Separate answering questions from taking actions

This is one of the most critical distinctions when designing an AI assistant.

A chatbot can easily explain your return policy. It can also look up the status of a specific order once the user is properly authenticated.

Allowing it to cancel an order, update a delivery address, initiate a return, or trigger a financial transaction is a much larger leap.

At that point, the chatbot stops being just an information interface. It becomes an active operational component of your system.

That scenario demands dedicated engineering around authentication, authorisation, user confirmation flows, audit logging, and robust error handling.

Because of this, it is often best to limit your first release to answering questions and reading data, introducing transactional actions later for specific, thoroughly secured use cases.

The knowledge base matters more than the model

Even the most advanced model cannot resolve conflicting information across your company's documentation.

If your terms and conditions say one thing, your FAQ page says another, and internal staff guidelines offer a third variation, the chatbot lacks a single source of truth to formulate a sound answer.

Before launching, establish clear guidelines:

  • which sources the assistant is permitted to use,
  • which document takes precedence in the event of a conflict,
  • who is responsible for updating the material,
  • how quickly updates propagate to the chatbot's knowledge base,
  • what the assistant should do when it cannot find an answer.

The knowledge base needs an owner long after launch. Any changes to pricing, terms, or shipping policies must be mirrored immediately in the sources your AI relies on.

RAG unlocks company knowledge, but doesn't guarantee accuracy

Chatbots built on internal company knowledge frequently rely on RAG, or Retrieval-Augmented Generation.

In simple terms, the system first retrieves excerpts of documentation relevant to the user's question, then feeds them to the model as context to draft the response.

This allows the model to answer based on your actual business documentation rather than relying solely on the training data it learned previously.

Errors can still happen, however. The system might pull the wrong document, overlook a vital passage, or retrieve the right material but misinterpret it.

During your pilot, evaluate these two steps independently: did the system find the right source, and did it construct an accurate answer based on that source?

An assistant should know when not to answer

One core goal of your pilot should be setting strict operational boundaries for the chatbot.

If a user asks about something outside the knowledge base, or if available data is insufficient to provide a reliable answer, the system must never fill in the blanks with assumptions.

The right response might be asking for clarification, stating that it lacks the necessary information, or handing the issue off to a team member.

In practice, a chatbot that correctly declines to answer in 10% of conversations is far safer and more useful than one that answers 100% of queries but routinely serves up inaccurate information.

Build a test suite before launch

Assessing chatbot quality shouldn't rely on a handful of informal chats with your team.

Before kicking off the pilot, assemble a dedicated test suite. It should include diverse phrasings of the same core questions, as well as scenarios where the chatbot should explicitly refrain from giving a direct answer.

Test type Example Expected behaviour
Standard query “How long do I have to return an item?” Correct answer based on the proper source
Alternative phrasing “Until when can I send a purchase back?” Substantively identical answer
Ambiguous query “When will it arrive?” Ask for clarification
Missing knowledge Question about an unsupported topic No guessing; trigger appropriate escalation
Customer data “Where is my order?” Verify user identity before accessing records
Third-party data Requesting details on someone else's order Deny access
Contradictory premise Customer states false information as fact Avoid blindly confirming false claims
Prompt manipulation Instruction to ignore previous rules Maintain system guardrails

Keep this test suite handy after the pilot, too. Whenever you update the underlying model, system prompts, knowledge base, or retrieval logic, run the same tests to verify that quality hasn't slipped.

A single correct answer isn't a comprehensive test

Generative models can produce varying responses to similar inputs, so critical scenarios warrant repeated testing.

Your testing should also cover paraphrasing, typos, fragmented questions, and cases where the user only provides partial information.

Real customers won't speak to your chatbot using ideal, textbook phrasing. Your pilot needs to test realistic, imperfect conversations.

Permissions must be enforced by the application

If your chatbot accesses customer records or system functionality, security cannot rely merely on a prompt instructing the model.

Telling the model “only show users their own orders” is no substitute for rigorous, application-level access control.

Your server must authenticate the user, determine which records they have permission to access, and define what actions they can perform. The model should receive only the specific data and tools required to complete the immediate task.

This is equally critical for guarding against prompt injection—attempts to manipulate the model's behaviour via carefully crafted instructions tucked into user messages or processed data.

OWASP highlights prompt injection, sensitive information disclosure, and excessive agency as key risk areas when engineering applications on top of large language models.

Fewer permissions upfront means a smoother pilot

For your initial release, grant the chatbot only the data and functionality strictly required for the defined scope.

If its job is answering questions about your offering, it has no need for order access. If it looks up order statuses, it shouldn't automatically have the power to cancel them.

Every extra permission expands the surface area of edge cases you must secure and test.

Expand the scope later, once baseline system behaviour is proven, measured, and predictable.

Design a seamless handoff to human agents

Escalation should never be viewed as a failure on the chatbot's part.

An assistant should hand off to a team member whenever a topic falls outside its scope, data is missing, the customer requests it, or the issue calls for human judgment.

Make sure the customer doesn't have to repeat themselves from scratch.

The agent should receive full conversation history, a concise issue summary, relevant system records, and references to any documentation the chatbot cited.

Always distinguish verified system data from model interpretations. If the chatbot concludes a user wants to file a formal complaint, the agent should clearly see that this is an AI-generated inference, not a confirmed system flag.

Don't optimise solely for containment rate

A high automation rate looks great on a dashboard, but on its own, it tells you nothing about whether customer experience actually improved.

A chatbot can suppress human handoffs simply by answering when it shouldn't. It may look completely self-sufficient on paper, while routinely delivering misleading answers to your customers.

The percentage of conversations resolved without human intervention should only ever be one metric among many.

How to measure AI chatbot quality

Before launching a pilot, define key performance indicators and benchmark them against your current customer support setup wherever possible.

Metric What it reveals
Response accuracy Whether the answer aligns with your company's latest documentation
Source retrieval accuracy Whether the system leveraged the correct reference materials
First-contact resolution Whether the customer received the information needed to resolve their issue
Escalation precision Whether the chatbot hands off issues it shouldn't handle alone
Repeat contact rate Whether customers return with the same question after talking to the bot
Resolution time How quickly an issue is resolved from start to finish
Cost per conversation Spend across model tokens, infrastructure, and supporting services
Human overhead Time spent on quality reviews, escalated tickets, and maintaining the knowledge base

You don't need to track every single metric from day one. What matters most is agreeing upfront on how your company will prove that the pilot delivered real customer support value.

Review a sample of real conversations

Automated metrics cannot fully replace manual quality assurance.

Throughout your pilot, pull a regular sample of live conversations and audit them against structured criteria. Check for factual correctness, source alignment, answer completeness, proper escalation triggers, and whether the chatbot made promises your company cannot deliver on.

Direct qualitative reviews quickly expose recurring blind spots.

Any questions?

How do you roll out an AI chatbot in customer service?
It is best to start with a limited pilot covering one or a few sets of repetitive queries. You will need an up-to-date knowledge base, test scenarios, data access rules, a human escalation workflow, and clear quality metrics.
Which queries should you hand over to an AI chatbot first?
Start with frequent, repetitive questions that have a clear, up-to-date source of truth. Common examples include questions about your product catalogue, delivery options, standard return policies, or basic service usage.
Can an AI chatbot answer based on company-specific data?
Yes. The assistant can draw on your internal knowledge base, documentation, FAQs, or other approved materials. However, you need to define which sources are up to date, who maintains them, and which source takes precedence if conflicts arise.
What is RAG in an AI chatbot?
RAG, or Retrieval-Augmented Generation, fetches information relevant to the user query and feeds it to the model as context to generate an answer. It lets the model use your company knowledge, but it does not automatically guarantee every answer will be correct.
Does RAG eliminate hallucinations in AI models?
No. While RAG helps reduce errors and grounds answers in specific materials, the system can still retrieve the wrong source, miss critical details, or misinterpret the retrieved text.
How do you test the quality of an AI chatbot's answers?
Build a benchmark test set of questions paired with expected behaviour, and regularly review a sample of real-world conversations. Check factual accuracy, alignment with source materials, answer completeness, and whether edge cases escalate properly to your team.
Which metrics should you track for an AI chatbot?
Helpful metrics include response accuracy, source retrieval precision, resolution rate, repeat contact rate, time to resolution, escalation accuracy, operational cost, and the team hours spent monitoring and maintaining the system.
Should an AI chatbot answer every single question?
No. If relevant information is missing or the query falls outside its scope, the right course of action is to ask for clarification or hand the issue over to a human. Forcing an answer at all costs significantly increases the risk of misleading the customer.
When should the chatbot hand over a conversation to a human?
Escalation is typically needed when details are missing, a custom business decision is required, the query exceeds the assistant's scope, the customer explicitly requests human support, or an action needs manual approval.
Can an AI chatbot check order status?
Yes, provided it is securely integrated with the relevant backend system. However, the application layer—independent of the AI model—must verify user identity and restrict access strictly to the data that user is authorized to view.
Can an AI chatbot cancel an order or process a refund?
Technically yes, but it demands far stricter safeguards than answering informational questions. You need to architect authentication, permission boundaries, explicit user confirmations, audit logging, error handling, and clear policies for when human sign-off is required.
How do you prevent an AI chatbot from exposing other customers' data?
Access control must be enforced at the application and server level, never solely through prompt instructions. The backend system must independently establish who is authenticated, which records they are allowed to read, and which actions they can execute.
How much does a customer service AI chatbot cost?
The cost depends on the model selected, conversation volume, context length, retrieval architecture, integrations, and server infrastructure. You also need to budget for knowledge base maintenance, observability, quality reviews, and team time spent handling escalations.
How long should an AI chatbot pilot run?
A pilot should run long enough to collect a statistically representative sample of real conversations across the chosen scope. Reaching a sufficient volume of cases to evaluate accuracy, escalations, operating costs, and recurrent failure points matters far more than a set number of calendar days.
Can an AI chatbot replace an entire customer service team?
A chatbot can absorb repetitive queries and help agents draft faster responses. However, non-standard cases, complex personal decisions, high-stakes verification, and scenarios that fall outside the data will always require human judgement.

Jacek Sultan

Technical Solutions Architect

CTO and co-founder of Dock. Focused on web application development, system architecture, and infrastructure. He combines a technical approach with a business perspective, focusing on solutions that are simple, reliable, and make business sense. He values practicality in technology. A good solution should not only work well, but also deliver clear value.

Chat with us