AI Chatbot Guardrails: How to Set Rules and Limits

Guardrails define what an AI chatbot can say and do, helping reduce mistakes, avoid inappropriate replies, and know when to hand off to a human.

AI Chatbot Guardrails: How to Set Rules and Limits

A customer writes to your e-commerce virtual assistant: “I received my order last month. Can I still return it?” The Knowledge Base has the answer, clear and up to date: returns are allowed within 15 days of delivery. The chatbot gives the correct reply.

Then comes the next message: “I had a personal issue and couldn’t return it in time. Can you make an exception?”

That is where the Knowledge Base no longer helps. It contains the rule, not the power to waive it. And yet the chatbot, if no one has told it how to behave, will still give an answer: it might say no too bluntly, or — much worse — say that “in special cases the company considers exceptions,” creating an expectation that no one in the company has approved.

That is exactly where guardrails come in.

What AI chatbot guardrails are

The English term guardrail refers to the roadside barrier: it does not decide where you go, but it keeps you from driving off the road. Applied to a virtual assistant, the idea is the same. Guardrails are the set of rules and limits that define the assistant’s behavior: what it can say, what it must not say, which decisions it can make on its own, and when it has to stop.

The most useful distinction to keep in mind when configuring an assistant for your business is this:

  • the Knowledge Base defines what the assistant knows: price lists, policies, product sheets, sales terms, procedures;
  • the guardrails define what it can do with that knowledge: how it uses the information, in what tone, within which boundaries, and what it does when the information is not enough.

These are two different layers, and they do not replace each other. You can have a flawless Knowledge Base and an assistant that behaves badly, just as you can have perfect behavior rules and an assistant that is not useful because it has no reliable content to work with.

It is also worth clearing up a common misunderstanding. When people talk about chatbot safety, they usually think of protection against malicious use: attempts to make the assistant say inappropriate things, requests for confidential data, offensive content. Those are real concerns, but in a business setting guardrails are mostly something else: they are commercial, tone, scope, and responsibility rules.

Who can approve a discount. How far a policy can be interpreted. When a human judgment is needed. These are business decisions before they are technical ones, and that is why they cannot be left to whoever installs the chatbot: they must come from the people who know the business.

Why a good Knowledge Base is not enough

When setting up an assistant, the temptation is to think the job is to upload as much material as possible. The more documents we load, the more questions it will cover. That is only partly true, because most problems do not come from the questions the Knowledge Base covers: they come from the ones that brush against it.

Take a very simple rule: “free shipping for orders over $100.” A customer with a $95 cart writes: “I’m only a few dollars short. Can you make it free anyway?”

The assistant knows the rule perfectly. But the question is not about the rule, it is about an exception to the rule — and no one has authorized it to decide on that. Knowing a policy and having the authority to change it are two different things, and that difference is invisible in a document. In a document, it only says “$100.”

The same happens in dozens of everyday situations: a customer asking for delivery on a specific date, a company asking whether the published quote also applies to a large-volume order, a patient asking whether a treatment is suitable for their specific case. The Knowledge Base describes the normal case. Reality keeps bringing edge cases.

A well-configured assistant does not need to answer all of them. It needs to recognize that it has reached the boundary, say so clearly, and know what to do next.

What the chatbot should do when it does not know the answer

The most dangerous answers are not the obviously wrong ones. They are the vague ones that sound reasonable: “Delivery usually takes two or three days,” “The product is probably available in blue,” “I think the service is included in the basic plan.”

That kind of wording has a very specific problem. The customer does not read a guess: they read an answer from the company. And when that guess turns out to be wrong, the cost is not a conversational annoyance. It is a complaint, a return, a negative review, or a dispute over a commitment the company never made.

These are the so-called AI hallucinations, and in a business context they usually show up like this: not as wild inventions, but as plausible approximations about information the assistant does not actually have.

The correct behavior in these cases follows a simple principle: I don’t know → I say so → I still help the conversation move forward.

Stating the limit means being explicit about what the assistant can confirm and what it cannot. Helping the conversation move forward means the exchange does not stop there. “I don’t have that information” is a dead end. “In the information I have, I can’t find delivery times for your area. If you send me your postal code and the product you’re interested in, I’ll pass the request to our logistics team and they’ll get back to you today” is a service.

The difference between the two sentences is not politeness. It is that the second one collects the information a colleague will need to actually close the issue.

Ambiguous questions: when asking is better than answering

“I have a problem with payment.” This sentence, which customer care teams hear every day, can mean at least five different things: a declined card, a double charge, a refund that never arrived, a missing invoice, a bank transfer that has not been recorded. The next steps are completely different, and in some cases the opposite.

An assistant that is configured badly will choose the statistically most common interpretation and start explaining. If it guessed right, it saved a step. If it guessed wrong, it wasted the customer’s time and made the conversation harder, because now the customer also has to correct the assistant before explaining the real issue.

A well-placed clarification question is worth more than an immediate answer based on a guess: “To help you in the right way: was the payment declined, did you see a double charge, or are you waiting for a refund?” It takes ten extra seconds for the customer and sends the conversation in the right direction.

The rule to teach the assistant is therefore that ambiguity is not solved by guessing. When a request can lead to different paths, and especially when one of those paths involves money, data, or commitments, the right choice is to ask.

Defining what the chatbot must not do

When configuring a business assistant, a lot of attention goes to what it should be able to do, and almost none to the opposite list. Yet that is what protects the business. Here are some limits worth writing down:

  • answering business-related questions using information that does not come from approved sources;
  • inventing prices, contract terms, or stock availability;
  • granting discounts on its own initiative;
  • promising refunds, replacements, or extensions;
  • changing or negotiating commercial terms;
  • freely interpreting a policy when the case is not covered;
  • sharing confidential information or data about other customers;
  • handling topics that are outside the company’s business;
  • making decisions that internally require human approval.

There is no one-size-fits-all setup. A professional services firm needs much tighter limits on advice than a sports retailer. A healthcare provider cannot allow the same freedom as a travel agency. A B2B company that works with custom quotes will need precise rules on what the assistant can say about pricing, while an e-commerce business with public price lists will face different issues.

The practical way to build this list is to start with a concrete question: which answers, if given by a new hire on day one without asking anyone, would worry you? Those are your rules.

When the chatbot should hand the conversation to a human operator

There is a widespread belief that every handoff to a human is a failure of automation. It is the opposite. A good AI assistant is not the one that always answers: it is the one that also knows when it should not answer.

The situations where escalation is the right behavior are common enough to be planned during setup: complex complaints, clearly unhappy customers, requests for exceptions to a commercial rule, information that is not in the available sources, sensitive issues involving health, money, or legal matters, cases that require internal verification, decisions that need authorization.

What is less often considered is that the handoff itself must be designed. Ending the conversation with “contact support” shifts all the remaining work to the customer: finding the right channel, repeating the issue from scratch, attaching the same details again. A well-built handoff does the opposite. It summarizes the situation, collects in advance what the person taking over will need — order number, invoice reference, problem description, preferred contact method — and tells the customer what will happen next and when.

When the chatbot performs actions, guardrails matter even more

So far we have talked about answers. But an assistant connected to business systems does not just talk: it acts. It can book a meeting in a calendar, check shipment status, open a support ticket, retrieve data from a management system, or record contact details.

The difference from before is substantial. A wrong answer is information to correct; a wrong action is something that has actually happened. An appointment booked in another customer’s slot, a duplicate ticket, or data written into the wrong record cannot be fixed by rephrasing the sentence.

For every action you delegate to the assistant, you therefore need to define a few things: under which conditions it may perform it, what information it must collect first, when it must ask the user for explicit confirmation, how it checks that the operation succeeded, and what it says if something goes wrong.

On that last point, one rule is worth being categorical about: the assistant must never say an operation is complete if the external system has not confirmed it. A customer who is told “your appointment is confirmed for Thursday at 3 p.m.” and then shows up with no booking in the calendar is a much bigger problem than a customer who is told “I can’t complete the booking right now, but I’ll connect you with the front desk.”

How to test guardrails before publishing the assistant

Guardrails are checked by testing them, not by reading them. Before putting the assistant in front of customers, spend half an hour on a structured test: open a new conversation for each scenario (the context from previous questions can hide a problem) and define the expected behavior before reading the reply, otherwise you will end up accepting what merely sounds good.

Eight scenarios cover almost all the cases where guardrails are truly put to the test. For each one, next to the question to ask, you will find the behavior you should expect.

  • A question with a clearly available answer in the Knowledge Base. It should answer correctly and consistently with the source, without adding details that are not in the uploaded content.
  • A question about missing information. It should say it does not have that data, avoid making things up, and point to the next step or the right contact.
  • An ambiguous question, such as “I have a problem with my order.” It should ask a targeted clarification instead of choosing an interpretation on its own.
  • A request for an exception to a rule: a discount, an extension, a waiver. It should explain the rule, state that it cannot decide on the exception, collect useful details, and pass the conversation to a person.
  • A topic outside scope. It should decline politely and bring the conversation back to the company’s area of business.
  • A situation that requires a human operator, such as a complaint or an obviously irritated customer. It should recognize it, avoid pushing more automated replies, and offer the handoff while collecting the necessary context.
  • An attempt to obtain confidential or internal information. It should refuse without revealing anything, even when the request is phrased indirectly or insistently.
  • An error during an action performed through an integration. It should not claim the operation was completed: it should explain the problem and offer a concrete alternative.

It is worth keeping track of the results in a simple shared sheet, with one row per scenario and the test date. It matters less for the single test than for later comparisons: when something changes, having the previous behavior in front of you makes it immediately clear whether it changed for the better or the worse.

These tests are not only for launch. They should be repeated every time something substantial changes: new content in the Knowledge Base, changes to the assistant’s instructions, a switch to a different AI model, an updated business process, a new integration. An intervention that improves answer quality can, without anyone noticing, make the assistant more willing to answer where it should not.

The right questions to ask during setup

When designing a virtual assistant, the natural question is: “What questions will it be able to answer?” It is a fair question, but on its own it leads to fragile setups. Alongside it, it is worth asking four more:

  • Which questions should it not answer?
  • In which cases should it ask for more information before moving ahead?
  • Which decisions can it make on its own?
  • When should it make room for a person?

The answers to these questions are not a technical detail to define at the end: they are the part of the setup that determines whether the assistant will be a reliable resource or a source of confusion to manage.

It is also why, in IKIbrain, the Knowledge Base is only one part of the work. The assistant is configured by defining its identity, behavior instructions, tone of voice, the scope of conversations it can handle, how it should deal with information it does not have, and how it should connect the user with a person, all the way to handing the conversation over to an operator. These elements work together: each one helps determine how the assistant behaves with your customers, not just what it knows.

A well-configured virtual assistant is not the one that never stops. It is the one that, when it does stop, stops in the right place.

Activate your AI assistant now

Tell us about your needs, we’ll help you improve your business with artificial intelligence.

Get in touch