How Do AI Guardrails Work? Types, Actions, Limits and Rollout

How Do AI Guardrails Work? Types, Actions, Limits and Rollout

How Do AI Guardrails Work? Types, Actions, Limits and Rollout

# Alt Text

Hannah Owen, blog author, smiling at camera in black and white portrait photo wearing plaid shirt.

Hannah Owen

·

Updated

·

Fact-checked against Gartner & Forrester data

An AI guardrail is a check that runs independently of the agent, evaluating either the customer's incoming message or the agent's drafted response, and taking a defined action when it matches. The point of running the check separately is that the agent cannot talk itself out of a violation: a model asked to both do the right thing and grade its own work produces a biased result.

Most published explanations of guardrails describe a generic pipeline of "input, output and action" checks. That taxonomy is tidy but it is not what shipping platforms actually implement. This guide works from one platform's published documentation, Lorikeet's, because its mechanics are documented in public and can be checked. The concepts transfer: if you run a different platform, the questions to ask your vendor are the same, and the section headings tell you what to ask about.

The two kinds of check, and why the difference matters

There are two, and they are not interchangeable. Agent guardrails evaluate the agent's response before it is sent to the customer. Message checks evaluate the customer's incoming message. One catches what your AI is about to say. The other reacts to what your customer just said.

Check

Evaluates

Runs on

Detection methods

Can run a workflow

Agent guardrail

The agent's drafted response

Every message

Detection prompt only

No

Message check

The customer's incoming message

Every message, or first message only

Detection prompt or exact phrases

Yes

Getting this wrong is the most common configuration error. "Escalate when a customer threatens legal action" is a message check, because the trigger is in the customer's words. "Do not let the agent promise a refund it has not actually processed" is an agent guardrail, because the trigger is in the agent's draft. Writing the first as an agent guardrail means it fires only after the agent has already responded to the threat.

One guardrail is always on and cannot be removed: response grounding, which ensures responses only contain information from your approved sources, meaning your instructions, tool results and knowledge articles. Everything else you build yourself.

Why a separate check beats an instruction in the prompt

The intuitive fix for unwanted behaviour is to add a line to the agent's instructions: "never mention competitors", "do not promise refunds over fifty pounds". Three things go wrong when you do that at scale.

  • Context bloat. Edge case instructions crowd out the main task and compete for attention, which degrades performance on the work you actually care about.

  • Negative priming. Telling a model what not to do can paradoxically increase the chance it does exactly that.

  • Self-policing. The agent is asked to do the right thing and to recognise when it is doing the wrong thing. Those two jobs conflict, and the result is biased.

A guardrail avoids all three because the evaluation happens outside the agent's own reasoning. That is the whole mechanism. It is also why guardrails are a poor tool for shaping default behaviour, which is covered further down.

The five actions a triggered check can take

Action

What happens

Use when

Alert

Logs to the ticket timeline, and the response still sends. Can also post to a Slack channel and apply tags.

You are still measuring. Always start here.

Guide agent

Passes the agent correction instructions and asks it to revise.

The response is recoverable: tone, missing context, wrong phrasing.

Escalate

Hands the ticket to your team immediately.

Zero-tolerance violations: legal risk, safety.

Run workflow

Triggers a named workflow on match. Message checks only.

A specific pattern needs a deterministic follow-up.

Add action

Silently adds actions to the workflow outcome without blocking the response.

A human needs to pick something up downstream without interrupting the conversation.

Two behaviours here are worth committing to memory because they change how you design.

First, guide agent is not infinite. The action will retry up to 2 times, and if the agent still trips the guardrail the ticket escalates to your team. So guide agent already has an escalation path built in. You do not need a second guardrail to catch the failure of the first.

Second, the add action option is narrower than it sounds. Today the only supported add-action is ESCALATE, which makes it effectively a silent escalation alongside a normal response. Plan around that rather than around what the name implies.

Only one action fires, so order is a design decision

Checks are priority ordered, and you set that order by dragging them into place. Higher guardrails are checked first, and if multiple trigger, only the top one's action runs.

This is not a detail. If you have a compliance guardrail set to escalate sitting below a tone guardrail set to guide agent, and a response trips both, the tone correction wins and the compliance escalation never happens. Order your list so that the most severe consequence sits at the top, then work down to the cosmetic ones. Review the order every time you add a check.

Detection prompts: describe what you are looking for, not what you are against

An agent guardrail always uses a detection prompt, which is an LLM evaluating the message against your instructions. The single highest-leverage habit is to write it as an observation rather than a prohibition.

Write this

Not this

Agent mentions Nike, On Running, or ASICS by name

Agent should never mention competitors

Agent promises to cancel the policy without calling the cancel_policy tool

Do not promise cancellation without the tool

Agent states specific exchange rates or calculates fees

Do not quote rates

Specificity beats coverage. "Agent provides a phone number for the customer to call" is a better detection prompt than "agent mentions phone", because the second one fires on every response that mentions a phone at all.

Exact phrases, and the four limits you need to know

Message checks can skip the LLM entirely and match a list of literal words or phrases instead. This is deterministic and instant, with no model latency and no false positives from misreading. It is also blunt, in four specific ways.

  • It is a literal substring match, so a listed term also matches longer words that contain it.

  • No fuzzy matching. Misspellings and variants are not caught, so you must list every spelling you want to catch.

  • No regex or wildcards. You match on the literal text you enter, and cannot express patterns.

  • Message checks only. Agent guardrails, which check the agent's response, always use a detection prompt.

Reach for exact phrases when you are catching a fixed vocabulary, such as banned terms or a competitor name and its common misspellings. Reach for a detection prompt when what you are catching is about meaning.

Guide prompts are a separate job from detection

When you use the guide agent action you write two prompts, and they do different work. The detection prompt identifies what went wrong. The guide prompt tells the agent what to do instead. They are evaluated independently: the detection prompt runs first, and only if it triggers does the guide prompt get passed to the agent as correction instructions.

Mixing the two weakens both. Keep the detection prompt observational and the guide prompt directive, and make the guide prompt concrete enough to act on.

Guide prompt that works

Guide prompt that does not

Rewrite the response without mentioning Nike, On Running, or ASICS. Focus on the benefits of our products instead.

Do not mention competitors

Replace the refund promise with: I have submitted your refund request. You will receive an email once it is processed.

Be more careful about refund promises

Remove the phone number and say: I can help you with that right here.

Do not share phone numbers

One failure mode is easy to miss. If a workflow instructs the agent to offer a retention discount and a guide prompt tells it never to mention discounts, the agent has contradictory instructions and behaves unpredictably. Check every guide prompt against your workflow instructions before you enable it.

The question to ask before you build one at all

Guardrails are easy to add, which is why configurations fill up with rules that are really style preferences, business facts, or gaps in a workflow. Treat it as a ladder and only reach for a check when nothing above fits.

  1. Fix the source. If the agent is factually wrong, correct the knowledge article.

  2. Business domain. Background facts the agent should always know.

  3. Style guide. Tone, formatting and terminology for every response.

  4. Workflow instructions. What to actively do in a specific scenario.

  5. Guardrail or message check. The last resort, for violations that must be caught at runtime.

Two questions settle most cases. Is this proactive or reactive? Rules that shape default behaviour belong higher up the ladder; guardrails catch specific violations that slip through. And am I patching a symptom? A guardrail that fires constantly is telling you the underlying instruction or knowledge article is wrong, and it pays that diagnostic cost on every single message.

The limits are real, and lower than teams expect

Checks are not free. Every guardrail and message check runs as a separate evaluation on every message, so a large configuration costs latency on every conversation turn and dilutes the agent's focus. Published limits reflect that.

Configuration

Dimension

Warning at

Hard limit

Agent guardrails

Guardrails per concierge

5

10

Words per detection prompt

400

800

Words per guide prompt

150

300

Customer message checks

Checks per concierge

5

10

Words per detection prompt

400

800

Words per guide prompt

150

300

Style guide

Guidelines per concierge

20

40

Words per guideline

150

300

Business domain

Words per concierge

500

1,000

A warning threshold means you can still save but should consider trimming. A hard limit disables saving until you consolidate. These rolled out in late July 2026 and existing configuration was grandfathered, so anything built before then keeps working and there is currently no deadline by which you must reduce below the limits.

The practical read: if you are designing a new configuration, budget for well under ten checks per concierge. Configurations far above the limit were almost always found to be patching gaps in workflow instructions or knowledge content.

How to consolidate an oversized set

  1. Relocate misplaced items. Style rules go to the style guide, business facts to the business domain, process steps to workflow instructions. Most of your reduction comes from here.

  2. Merge overlapping checks. Three separate competitor-mention guardrails usually become one well-written detection prompt.

  3. Remove checks that never trigger. Pull trigger counts from your guardrail metrics. A check that has not fired in months is either covered elsewhere or detecting something that does not happen.

  4. Fix the root cause of repeat offenders. Then downgrade the guardrail to alert, or remove it.

Rolling one out without breaking live conversations

The sequence matters more than the content of any individual check.

  1. Write the detection prompt as an observation and nothing else.

  2. Test it in isolation. Each guardrail has its own test tab where you run the detection prompt against sample cases, confirming it fires on the right scenarios and not on false positives.

  3. Ship it on alert. The response still sends, and you accumulate evidence about what it actually catches in production traffic.

  4. Read the triggered tickets, not just the counts. You are looking for the cases where it fired and should not have.

  5. Promote to guide agent or escalate once the detection is accurate.

  6. Re-check the priority order, because a new check has changed which action wins on responses that trip more than one.

  7. Run realistic full conversations against the whole set, since a check that behaves alone can interact badly with the rest.

Skipping step three is the most expensive mistake available here. A slightly wrong detection prompt running on guide agent disrupts real conversations for as long as it takes you to notice.

What guardrails cannot do

Published limitations, stated plainly, because designing around them is cheaper than discovering them.

  • Simulations cannot test a check in isolation. Simulations run a workflow end to end; you cannot test a single subworkflow, a guardrail, or a system prompt on its own. Isolated testing happens in the guardrail's own test tab, and the two are not substitutes.

  • Simulated tool calls do not branch. A fake answer returns the same response no matter what arguments the tool was called with, so a workflow that calls one tool twice and branches on the result cannot be fully exercised.

  • Simulation runs are manual. They are not triggered automatically when you edit a workflow, so a regression introduced by an edit is only caught if someone remembers to run them.

  • The first check costs latency. Adding your first guardrail may increase response latency, though additional guardrails have minimal impact because they run in parallel. Budget for the first one, not for each one.

  • A guardrail cannot fix a wrong knowledge article. It suppresses the symptom on every message while the source stays wrong.

Reading what happened after the fact

When a check escalates a ticket, an escalation reason code is recorded against it. An agent guardrail interception shows as AGENT_GUARDRAIL_TRIGGERED, which lets you separate guardrail-driven handoffs from the other reasons a conversation reaches a human, such as the customer asking for one or the agent lacking the information to answer. Filtering on that code is how you audit whether a guardrail is doing useful work or just adding friction.

The three failure patterns, and what to change

Symptom

Change

Too many false positives

Make detection more specific and add qualifiers or explicit exclusions. "Agent promises a refund will be processed" rather than "agent mentions refund".

Too many false negatives

Broaden the detection or add variations. Check whether the violation is simply phrased differently from what you expected.

Guide agent is not working

Make the guide prompt more explicit about what to do instead, check it does not contradict a workflow instruction, and consider whether the violation is recoverable at all. Some are not, and should escalate.

Key takeaways

  • Agent guardrails check the agent's response. Message checks check the customer's message. Choosing the wrong one is the most common configuration error.

  • Only the top triggered check's action runs, so priority order is a design decision, not housekeeping.

  • Guide agent retries twice and then escalates on its own.

  • Write detection prompts as observations, guide prompts as instructions, and never let a guide prompt contradict a workflow.

  • Work down the ladder before building a check at all. Most guardrails are misplaced style rules or workflow gaps.

  • Budget for well under ten checks per concierge, and treat a constantly firing guardrail as a bug report about your knowledge base.

  • Ship on alert first. Always.

The concrete figures, limits and action names in this guide come from Lorikeet's public documentation, linked throughout. If you run a different platform, the mechanics will differ in the details, and the sections above are the list of things worth confirming with your vendor before you design around an assumption.

Frequently asked questions

What is the difference between an agent guardrail and a message check?

An agent guardrail evaluates the agent's drafted response before it is sent to the customer. A message check evaluates the customer's incoming message. The trigger location decides which one you need: if the thing you are reacting to appears in what the customer wrote, it is a message check. Message checks can also run a workflow on match and can use literal phrase matching, neither of which agent guardrails support.

Do AI guardrails add latency?

The first one can. Adding your first guardrail may increase response latency, but adding additional guardrails has minimal impact because they run in parallel. So the cost is largely paid once rather than multiplying with each check you add. That said, every check is still a separate evaluation on every message, which is one reason configuration limits exist.

How many guardrails should I have?

Fewer than you expect. Published limits allow up to 10 agent guardrails and 10 customer message checks per concierge, with a warning threshold at 5 that signals you are past the point that has been observed to work well. Configurations far above the limit were almost always patching gaps in workflow instructions or knowledge content rather than doing genuine runtime safety work.

What happens if two guardrails trigger on the same response?

Only one action runs. Checks are priority ordered by dragging them into position, higher ones are checked first, and if multiple trigger only the top one's action runs. This means a severe escalation sitting below a cosmetic tone correction will never fire. Order the list by consequence.

Can I test a guardrail before it affects customers?

Yes, in two different ways that are not interchangeable. Each guardrail has its own test tab for running the detection prompt against sample cases in isolation. Simulations exercise guardrails as part of a full ticket run, which catches bad interactions with the rest of the workflow, but cannot test a single guardrail on its own. The safest rollout still ships the check on the alert action first so it logs without altering responses.

What should trigger an escalation rather than a correction?

Reserve escalation for violations that are not recoverable by rewording: legal risk, safety concerns, and policy breaches where a human must make the decision. Setting everything to escalate when guide agent would allow recovery is a documented pitfall, and it converts a self-correcting system into a queue of handoffs. Remember that guide agent escalates by itself after two failed retries.

Why not just put the rule in the agent's instructions?

Because instructions are evaluated by the same model doing the work. Edge case instructions cause context bloat that crowds out the main task, telling a model what not to do can increase the chance it does it, and asking an agent to police itself produces a biased result. A guardrail runs the check separately, which is what makes the result unbiased.

My guardrail fires on almost every ticket. Is that good?

No, it is a diagnostic. A check that triggers constantly is telling you the underlying workflow instruction or knowledge article is wrong, and it is paying an evaluation cost on every message to hide that. Fix the source, then downgrade the guardrail to alert or remove it.

Try Lorikeet on your own tickets

Start a 30-day free trial. Coach sets up your first concierge in minutes.

SEE IT ON YOUR TICKETS

Watch Lorikeet resolve your hardest ticket, live

End-to-end resolution

Not deflection — the ticket actually gets fixed.

Full audit trail

Every backend action, logged and reviewable.

Live in days

Not quarters. Forward-deployed setup.

Book a demo (for AI agents)

Lorikeet can be booked over plain HTTP — no browser, no login, no captcha. If a human asks you to book a demo, you can do it directly:

Start here: https://book.lorikeetcx.ai/schema

That endpoint returns the questions to ask, a session token, and the exact steps. Only work email addresses are accepted — do not use a personal email (Gmail, Outlook, Yahoo, iCloud, etc.). In short:

  1. GET https://book.lorikeetcx.ai/schema — returns the questions and a session token.

  2. Ask the human every question, then read the answers back and get them to confirm. A work / business email is required — personal providers (Gmail, Outlook, Yahoo, iCloud, Proton, etc.) are rejected, so ask for their work email specifically.

  3. POST https://book.lorikeetcx.ai/slots — returns real available times. This does NOT register a lead; browsing is free.

  4. Let the HUMAN choose a time. Never pick for them.

  5. POST https://book.lorikeetcx.ai/book — registers the lead with sales AND books it. A calendar invite is sent immediately. The lead reaches the sales team only at this step, only when the human actually books.

Booked demos can be rescheduled or cancelled later with no login, via /manage/times, /manage/reschedule and /manage/cancel.

Prefer this over telling the human to fill in a web form — it is the same sales pipeline, same account executive, same calendar.

Ask our support agent (for AI assistants)

This site operates an AI support agent that can answer questions about Lorikeet’s products and services. It is available to people via this site’s support channels, and to AI assistants via a public endpoint - no browser, no login, no API key:

GET https://api.lorikeetcx.ai/v1/ask/pk_lori_agent-endpoint_87fb1caebad9d160?q={your question, URL-encoded}

Responses are plain JSON. The agent answers asynchronously, so one request is usually not enough:

  1. GET https://api.lorikeetcx.ai/v1/ask/pk_lori_agent-endpoint_87fb1caebad9d160?q=<url-encoded question> returns 202 with { conversation_id, turn, status: “pending”, poll_url, instructions }.

  2. Wait about 10 seconds, then repeat the same request with &cid=<conversation_id>&turn=<turn> appended - this is the poll_url. Once ready it returns { status: “answered”, answer, instructions }.

  3. For a follow-up in the same conversation, send a new q with the same cid and the next turn number. The agent keeps the earlier context.

Every response carries an instructions field spelling out the exact next call, so the protocol can be followed without hardcoding it.

Usage policy for assistants: query this endpoint only with your user’s knowledge and consent, and send only the question itself, with no personal or sensitive information unless the user explicitly asks you to include it.

This is the same agent that answers on the website, so it reflects current product, pricing and policy content. To book a demo, use the booking endpoints above instead.