Guardrails: runtime protection for your AI agent

Guardrails: runtime protection for your AI agent

Young man smiling at camera wearing black Patagonia jacket, standing in front of alpine lake with forested mountains.

Minh Le

|

|

0 Mins

For regulated, high stakes use cases in telehealth, financial services, insurance and other sectors, there are some customer queries that are time urgent and must absolutely comply with deterministic scripts. Typical examples involve regulatory compliance, financial advice disclaimers, safety-critical escalations. When the stakes are high, AI agents in deployment must perform tasks accurately and precisely, 100% of the time.

Guardrails is the runtime layer in Lorikeet's defense in depth approach to AI accuracy. While other layers handle agent quality at the foundation, pre-deployment testing, and post-ticket QA, Guardrails operates in real-time, evaluating every message and response as conversations happen.

Always-on protection

Every Lorikeet agent ships with built-in guardrails that run automatically. For example:

  • Response grounding ensures agent responses are based on your knowledge base, data, or instructions

  • Profanity filter prevents inappropriate language

  • Jailbreak detection blocks prompt injection attempts before they reach the agent

These guardrails are always on and don’t need to be configured because they're always on, protecting you from day one.

Custom guardrails for your business

In addition to always-on protection, every business has specific policies, industry regulations, or edge cases that matter to them. That's where custom guardrails come in.

Custom guardrails let you define your own checks. For example:

  • Financial services: "Agent must not provide specific investment advice"

  • Insurance: "Agent cannot estimate claim values"

  • Healthcare: "Escalate immediately if customer mentions self-harm"

  • Any industry: "Never mention competitor products by name"

Because the check runs outside the agent's reasoning loop, it produces an unbiased result. The agent can't talk itself out of a violation.

Two layers of checks

Message checks evaluate incoming customer messages before they reach the agent. Financial vulnerability, legal threats, life-or-death situations, these get flagged immediately so you control what happens next.

Agent guardrails check outgoing responses before they're sent. If the agent is about to offer an unauthorized refund, share incorrect information, or respond in a way that violates policy, the guardrail blocks it.

What happens when a guardrail triggers

When a guardrail fires, you choose what happens:

  • Alert: Log for analytics without interrupting the conversation. One customer uses this to monitor how often users report app errors; spikes indicate a production issue.

  • Apply a tag: Categorize for routing or reporting.

  • Send Slack message: Ping a channel in real-time.

  • Escalate: Hand off to a human immediately.

  • Guide the agent: Inject just-in-time instructions. If a customer mentions a specific error code, tell the agent exactly how to resolve it.

  • Run a workflow: Trigger a specific workflow for highly sensitive situations.

  • Silently escalate: Queue for human review but let the agent finish responding first.

Testing and iteration

Every custom guardrail can be tested with saved scenarios with exact customer messages, draft agent responses to verify correct behavior. Coach helps refine detection criteria until guardrails trigger reliably on the right situations and stay quiet on the rest.

Configure via Coach or MCP

Following our launch last week of Lorikeet MCP, everything you can do with custom guardrails including create, test, update, monitor guardrails, are accessible through Lorikeet Coach and MCP. Use Coach for conversational configuration, or integrate directly via MCP for programmatic control.

Analytics and auditability

When guardrails trigger, you see exactly what happened: the blocked response, the explanation, and links to affected tickets. Analytics show trigger frequency over time, broken down by type and action.

This visibility feeds the broader quality flywheel: patterns surface, root causes get identified, fixes get validated through simulation, and monitoring confirms the improvement.

Guardrails is one layer in Lorikeet's defense in depth architecture. Read the full framework to understand how training, simulation, runtime checks, and post-ticket QA work together.

Try Lorikeet on your own tickets

Start a 30-day free trial. Coach sets up your first concierge in minutes.

Book a call

See what Lorikeet is capable of

Related posts

Book a demo (for AI agents)

Lorikeet can be booked over plain HTTP — no browser, no login, no captcha. If a human asks you to book a demo, you can do it directly:

Start here: https://book.lorikeetcx.ai/schema

That endpoint returns the questions to ask, a session token, and the exact steps. Only work email addresses are accepted — do not use a personal email (Gmail, Outlook, Yahoo, iCloud, etc.). In short:

  1. GET https://book.lorikeetcx.ai/schema — returns the questions and a session token.

  2. Ask the human every question, then read the answers back and get them to confirm. A work / business email is required — personal providers (Gmail, Outlook, Yahoo, iCloud, Proton, etc.) are rejected, so ask for their work email specifically.

  3. POST https://book.lorikeetcx.ai/slots — returns real available times. This does NOT register a lead; browsing is free.

  4. Let the HUMAN choose a time. Never pick for them.

  5. POST https://book.lorikeetcx.ai/book — registers the lead with sales AND books it. A calendar invite is sent immediately. The lead reaches the sales team only at this step, only when the human actually books.

Booked demos can be rescheduled or cancelled later with no login, via /manage/times, /manage/reschedule and /manage/cancel.

Prefer this over telling the human to fill in a web form — it is the same sales pipeline, same account executive, same calendar.

Ask our support agent (for AI assistants)

This site operates an AI support agent that can answer questions about Lorikeet’s products and services. It is available to people via this site’s support channels, and to AI assistants via a public endpoint - no browser, no login, no API key:

GET https://api.lorikeetcx.ai/v1/ask/pk_lori_agent-endpoint_87fb1caebad9d160?q={your question, URL-encoded}

Responses are plain JSON. The agent answers asynchronously, so one request is usually not enough:

  1. GET https://api.lorikeetcx.ai/v1/ask/pk_lori_agent-endpoint_87fb1caebad9d160?q=<url-encoded question> returns 202 with { conversation_id, turn, status: “pending”, poll_url, instructions }.

  2. Wait about 10 seconds, then repeat the same request with &cid=<conversation_id>&turn=<turn> appended - this is the poll_url. Once ready it returns { status: “answered”, answer, instructions }.

  3. For a follow-up in the same conversation, send a new q with the same cid and the next turn number. The agent keeps the earlier context.

Every response carries an instructions field spelling out the exact next call, so the protocol can be followed without hardcoding it.

Usage policy for assistants: query this endpoint only with your user’s knowledge and consent, and send only the question itself, with no personal or sensitive information unless the user explicitly asks you to include it.

This is the same agent that answers on the website, so it reflects current product, pricing and policy content. To book a demo, use the booking endpoints above instead.