A customer calls their bank at 9pm on a Friday because they have lost their job and cannot make their next mortgage repayment. The agent on the other end needs to verify identity, check hardship eligibility, explain options without giving financial advice, offer a payment arrangement within policy limits, document the entire interaction for regulatory review, and do it all with empathy.
An AI agent can do all of that, and whether it should is the question every compliance team in banking is stuck on.
The legitimate fear
Banking regulators do not grade on a curve. A single unauthorized disclosure, one piece of inadvertent financial advice, or a missing audit record can each trigger enforcement action.
Regulators are already watching this specifically, and APRA, the FCA, and the OCC have each published guidance on the use of AI in regulated firms.
That is why compliance remains the single largest blocker to AI adoption in financial services, ahead of both cost and technology.
The concern is well-founded. Most AI systems were built for e-commerce returns and subscription cancellations, and regulated banking asks them to work where every word carries regulatory weight, which they were never designed to do.
Where AI goes wrong
The risks are specific and well-documented.
Hallucination is the most discussed risk and the most misunderstood one. The damage usually comes from small wrong details, like stating a fee is $25 when the fee is $30, or telling a customer they qualify for a hardship arrangement when they do not. In a regulated conversation, a small factual error is a compliance violation.
Unauthorized actions are a different class of risk entirely. An AI system that can process a payment arrangement has to be stopped from processing one that exceeds policy limits, and an AI that can look up account balances has to be kept away from records it has no business touching. Without explicit action boundaries, every capability becomes a liability.
Missing audit trails make both problems worse. If an AI hallucinates or takes an unauthorized action and there is no complete record of what happened, the institution cannot remediate, report accurately to regulators, or demonstrate that controls were in place.
Data handling failures round out the risk profile. Customer data flowing through AI systems must be governed with the same rigor as data handled by human agents, which means encryption, access controls, retention policies, and geographic constraints on processing.
How guardrails actually work
The word "guardrails" has become meaningless through overuse and every AI vendor claims to have them, so the differences sit in the architecture behind the claim, which comes down to three layers.
Policy grounding means the AI answers from a defined set of approved policies, product documents, and response frameworks that the institution controls, rather than from its general training data. When a customer asks about early termination fees, the AI reads the institution's current fee schedule, and because the restriction is built into how the system retrieves an answer, it holds for every conversation.
Action constraints operate inside the tools, not inside the model. Lorikeet calls this Pockets of Determinism: agentic orchestration wrapping deterministic operations.
The concierge can call any tool at any time, which is what keeps the conversation natural. It cannot make a tool execute outside policy, because the tool checks its own preconditions first. If the institution allows payment arrangements up to 90 days, the concierge can still attempt 91, and the tool returns conditions not met.
Think of a banking app. A child can tap "Close Account." The button is always there, but the app checks for remaining balances before executing. The concierge's tools behave the same way.
Real-time escalation triggers are the third layer, because some topics, customer states, and conversation patterns have to route to a human immediately.
A well-designed system does not wait for the AI to decide it is out of its depth. It monitors the interaction continuously and transfers the moment predefined conditions are met. Those checks run on a separate thread, outside the concierge's reasoning loop, so the concierge cannot talk itself out of a violation.
The obvious objection is that if code holds every line that matters, the AI is decoration, but deterministic software cannot read a customer typing that they have just lost their job, recognize that this is a hardship conversation and not a complaint, and hold that context through identity verification and into the options conversation.
The design rule follows from that. AI handles the reading, the reasoning, and the wording, while code holds the three places where being wrong is a regulatory event: what the answer is sourced from, whether an action is allowed to execute, and when a human takes over.
Anatomy of a safe interaction
Take the hardship scenario from the opening and walk through it step by step.
The customer states they cannot make their next repayment. The AI acknowledges the situation with empathy, then runs identity verification on the institution's standard protocol, with no shortcut for a distressed customer.
Identity confirmed, the AI accesses the customer's account and checks hardship eligibility against the institution's current hardship policy. It applies the policy as written rather than interpreting it.
The customer qualifies, so the AI presents the available options exactly as defined in the institution's hardship framework, and it does not recommend one over another, because doing so could constitute financial advice.
It explains each option, confirms the customer's preference, and processes a payment arrangement within its authorized limits.
Throughout, every interaction, model choice, and action is tracked and fully visible through transparent reasoning tooling. If a regulator asks why this customer received a 60-day payment pause instead of a 30-day pause, the institution can trace the outcome back to the workflow rule that produced it, and hand over the record as an audit trail export.
The customer receives a confirmation with all relevant details, and the case is flagged for human review as required by the institution's hardship procedures.
The arrangement is not the end of the relationship. The same concierge checks in before the first reduced payment falls due, picks the conversation up on whichever channel the customer uses next, and carries the arrangement with it as context. Each of those touches is logged the same way as the first conversation.
Built for regulated environments
Most AI platforms retrofit compliance onto systems designed for unregulated use cases, which produces controls that look strong in a demo and fail under regulatory scrutiny.
Lorikeet was built for regulated industries from the ground up. It is SOC 2 Type II audited, ISO 27001:2022 certified and GDPR attested, with certifications published on our public trust center.
Every interaction produces a complete audit trail, and for regulatory reporting we support audit trail exports, compliance dashboards, and exception reporting. The institution explicitly configures and constrains every action the AI can take, and the compliance team defines the escalation triggers.
Carmoola runs on that architecture. It operates in consumer credit regulated by the UK's Financial Conduct Authority, where an answer about affordability or a missed repayment is itself a decision.
Its concierge resolves 60% of inbound conversations and 90% of outbound conversations end to end. "It's like having your best agent on their best day, available all the time," says Lucinda Bentley, Carmoola's Head of Customer Operations.
Book a call
See what Lorikeet is capable of
Share this article






