The determining question in financial services is not what the system can do. It is whether you can reconstruct, months afterwards, why it did it — with evidence a regulator accepts.
The constraint that shapes deployment
Model risk management frameworks already exist in most regulated firms, and AI systems fall under them. That means documented validation, ongoing monitoring, and demonstrable governance. Firms treating AI as outside that perimeter discover otherwise at examination, which is an expensive moment to find out.
Where agents work well
- Document processing — extracting from statements, applications and disclosures arriving in inconsistent formats.
- KYC and onboarding — assembling and cross-checking, with decisions escalated rather than made.
- Exception triage — surfacing anomalies for human review rather than resolving them.
- Internal knowledge — answering staff questions from your own policy documents.
- Reporting assembly — collecting and reconciling ahead of human sign-off.
Where autonomy stops
Anything constituting advice, a suitability determination, a credit decision or a compliance judgement requires human accountability. These are not areas where careful prompting suffices — the constraint must be structural, so the system is incapable of issuing the determination rather than instructed not to.
Why generic AI advice fails here
Most AI implementation guidance assumes you can iterate in production and fix issues as they surface. In a regulated firm, an issue that reaches a client is a reportable event. That inverts the sequence: validation before deployment, monitoring from day one, and a conservative scope that expands on evidence.
What this means for your firm
Build auditability into the architecture rather than adding logging later, keep suitability and advice human, and treat AI systems as within your existing model risk perimeter. Our governance guide covers the review and audit structure.