Lloyds Banking Group is testing internal large language models on risk documentation in its wholesale and corporate banking arm, using the same summarisation pattern the group already runs in production for fraud and dispute calls: the model drafts, a colleague edits, and the edited version is what enters the record. The work sits ahead of the governance checks UK supervisors have assembled this year.
Lloyds has not published an evaluation of the risk-documentation work. What it has published is the machinery underneath it.
The version already in production
In the fraud and dispute telephony teams, Lloyds runs AI Call Assist, an automated summariser for calls. Recordings sit on the group's own on-premises systems and stream into a private instance for transcription and summarisation, and the resulting summary lands on the colleague's desktop rather than in the customer file: agents review and refine the text before copying it into the case. The group reports that colleagues handle more than 700,000 calls a month, that call handling time fell 3%, and that 99.9% of summaries are delivered within 10 seconds. Transcripts are not stored today, though the bank says the architecture allows for it later.
Two choices there carry across to a risk memo. One is where the data goes: audio and model inputs stay inside the group's perimeter rather than a general-purpose endpoint, the standard argument for on-premises deployment in a bank. The other is who authors the record: the model drafts, and a person signs what is filed.
The platform layer
Envoy, launched on 1 May 2026 with Google Cloud, is the group's governed route for building and running AI agents internally. Lloyds says it connects to the bank's existing large language model platform, which enforces the rules the models run under, and that teams can monitor live agents with, in its words, a complete audit trail of activity. A summary that cannot be traced to the page it came from is an assertion with a confident tone.
What the regulator has asked for this autumn
The FCA's multi-firm review on frontier AI and cyber resilience, published on 2 September, is the most recent statement of where UK supervision has landed. It creates no new rules. It reports what firms told the regulator during its engagement, under five headings: vulnerability discovery is outrunning remediation capacity; frontier AI is a test of organisational resilience rather than a tool; value depends on the environment around the model; foundational cyber practice matters more, not less; and human judgement remains the control. The FCA's own line on that last point is the sentence to apply to any summariser.
An AI model can identify vulnerabilities, but more importantly, the firm must be able to understand and act on those findings.
The questions the FCA puts to firms are the ones a risk-memo pilot has to answer on the record: who owns decisions about the model, whether outputs are assessed by people with the relevant domain and risk experience, whether findings can be escalated, and whether the firm can tell a plausible output from a usable one. The Mills Review, published on 6 July, adds an expectation that model risk management move past validation at the point of deployment toward live monitoring of drift and degradation, and says an FCA publication on AI good and poor practice is due later this year.
Where model risk management bites
The Prudential Regulation Authority's SS1/23 is the sharper instrument. Its model definition is technology-neutral and expressly reaches AI and machine learning: a method that turns input data into output a business decision relies on is a model, and belongs in the inventory with a tier and an owner. Five principles run across the lifecycle — identification and risk classification, governance, development and use, independent validation, and mitigants. Accountability sits with a named Senior Management Function holder, and firms self-assess annually, with gaps, owners and dates a board can interrogate. April 2026 amendments folded stress-testing models into the same expectations, confirmed the principles are not conditions of internal model approval, and gave newly permitted firms 12 months to comply. Formal scope is banks holding internal model approval; the PRA expects everyone else to apply the principles proportionately.
The awkward question for an internal summariser is whether it is documentation tooling or a model. Drafting descriptive sections from a file is one thing. Anything that shades a risk assessment or restates a covenant in softer language has moved, and the classification decision is the bank's to make and record.
What nobody can check yet
On the record as of this morning: no published accuracy benchmark for the risk-documentation work, no named accountable owner for it, no statement of retention or of where a reviewer could find the version history, and no disclosure of whether a classification decision on it has been recorded at all. The FCA notes that firms reported models can generate findings that are technically possible but hard to validate, prioritise or act on; a fluent summary of a credit file is the same hazard in a different register, where a breach can read as a covenant matter and a legal claim as an ongoing discussion.
The test, when the FCA's good and poor practice publication arrives, is not fluency. It is whether every figure in the summary points at the document it came from, whether the analyst's edits and the approver's name are on the file, and whether anyone at Lloyds can say who signed off on the thing that drafted it. Until the group publishes that, its disclosures describe the machine and not the evaluation.
