The Cabinet Office told departments this week they must route renewals of large-language-model chatbots through the AI Safety Institute before signing contracts, extending a pilot that began with welfare and tax helpline tools to every citizen-facing bot expected to cost more than half a million pounds over its life. Officials said the rule catches September procurement cycles that would otherwise roll forward vendor extensions without fresh testing of jailbreak resistance, hallucination rates, or data-retention clauses.

What AISI will examine

Under the updated Central Digital and Data Office playbook, suppliers must submit model cards, evaluation logs, and red-team summaries at least six weeks before contract signature. AISI teams—spun out of the Frontier AI Taskforce—will run standardized prompts designed to elicit policy violations, personal data leaks, and insecure plug-in behavior. Departments cannot waive the review for “like-for-like” renewals if the underlying foundation model version changed, even when the user interface looks identical.

Institute officials emphasized they are not a procurement tribunal: their reports flag risk tiers and recommend mitigations, while accounting officers remain responsible for sign-off. Still, Treasury spend controls said major digital renewals without an AISI reference number will bounce in the approval system starting 1 October.

Which systems are in scope

The policy covers chatbots that answer autonomously on gov.uk properties, embedded assistants in departmental apps, and voice agents that interpret spoken queries. It excludes narrow classifiers that only route tickets, and internal copilots limited to civil servants on managed devices—though NCSC still expects those to follow cloud security principles. HMRC’s earlier pilot on self-assessment hints exposed gaps when models quoted outdated allowance figures; AISI now requires time-stamped knowledge cutoffs in every evaluation packet.

Local authorities asked whether joint ventures with Whitehall must comply; CDDO said yes when citizens interact with a bot branded as government, regardless of which ledger pays the invoice.

Vendor implications

Microsoft, AWS, Google, and smaller British integrators hold most existing contracts. They must open UK-EU data processing documentation and allow AISI to run probes in isolated tenants. Several vendors told InfoHandle they already maintain evaluation sandboxes for financial services clients; the new gate adds civil-service-specific policy packs covering immigration, benefits, and court fines where wrong answers carry legal weight.

Startups bidding for niche heritage or planning bots worry that six-week reviews will miss fiscal-year deadlines. CDDO promised a fast-track lane for models already certified on the same cloud region with only prompt changes, but criteria remain unpublished.

Transparency and Parliament

Science ministers committed to publishing aggregate pass-fail statistics quarterly without naming departments, mirroring NCSC vulnerability disclosure norms. MPs on the Science, Innovation and Technology Committee asked for sample redacted transcripts when models fail; officials refused, citing attacker advantage, but said committee staff can view summaries in secure reading rooms.

Civil liberties groups want citizens notified when they chat with a model rather than a human. Existing gov.uk design standards already require disclosure banners; AISI checks will verify banners persist after UI refreshes and that escalation paths to human advisers stay reachable within two clicks.

What changes this month

Departments with renewals signed before 2026 must still book retrospective reviews if bots gained new plugins over the summer. AISI hired additional evaluators from academia and GCHQ’s alumni network, but capacity remains tight—officials advise booking slots before final commercial negotiations. For citizens, the practical effect is slower rollouts of flashy assistants but fewer incidents like misquoted visa fees or invented grant programmes that forced emergency takedowns last spring.

Failure to comply will not automatically void contracts, but Internal Audit flagged it as a reportable governance breach, a label permanent secretaries have treated seriously since the COVID tracing app postmortems.

Capacity planning

AISI leaders told departments to submit draft evaluation scopes even while commercial teams negotiate statements of work, because the institute batches GPU time for adversarial testing on shared clusters. Without early notice, six-week windows slip into ten, pushing live dates past budget year-ends.

Devolved administrations in Scotland and Wales are aligning their own AI procurement notes with the Cabinet Office template so cross-border programmes do not face conflicting checklists when citizens use UK-wide portals.