IntegrityAI

Insights

FCA policy

What the Mills Review means for compliance file checking

The FCA’s review of AI in financial services proposes no new AI-specific rulebook. That is not a relaxation. It puts the weight back on governance the regulator already expects firms to have, and governance is where an AI file check either holds up or falls over.

The FCA published the Mills Review, led by its executive director Sheldon Mills, on 6 July 2026. Anyone who opened it hoping for a definitive list of what an AI system may and may not do inside a regulated firm will have closed it slightly unsatisfied. The review concludes that no new AI-specific rulebook is needed. Its reasoning is that the Consumer Duty and the Senior Managers and Certification Regime already supply the right anchors, provided firms build accountability, auditability and human oversight in from the point of design.

Read that last clause again. A review that declined to write new rules has told firms in the same breath that the existing rules bite at the design stage. They do not wait for deployment. You cannot buy an AI tool, run it for six months and then retrofit an audit trail onto it. Either the accountability was designed in or it was not, and the file will show which.

No new rulebook is the harder answer

Some readers will treat “no new rules” as a green light. We read it the other way round. A dedicated AI rulebook would have handed firms a checklist, and checklists are comfortable: satisfy the list, evidence the list, move on. By pointing firms at the Consumer Duty and SM&CR, the review swaps that for a harder question. Can a named individual explain and defend what this tool did to a specific client file? Volume does not make that question any easier to answer. Neither anchor is new, either. Both applied to the firm before it bought anything.

The review frames the problem with an autonomy spectrum. It runs from AI that assists a person, through genuine human-machine collaboration, to full delegation of a task to the system. As the review puts it, expectations tighten as you move along that spectrum. Where a firm sits on it is a design decision, and plenty of firms have made that decision by default.

Compliance
file review

Assist Collaborate Delegate

More autonomy, tighter governance expectations →

Schematic. The three stops are the review’s; the wedge is a shape, not a measurement, and placing compliance file review at the assistive end is our reading.

Compliance file checking is not a candidate for the delegation end of that spectrum, and should not be built as though it were.

We put file review firmly at the assistive end. A file check produces a judgement about whether advice given to a real person was suitable and well evidenced. It feeds supervision, T&C records, remediation decisions and, on occasion, redress. Nothing in that chain should be recorded because a model said so.

The four governance expectations, in practice

Strip the governance framing back and firms are being asked for four things: a named accountable person, an evidence trail back to source documents, human sign-off before an outcome is recorded, and ongoing testing for drift. Frank, the reviewer inside our platform, reads the file, grades it and shows the reader what he read. He was not built against the Mills Review; the review did not exist when the platform was designed. Those four still describe how the product works, so here is what each one looks like in a live system.

A named accountable person

Every review in Frank belongs to a firm, an adviser and a reviewer, and the platform will not produce an ownerless output. Under SM&CR the accountability was always going to land on a person. The system’s job is to record who, at the time, so that nobody has to reconstruct it afterwards.

An evidence trail back to source

Frank shows his reasoning against the documents he read. Every finding points at the part of the file it came from, so the reviewer can go and look. That is also what makes disagreement productive: a reviewer who thinks a finding is wrong can see straight away why the system thought otherwise. Exhibit A is one finding and the page it cites.

Human sign-off before anything is recorded

Frank returns a RAG-rated suitability review, Red, Amber or Green, and a human reviewer confirms, amends or overrides every grading. Nothing is auto-recorded. The reviewer’s promotion or demotion of a grade is the outcome; the model’s opening view is an input to it.

Testing for drift

Grading behaviour is a calibration, and it has to be held still on purpose. Prompt changes are versioned and archived, the model runs at a fixed configuration, and a change to how one product type is assessed must not silently move the grading of another. Drift here is a well-intentioned edit that quietly changes what a whole book of files scores.

Exhibit A A finding, and the page it cites
Fact-find extract NFP-2026-0418, p.6

Northgate Financial Planning LLP

Section 4, attitude to risk

Risk profile questionnaire completed
14 Dec 2023
Profile recorded
4 of 7, balanced
Reviewed at advice date
Not recorded
Advice date
14 Feb 2025

Adviser: R. Callender. Client signature on file.

Finding 03 of 11 Amber

Attitude to risk: questionnaire currency

The client’s attitude to risk questionnaire was completed in December 2023, fourteen months before the advice date, and the file does not record that it was revisited when the recommendation was made. Either refresh the questionnaire or record in the file why the December 2023 profile still held in February 2025.

Evidence Fact-find, section 4, page 6
Reference COBS 9.2, suitability

A reviewer confirms, amends or overrides before this grade is recorded.

A reconstruction, drawn for this piece. The firm, the adviser, the file reference and every date are invented, and no real client file is shown. The point is the citation, which is what lets a reviewer disagree with Frank on the evidence.

What this does not mean

None of that solves a firm’s AI governance problem, and we are not going to claim it does. Frank grades leniently on purpose: only clear omissions grade AMBER, and borderline cases grade GREEN. We calibrated it that way because a reviewer handed a wall of amber stops reading the ambers. That is a decision we made and can defend, not a limit of the model, and it is the reason a GREEN is a narrower statement than it looks. A green file is a file where nothing clearly deficient was found. It is not a file certified as beyond criticism.

An AI file review does not catch everything, does not replace a file checker, and is not compliance sign-off. What it does is read the file quickly and consistently, surface what is missing, and show its working. The judgement stays with the reviewer. That is a limitation, and on the review’s own framing of the autonomy spectrum, it is also the point.

The other half of the review

The Mills Review made seven priority recommendations to the FCA Board, and the FCA notes that it drew on a call for input and a survey of more than 5,000 UK consumers. That consumer strand matters for context. Alongside the review, the FCA operates an AI Live Testing service that firms can use, and it has said it is conducting a perimeter review, over roughly three to six months, into consumers’ use of general-purpose large language models for savings, investment, pensions, mortgage and debt decisions.

For advice firms, the perimeter question may be the more consequential one. It asks what happens when a consumer takes a financial decision on the strength of a general-purpose chatbot that sits wholly outside the regulatory perimeter. Set that against a purpose-built, firm-operated, human-supervised tool and the contrast is stark. A firm that can describe the contrast in its own governance documentation, and not in a vendor’s brochure, is in a better position than a firm that cannot.

Where this leaves a firm

As we read it, the interesting question has moved. Permission is settled: a firm may use AI in its file-checking process. What is left is whether the firm can answer four questions about the process it already runs.

Four questions to ask of your own file-checking process

  1. Who owns this output?
  2. What is it based on?
  3. Who signed it?
  4. How do we know it still behaves the way it did last quarter?

Ask them of the process the firm runs today, because a tool it has not bought yet cannot answer any of them. A firm that can answer all four has most of what the review asks for. A firm that cannot has a governance gap that no amount of model quality will close.

A note on this piece

This is commentary, not regulatory advice, and it does not constitute compliance sign-off for any firm or file. Descriptions of the Mills Review and related FCA activity are our reading of what the FCA has published; firms should refer to the FCA’s own materials and take their own advice before acting. If you would like to talk about how any of this applies to your file-checking process, we are happy to have that conversation.