Zubin Deepak Rajasekar
Back to home page

Precedent

Category

In progress

Year

2026

Checking complaint decisions against how the Financial Ombudsman has actually decided comparable cases. Started August 2026.

Context

UK firms must resolve complaints within eight weeks under DISP. Get the call wrong and it costs redress, Ombudsman case fees, and Consumer Duty scrutiny. Handlers are deciding against precedent they cannot see, because the Ombudsman's reasoning sits in hundreds of thousands of individual published decisions.

The thinking

The obvious build here is an uphold predictor: will the Ombudsman find against us. It is easier and it is the wrong product. Its most natural use is working out which complainants are unlikely to escalate, which is the opposite of a Consumer Duty outcome.

So the question is consistency, not risk. Does the proposed handling of this complaint match how comparable cases have actually been decided, and when no comparable cases exist, does the system say so rather than guess.

What I am building

Fact extraction over roughly 2,000 published Ombudsman decisions in one product line: affordability and irresponsible lending on credit cards and unsecured loans.

Retrieval of comparable decisions with the reasoning attached, so the handler sees precedent rather than a verdict with a score on it.

Three outputs, not two: aligned, misaligned, and no comparable precedent found. The third is the one the project is really about.

Deliberately excluded: any feature predicting whether the customer is likely to escalate. That single feature is what turns this into the product described above.

How I will know it works

The thresholds are set and published before the system exists, which is the whole point of the exercise. Two of them, both relative rather than absolute: it has to beat the base uphold rate for the product line, and it has to beat the firm's current overturn rate. A model that predicts the majority class scores well and is worthless, and nobody measures the human baseline it is supposed to be replacing.

Published results: accuracy against base rate, a reliability curve, a coverage against accuracy trade off, abstention behaviour on a held out sub-category, and the run to run noise floor.

The open question

FirmMemory had a retrieval thinness gate that I calibrated and then removed, because no defensible threshold existed and there were no labels to find one with.

Here there are labels. So the coverage floor gets calibrated properly. If a defensible threshold exists I will ship it and show the calibration. If it does not, I will remove it again and publish that instead. I do not yet know which.

The limit

In progress, and nothing is measured yet. Published Ombudsman decisions are the outcomes of escalated complaints, not of all complaints, so the corpus is skewed towards disputes that went the distance. Everything above is a plan with thresholds attached, which is a weaker claim than a result and should be read as one.

More projects