Photo captured at joining, March 2019, and never updated.
| EMP NO | 0047812 |
| NAME | ARCHER, SAMANTHA J |
| START | 11/03/2019 |
| GRADE | C2 · FINANCE SYSTEMS |
| COST CTR | FIN-0210 PAYROLL SVCS |
| SALARY | 41,200.00 |
| STATUS | LEAVER 27/02/2026 |
Working paper · Why data quality decides whether AI works
Sam Archer worked at the Ashworth Group for seven years. Human resources processed her exit correctly, on time, by the book. Seventy-four days later her finance account approved a supplier invoice, because no system in the building knew that three different records were the same person.
ARCHER, SAM
Status: LEAVER
Last working day: 2026-02-27
Exit checklist: complete
12/05/2026 09:41
INV-208441 · Kestrel Fleet Services Ltd
£18,650.00 · APPROVED
User: s.archer2
Every figure on this page comes from one fabricated, internally consistent case file. Cross-reference any two numbers and they will agree. Sam Archer is a fictional composite; the failure pattern is not.
Everyone is wiring AI into their operations this year: approval workflows, HR automation, agents that read systems and act on what they find. All of it inherits one silent assumption: that the records underneath are right. When they are not, nothing crashes. The automation runs happily, confidently, and does the wrong thing at machine speed, with an audit trail that says everything worked.
An agent given reach into the Ashworth Group's three systems would not have caught Sam Archer's surviving approval right. It would have used it. That is the difference between a data problem and an AI problem: almost every AI governance failure starts life as a data quality failure, and it is decided long before any model is involved.
This page is how Nimble AI thinks about data: how it is captured, processed, understood, prioritised, how it goes wrong, and how we catch it. Seven questions, asked of one leaver's records. Every part is something you can operate, and every part ends with what it means for your data.
The case file
The Ashworth Group runs payroll on a green-screen system installed in 1994, HR on a cloud platform migrated in 2022, and finance on an ERP configured by a contractor. Each one holds a record for Samantha Jane Archer, joiner in 2019, mover in 2022 and 2024, leaver in 2026. Each photographed her in a different year, keyed her differently, and disagrees with the other two. Look closely: the five highlighted fields are the five defects this page is about.
The Ashworth Group, its systems and every person named on this page are invented for this worked example. No real customer data appears anywhere. The portraits are AI-generated images of a person who does not exist.
Photo captured at joining, March 2019, and never updated.
| EMP NO | 0047812 |
| NAME | ARCHER, SAMANTHA J |
| START | 11/03/2019 |
| GRADE | C2 · FINANCE SYSTEMS |
| COST CTR | FIN-0210 PAYROLL SVCS |
| SALARY | 41,200.00 |
| STATUS | LEAVER 27/02/2026 |
Photo recaptured at the 2022 migration, three years into her career.
| Employee ID | SA-1178 |
| Name | Sam Archer |
| Hire date | 2019-03-11 |
| Role | Finance Systems Analyst, grade C2 |
| Department | Finance Systems, FIN-0455 |
| Manager | P. Lindqvist |
| Status | Leaver, last working day 2026-02-27 |
Security pass photo, October 2024, when the approval right was granted.
| User | s.archer2 |
| Display name | S J ARCHER |
| Created | 10/03/2024 (US locale) |
| Role | AP invoice approver |
| Approval limit | £25,000.00 |
| Cost centre | FIN-0455 |
| Status | ACTIVE · last action 12/05/2026 |
Three photographs, taken years apart, of one person. The wardrobe changes because careers do: a payroll administrator in 2019, an HR operations analyst in 2022, a finance systems analyst in 2024. The systems never noticed it was the same career. No field is shared across all three records, not the name, not the identifier, not even the format of the dates.
How was it captured?
The surface you will one day have to govern is decided at capture, before anyone has asked what the work needs. Sam's 2019 joiner form asked for everything any downstream system had ever wanted.
Take everything, decide later. The joiner form collects eighteen kinds of personal data because eighteen systems each asked for one once, and removing a field needs a meeting while keeping it needs nothing.
Every field taken must now be classified, secured, justified and deleted on request, in three systems at once. When a subject access request arrives for Sam Archer, somebody has to find and reconcile all eighteen, in systems that cannot agree what she is called.
Separate the working set from the live surface, and make intake a decision rather than a default. In our own product work, source material is processed offline in a controlled working set, and the live surface accepts one field of personal data: an email address. Ask what has to cross each boundary, then take only that.
Before you ask whether your AI is safe, ask what your systems are holding and which of it anyone actually chose. The fields you can justify are the fields you can govern, and every field that never crosses a boundary is one you never have to.
Eighteen kinds of personal data on one joiner form. Not one of them has been questioned yet.
How was it processed?
Sam's exit was processed from a checklist. The checklist was built from memory, keyed on the HR identifier, and the three systems that key her differently were the three it missed.
Process by memory, not by inventory. A leaver is offboarded by checking the systems someone can name, in the order someone remembers them. Coverage is whatever the checklist happens to say, and nothing measures the gap between the checklist and the estate.
The gap surfaces months later as an account that still works. At Ashworth, the exit checklist closed nine systems keyed on SA-1178 and never saw 0047812 or s.archer2. The approval in May was not a lapse of process. The process ran perfectly over incomplete coverage.
Full coverage with deliberate overlap, and a citation for every fact. In our own product the whole source is split into overlapping segments and read in full, so nothing is lost at a seam, and every extracted claim cites the segment it came from. The same shape applies to a process: cover the whole inventory, overlap every handover, and record where each fact came from so coverage can be checked rather than trusted.
Ask one question of any process or supplier: which parts did you check, and can you show me? If the honest answer is a sample, the coverage is unknown and the confidence is borrowed.
Nine of twelve closed. The three in red key Sam Archer under a different identifier, so the checklist never saw them. Payroll was told to stop her pay, but her operator login on the terminal, her ERP approver account and her warehouse access were nobody's job to find.
Is it understood? · The whole argument is here
Anyone can store data. Every consultancy will make yours “AI-ready”. The gap between having data and understanding it is where governance actually lives, because you cannot govern, review, or hand to an agent a person your systems cannot even agree exists.
PAY-2000 · 1994ARCHER, SAMANTHA J EMP 0047812
START 11/03/2019
COST CTR FIN-0210
PeopleFirst · 2022Sam Archer ID SA-1178
Hired 2019-03-11
Dept FIN-0455
Ledgerline · ERPS J ARCHER User s.archer2
Created 10/03/2024
Approver £25,000
One person, resolved from three records that shared no key. Joined 11 March 2019, moved twice, left 27 February 2026. A fictional composite.
This is the view no system at Ashworth could produce: the whole person, with every grant in one place. Every red flag was invisible until the three records were resolved.
The same person exists as three records in three systems and nothing joins them. Every count is then wrong, and every question about the whole person is answered from a part. The access review passes, because no reviewer can see what they were never shown.
Access granted for a role two moves ago stays live, invisibly. At Ashworth that meant an active £25,000 approval right on a leaver's account, and roughly £133,000 of salary charged, across nearly four years, to a cost centre Sam left in 2022, both sitting in plain sight of anyone who could have seen all three records at once. Nobody could.
Resolve every entity to one canonical record before anything else is attempted. This is the mechanism we build first, and we proved it on the hardest material we could find: one entity named several ways is reconciled to a single record, links are derived from where things occur together, and a reference that does not resolve is dropped rather than left dangling, so an unresolved thing is visibly absent instead of quietly wrong.
Understanding is the step everyone skips and the one governance depends on. Two questions expose it: can you resolve a person to one record, and can you show what that record is joined to? If either answer is no, every review, workflow and agent downstream is operating on a part and calling it the whole.
Where the analogy stops. This part argues a principle: resolve the subject before you govern it. Nimble AI does not sell an identity-based entitlement system. The resolution mechanism in our own product work operates on records and text, and the joiners, movers and leavers setting is a worked illustration of the same discipline, not a product claim.
Why this is part three of seven and the centre of the page. Capture and processing decide what exists. Understanding decides whether anyone can see it whole. Every part after this one, prioritisation, the defects, the catches, depends on the resolved view existing first. Skip this step and the rest is decoration.
What gets looked at first?
Every review queue ranks, and every ranking cuts. The order gets argued about; the position of the line usually does not, and neither does what quietly falls beneath it.
Nothing decides what matters, so something incidental decides instead. Accounts are reviewed in the order the requests arrived, or by last month's activity, or starting with whoever chased hardest. None of those orderings has anything to do with risk.
Ranked by activity, s.archer2 is the quietest account on the list: one action in three months. It sits last, below the review line, never looked at. The one action was an £18,650 approval by a leaver. A queue with no risk ranking gives no signal that the wrong thing fell below the line.
Rank across the whole estate, not one system at a time, and make the omissions visible. In our own product, prominence is computed across a whole work and recomputed across a series, so something that matters cumulatively is not ranked out by one quiet volume, and a quality report is produced at build listing exactly what did not map. The gaps are surfaced, never silent.
Ask of any ranked queue: what decides the order, and can it show you what fell below the line? A list that cannot display its own omissions is telling you half of what it knows.
The line sits after rank 3 of 6. s.archer2 is rank 6: below the line, never reviewed. Its one action last quarter was the £18,650 approval.
The five defects
Scroll back up to the three records and you can find all five with your own eyes. None of them is exotic. Each one is small, boring, years old, and load-bearing. That is what real data defects look like.
A wrong or mismatched value is copied between systems until three systems agree with each other and all three are wrong. Nothing marks the first copy as unverified, so the next system copies the copy, and the defect simply becomes the record.
Each defect compounds the others. The mismatched keys broke the leaver sweep; the surviving account met the stale approval right; the ambiguous date format means even the forensics are arguable. Five small defects, one £18,650 approval by a leaver, and an access review that passed the whole time.
Assume defects exist and hunt them as a first-class activity, not a clean-up. Reconciliation surfaces the mismatches, the build produces a quality report of everything that did not map, and anything that cannot be traced to its source is treated as suspect. We do not certify data clean; we make its defects visible, cited and fixable.
Your estate has its own five. They are not where anyone is looking, because the places anyone looks are the places already keyed and reconciled. The question is not whether they exist; it is whether anything in your organisation is capable of showing them to you.
Payroll holds the legal name from her 2019 right-to-work documents. HR holds what she typed into the 2022 migration portal. Any automated match on name now needs fuzzy logic, and fuzzy matching is how s.archer2 got created in the first place: an S. Archer already existed.
Payroll writes day first. The ERP module was configured to a US locale by a contractor, so its dates read month first: 10/03/2024 is the 3rd of October, not the 10th of March. Every export joining these systems has to know that, and nothing records it anywhere except in one leaver’s head.
The HR migration re-keyed every employee in 2022. The ERP mints usernames from initials. No system carries another system’s key, so there is no join anyone can compute, only a join a human infers. The leaver sweep was automated; the inference was not.
When Sam moved teams in July 2022, HR updated. Payroll did not, and payroll is what drives the ledger. Roughly £133,000 of her salary was charged to a team she had left, across 44 months, and both departmental budgets were wrong in every review that ever used them.
The exit checklist closed everything keyed on SA-1178. The ERP account is keyed s.archer2, so it survived, with its £25,000 limit intact, and approved INV-208441 seventy-four days after her last working day. Whether Sam pressed the button or a colleague with her saved credentials did is a question the records cannot answer. That is the point.
How we catch them
This is not a picture of a check. It is a check, running in your browser: a deterministic, character-by-character comparison against the signed source, with no model in the loop and no judgement call.
A record is never checked against its authoritative source, or it is checked by the same kind of system that produced it. Values get retyped, tidied into house style, or assembled from two systems that never agreed. The result reads perfectly, which is precisely the problem.
A single altered field destroys trust in every field beside it, including the ones that were right. It is the one thing a regulator, an auditor or a former employee can verify independently, and the one thing that ends the argument if it is wrong.
Verify at the point of use, deterministically, and test the checks by breaking them. In our own product every quotation shown to a reader is machine-verified character for character against the source before it can ship: the check runs at build and again at deploy, and either failure blocks publication. And because a check nobody has seen fail is not known to work, we break things on purpose: we have deliberately corrupted an asset, watched the gate fail, restored it and watched it pass.
Honest scope. A character-for-character match proves a value was not altered in flight. It cannot prove the signed source was right in the first place, and it says nothing about derived or summarised data, which get human review instead. We would rather tell you exactly where a guarantee stops than let you assume it covers everything.
Source: the signed leaver form in Sam Archer’s case file. Fictional, like everything else in it.
Edit the line above, any character, anywhere. It fails the instant it differs and passes the instant it is restored, every time, with no delay and no discretion.
What that means
Run Sam’s seven years in ten seconds. Every move grants; nothing revokes; and the one view that would have shown it, the resolved person, never existed until part three of this page built it.
Access is treated as a setting configured once, rather than something that accumulates on every move and must be deliberately walked back. Nobody revokes what nobody can see, and nobody can see what three unjoined systems each hold a third of.
What a person, a workflow or an agent can reach only ever grows. At Ashworth the growth was invisible for seven years and survived the leaver process itself. The cost arrived later, on an invoice, with a clean audit trail attached.
Make change visible at the moment it happens. In our own product work, when something comes into scope the interface says so explicitly rather than disclosing it silently; the principle is that a grant, a move or a reveal should announce itself, because silent accumulation is how every estate on this page got the way it is. The engagement is done for you, and the human in the loop is the service.
Before wiring AI into your operations, buy the resolved view first: one record per person, every grant visible, every change announced. It is not the glamorous half of AI adoption. It is the half that decides whether the other half works.
Joins as Payroll Administrator. PAY-2000 operator access granted.
March 2019. One grant, one system, one identifier. Nothing wrong yet.
Three systems photographed three different years of one career, and never noticed. Sam Archer is a fictional composite; the portraits are AI-generated images of a person who does not exist, aged across roles on purpose.
The evidence · The same method, somewhere much harder
A worked example proves a method is coherent. Running it for real proves it works. We ran this exact discipline, capture, coverage, resolution, ranking, defect-hunting, deterministic verification, over a set of full-length literary manuscripts: no data owners, no classification scheme, no regulator, no IT policy to lean on. If the method holds there, the structured, well-owned data in your estate is the easy case. Everything below is a real screenshot of the shipped product, not an illustration.
Six authors, five genres, 1813 to 1902, plus one invented contemporary sample. These are the live titles published on Show Me My Book, so you can open any of them and check the output against the source yourself. The same mechanism runs on manuscripts we cannot publish. One schema, one engine, one set of deploy gates.
Jane Austen, Pride and Prejudice, 1813. Public domain.
A page that says “we do reconciliation” asserts competence. Here is what demonstrating it looks like. While building this capability we caught, in our own output: a fabricated in-world detail that propagated across several planning documents before anyone noticed, exactly the propagation pattern from part five; a hallucinated character detail that reached a draft marketing page; and a review step that approved a bad image render and reported it good. Each one was caught by the discipline on this page, verification at the point of publication rather than the point of authorship, not by anyone’s memory or diligence.
And because a check that has never failed is not known to work, we tested the gates by breaking them: we corrupted an asset on purpose, watched the gate fail, restored it, and watched it pass. Ask your own estate the same question: which of your controls have you ever actually watched fail? A control that has only ever returned green has not been tested. It has been assumed.
The offer
Not necessarily a leaver with a live approval right; perhaps a supplier that exists three times, a cost centre nobody has trued up, a date field two systems read differently. You will not find them by asking whether your AI is safe. You will find them by asking whether your data is understood, which is the question this whole page has been asking.
Nimble AI is an AI governance consultancy. The engagement is done for you: we apply this discipline to your data the way we applied it to our own, and we show you the working, including where the guarantees stop. Start with the free scorecard, or talk to us about your data directly.