A finished pale block arriving at the end of a conveyor rail facing an empty lit platform, with the production machinery behind it unlit and idle

What the Finance Analyst Does When the Work Becomes Adjudication

The reconciliation is already done when the analyst opens it. So is the coding, the matching, and a first pass at the variance explanation.

What is left is a decision — accept it, send it back, or escalate. Mostly accept.

That is not a lighter version of the old job. It is a different job, and it has a name. Judging work produced by something else, case after case, against a standard, is adjudication. It has been studied for decades, almost entirely outside finance, and what is known about it is not especially comforting.

The Title Did Not Change

There was no announcement that the analyst’s output would stop being work product and start being a verdict.

The role moved by accretion. Invoice capture in Dynamics 365 Finance took invoice registration. Matching rules took the easy half of the bank rec, and Copilot-assisted bank reconciliation in Business Central — in preview as this is written — is aimed at more of the rest. Those are two different products with different capability sets, so how far this has gone depends on which one you run. But the direction is consistent, and each step moved a little more of the day from producing to reviewing.

What did not move is everything around the role. The job description still describes preparation. The performance conversation still rewards throughput. The training plan, if there is one, teaches the system rather than the judgment. And the skills that make a person good at producing a reconciliation overlap with, but are not the same as, the skills that make them good at ruling on one.

The Caseload Is Selected for Difficulty

A stream of uniform pale tiles flowing past while an arm diverts only the jagged dark blocks up a narrow ramp to one small lit platform, beside an empty storage bay

The shape of the work is set by configuration — an approval limit, a routing rule, and increasingly a confidence threshold in whatever decides what an agent handles alone. The first two are stock ERP objects. The third usually is not: graduated confidence routing tends to get built in AI Builder or a custom flow rather than switched on in Finance, which makes it easy to change without telling anyone. It is one instance of a larger pattern in how a queue gets its shape, but what matters here is simpler. An item reaches a person because a machine judged it hard.

Lisanne Bainbridge described the consequence in 1983, writing about automated industrial process control (Automatica 19(6), 775–779). Automate the routine and the operator is left with the tasks that could not be automated, which are the difficult and unusual ones. Meanwhile the routine work that kept their skills sharp is gone, so competence decays precisely where it is most needed. Her conclusion was that heavily automated systems make the case for more operator training, not less.

Translate that into a finance function. The ordinary items an analyst used to grind through were not only output. They were practice — how a person learns what a normal accrual looks like, which suppliers habitually invoice wrong, and what a plausible variance feels like before anyone explains it. That library gets assembled at low stakes on easy cases, and it is what makes a judgment on a hard case worth anything.

Remove the easy cases and the library stops being built. The reviewer is asked to be the expert on exactly the items where expertise is scarcest, using intuition they no longer have a way to acquire.

Judging Leaves No Reasoning Behind

There is a second structural change, easy to miss because it looks like a saving.

When you produce the work, the work is the record of your reasoning. What you tied out, what you accepted, where you stopped — it is all in the artifact, and anyone can reconstruct the thinking from it.

When you judge work someone else produced, your reasoning leaves no artifact unless something forces one. A comment field exists in most approval workflows. It is optional, so it is usually blank. Two analysts can approve the same item for opposite reasons — one satisfied the variance was explained, one satisfied it was immaterial — and the record says the same thing both times: reviewed, approved, timestamped. That compounds, because the next reviewer facing a similar item has nothing to inherit. It is also why review that leaves a substantive trace matters for more than the auditor.

What Adjudicative Systems Have That Finance Review Does Not

Institutions that have judged things at volume for a long time tend to have some version of the following, though how much of each varies widely by forum:

  • Reasons, as part of the decision and not a courtesy.
  • A record of the reasoning, not only the outcome, so a decision can be examined later on its merits.
  • Consistency devices. Published criteria, worked examples, precedent — and, where decisional independence allows it, calibration between decision-makers.
  • An appeal route, which is also how systematic error becomes visible.
  • Review of the decision-maker, not only of the cases — someone looks at the pattern in one person’s decisions.
  • Procedural intensity scaled to stakes. A small claim does not get a full hearing, and what it does get is decided deliberately.

Against that list, most finance review has an approve button, a monetary threshold and a four-eyes rule. Not nothing — four eyes is doing real work — but a thin kit for what has become a judging function.

Consistency Is the Part That Does Not Solve Itself

Two identical chambers fed identical pale inputs, their output rails diverging sharply — one climbing into warm light, the other dropping into an unlit void

Consistency deserves attention first, because the evidence on how badly it goes unmanaged is unusually blunt.

Refugee Roulette, published in the Stanford Law Review in 2007 by Ramji-Nogales, Schoenholtz and Schrag, examined US asylum decisions and found that Colombian applicants in the Miami immigration court had a 5% chance of prevailing before one of that court’s judges and an 88% chance before another judge in the same building. Same law, same courthouse, same applicant nationality — though the study compared grant rates across cohorts rather than matching cases individually, which was the main objection raised against it.

The follow-up matters more than the headline. A US Government Accountability Office review in 2008 found significant variation across courts and judges. Disparity then narrowed for a couple of years and widened again: TRAC measured an average 27% increase in decision disparity over 2011–2016, and a second GAO review in 2016 found the variation persisted after controlling for case and judge characteristics. Some argue a degree of variation between decision-makers is unavoidable regardless, and they have a point.

The stakes are not comparable and the analogy should not be pushed — an asylum claim is not an invoice, this is not legal commentary, and questions about adjudicative process in a regulated setting belong with a qualified professional. What transfers is the structure. Where trained people judge similar cases without shared criteria, worked examples or calibration between them, variation is large, it survives controlling for the cases, and it stays invisible until measured.

Variance between reviewers is rarely measured in operational finance. It is routine in audit quality review — the discipline that has thought hardest about review quality is the one that checks its reviewers against each other.

The Failure Mode Is Agreement, Not Disagreement

The intuitive worry about human review of machine output is that people will resist it. The better-supported worry runs the other way.

Parasuraman and Manzey’s 2010 review in Human Factors (52(3), 381–410) gathered the empirical work on automation complacency and bias. The effect appears under multiple-task load, when other work competes for the reviewer’s attention — the ordinary condition of a close week. It shows up in experts, not only novices, though experts caught more failures than non-experts did. And the uncomfortable part is what does not fix it: extended practice did not remove the effect, and for automation bias, neither did explicit instructions to verify the aid’s recommendations. The authors note that some specific training designs showed promise; generic exhortation is not among them. Which is awkward, because “remind the reviewers to be careful” is the standard mitigation.

Endsley and Kiris found in 1995 (Human Factors 37(2), 381–394) that people supervising an automated task passively had lower situation awareness and took longer to decide once it failed, and that how much control the operator retained moderated the loss. Passive review is a weaker cognitive position than doing, not just a cheaper one.

Two adjacent pieces point the same way: explanations can increase acceptance regardless of whether the output is right, and review degrades rather than stops when load outruns capacity. An override rate drifting toward zero is not evidence that the agent got good.

Three Objections Worth Taking Seriously

Adjudication is, for a lot of analysts, a better job than production was. Deciding is more interesting than keying, and treating the shift as a loss patronizes people who mostly do not experience it that way.

The variance finding also points in the opposite direction. A deterministic rule applied to every case will be more consistent than a group of reviewers — though consistency is not correctness, and a rule can be reliably wrong. That holds for rules; it does not transfer cleanly to generative models, which vary between runs.

The hardest objection is proportionality. You cannot run a hearing for every invoice. Reasons cost time, and time is what automation was bought to save — demand a rationale on every decision and the bottleneck has moved up a layer while generating a pile of thin justifications nobody reads. The kit cannot be adopted wholesale. The usable version is choosing which classes of item get which treatment, and that choice is the actual work.

Six Things That Are Available Now

  1. Say plainly that the job changed, then change the job description and the performance measures to match. A role measured on throughput produces throughput, which is the wrong output for a judging function.
  2. Protect practice on purpose. Route some routine volume to people not because it is efficient but because it is how judgment stays calibrated. Sampling items the agent already cleared does the same job and yields a second-opinion signal as a by-product.
  3. Require a reason on a defined class of decisions — above a value, on a named exception type, on anything overriding the agent. The point is that the class is chosen rather than assumed.
  4. Compare reviewers against each other at least once. Take twenty borderline items from last month and give the same twenty to two reviewers who did not see them, working separately. Compare the two sets of answers against each other rather than against a correct answer, which mostly does not exist.
  5. Treat the override rate as a signal rather than a target. Falling override against rising volume is the pattern worth an alarm.
  6. Give the reviewer an escalation route that is not their own judgment. A second-opinion path is how systematic error surfaces.

Most of the six are design decisions about the review step.

Who Designs the Review Step

On most projects, no one does. The agent gets a specification. The workflow gets a specification. The human at the end of it gets a permission.

Which is the gap DAX Software Solutions works in as a Dynamics 365 implementation and support partner. Specifying an agent is mostly a question of boundaries — what it may touch, how far it may go. Specifying the review step is a question of standards: what the reviewer is checking for, what evidence arrives with the item, what they are expected to write down, and where it goes when they are not sure. All four are answerable while a process is being designed and awkward to retrofit once it is live. If agents are being scoped into your D365 estate, the review step is a deliverable and not a leftover. Ask us to scope it with the rest.

New capability will keep arriving on its own schedule, release by release. What the person at the end of it is meant to be checking for is not a platform question, and it will not turn up in the release notes.

Scroll to Top