Then the audit arrives, and someone asks a question the pilot never planned for: show me evidence that this control operated throughout the period.
That question is not hostile and it is not unusual. It is the question auditors have always asked. What changes when an agent is doing the work is not the standard — it is what the organization has available to answer with.
Auditors Test the Control, Not the Outcome
It is worth being precise about what is being examined, because this is where agentic reconciliation programmes most often misread the risk.
Whether the reconciliation was correct is not the whole question. The question is whether a control existed, whether it operated as designed, whether it did so consistently across the period, and whether the organization can demonstrate that after the fact. Where the reconciliation is a key control in a controls-focused audit, a correct ledger with no evidence of how it got that way can itself be a problem — and where the auditor is taking a substantive approach instead, the practical consequence is different: they simply decline to rely on the control and test the balances harder. Either way, evidence changes the cost of the audit.
For manual reconciliation the evidence set is well understood: a preparer, a reviewer who is not the preparer, dates that sit in a sensible order, the exception population, what was decided about each exception, and a sign-off. Finance teams have been producing this for decades, largely without thinking about it.
An agent does not produce that by default. It produces something else, and whether that something else is sufficient is a question worth answering before the audit rather than during it.
A note on scope before going further: what follows describes general practice, not audit or accounting advice. Requirements differ by jurisdiction, regulatory regime, auditor and entity. The fuller caveat is at the end, and it matters.
The Agent Changes What Kind of Control It Is
This is the part that surprises people, and it is partly good news — though as the later sections argue, only partly.
Audit approaches generally distinguish between controls a person performs and controls the application performs. Manual controls are typically tested by sampling — a number of instances across the period, examined individually. Automated controls can often be approached differently, because IT processing is consistent in a way human processing is not, and because software does not get tired or distracted. Establish that the control is configured and implemented as claimed, and that evidence — combined with sound general IT controls around change, access and operation — can support a conclusion about how it behaved across the period.
Two cautions belong immediately next to that. Testing a single instance is rarely enough on its own: logic can be correct in December and wrong in January, and a one-item test will not find that. And the reliance rests on those general IT controls being genuinely effective, which is a separate piece of work with its own evidence requirements.
The practical consequence is that moving reconciliation to an agent does not remove audit effort. It relocates it. Less sampling of individual reconciliations, more scrutiny of configuration, change history and access. In our experience the teams caught out are the ones that expected a lighter audit and changed nothing about how they manage configuration. Deciding which checkpoints stay with a person is a related question we have written about in where human oversight still belongs.
The “Throughout the Period” Problem

Reliance on an automated control rests on an assumption that is easy to miss: that the control was the same control all period.
If the agent’s thresholds were widened in March, or its scope extended to a second entity in May, or its underlying model version changed at some point without anyone recording when, then there may not have been one control operating for twelve months. Where the versions differ substantially, each may need to be considered — and evidenced — separately.
This is why configuration change history stops being IT housekeeping and becomes audit evidence. It is also why the absence of that history is worse than it looks. If nobody can say which configuration produced which result, reliance may fail not just around the change but across the whole period, because there is no basis for concluding anything about any part of it.
The practical version: if you cannot produce a dated, approved list of every configuration change made during the period, you are asking your auditor to take the period on trust.
Where Agentic Reconciliation Strains the Model
There is a genuine complication here, and it is more honest to name it than to imply the old categories fit neatly.
Part of why automated controls are efficient to audit is consistency: a conventional matching rule does the same thing every time, and unlike a person it does not have an off day. Test the rule, confirm nothing changed, and conclusions about the population follow. An agent that works from confidence levels, and that may reach a different conclusion on two similar items, offers a weaker version of that guarantee. The variability that makes agents useful is the same variability that makes the classic reliance argument harder to run.
The reasonable expectation, at least for now, is that agent-performed reconciliation calls for more evidence rather than less. Framework-level guidance on controls over generative and probabilistic systems has started to appear, and it makes much the same point about moving from deterministic tools to models with variable outcomes. What has not settled is how individual auditors apply it. Expect your auditor to be working out their approach, and expect that approach to differ from the firm down the road.
What the Auditor Will Ask For
The specific list will vary by auditor, entity and regulatory regime, but the shape of it is fairly predictable.
- What the agent is, what it is permitted to do, what it is not permitted to do, and who owns it by name.
- The configuration as it stood at the start of the period, and every change since, with date, approver and reason.
- Access rights over that configuration — who could alter thresholds or scope, and whether those people are segregated from the people who post entries.
- A complete record of what the agent did, item by item, with enough basis to understand why.
- The exception population, what escalated, what a person decided, when, and on what information.
- Evidence that human review was substantive, not merely that an approval was recorded. This is the hardest one, and the one most likely to be thin — and it is thin for a structural reason, which is that review capacity is finite and rarely counted. That is the subject of how many agents a finance team can actually supervise.
- Model or version changes affecting behaviour during the period.
- What happens when the agent is wrong — how that gets detected, corrected, and by whom.
The Audit Evidence Gap Nobody Anticipates

Most teams prepare to show what the agent did. Rather fewer can show what it decided not to do.
You cannot draw conclusions about a population you cannot define. If the log captures actions taken but not items the agent examined and passed over, the population is unknowable, and no amount of well-documented activity fixes that. Silent non-action leaves no trace unless the system was designed to leave one.
This is worth checking early, because it is usually a logging design decision rather than a policy problem, and far cheaper to fix before a period closes than after.
Designing for the Audit Before the Pilot
None of this argues against agent-performed reconciliation. It argues for deciding the evidence model at design time, when it is nearly free, rather than at year end, when it is not. Which process to put first is a related decision, covered in what to automate first.
- Walk your intended evidence set past your auditor before go-live. They cannot design it for you — independence rules see to that — but they can tell you early whether what you are planning to retain is likely to be enough.
- Version and freeze configuration, and put changes through proper change management with approval recorded.
- Log intent and basis, not just outcome. “Matched” is an outcome. Why it matched is evidence.
- Capture the null set — items considered and passed over, so the population is definable.
- Make human review leave a substantive trace: what was examined, what was changed, what was overridden and why.
- Keep segregation of duties intact across configuration, execution and review, as you would for any financial control.
- Rehearse it. Ask your own team to produce the evidence for a single month before an auditor does. What comes back will be instructive.
The Underlying Shift
Reconciliation is a good early candidate for an agent precisely because it is repetitive, high-volume and rule-shaped. That is also why it tends to be a well-controlled process that auditors look at closely. The two facts arrive together, and the second is easy to forget while building the business case for the first.
The organizations that will handle this comfortably are not the ones with the most capable agents. They are the ones that treated auditability as part of the deployment rather than a consequence of it — which mostly means deciding, in advance, what a person will be able to prove in twelve months’ time.
DAX Software Solutions works with finance teams on what comes before the pilot: whether the ERP, the data and the governance can support an agent at all, and what the evidence model around it needs to specify. In practice that is a narrower conversation than it sounds — which logs exist, what they capture, who can change the configuration, and what a reviewer leaves behind when they look at something. Those decisions are cheap to make at design time and expensive to reconstruct later. We have also written about where the close actually stalls, which is usually the process an agent is brought in to shorten.
Scope and caveat. This article describes general practice and is not audit, accounting or legal advice. Requirements differ by jurisdiction, regulatory regime, auditor and entity, and practice for probabilistic systems is actively evolving. Confirm the treatment with your own auditors and qualified advisers before relying on any of it.
If you are putting an agent anywhere near a financial control, the evidence conversation belongs at the start of the project. Get in touch if it would help to have it with someone who has had it before.

