A finance leader should automate an action only when the organization can define an acceptable result, detect important failures, and recover from them within the process’s risk tolerance. Repetition alone is insufficient. A recurring task can still contain ambiguous facts, consequential judgment, or authority that should remain with a qualified person.
The useful unit of analysis is the individual action. An invoice process can contain automated extraction, deterministic matching, human resolution of a disputed receipt, and a separately authorized payment. Labeling the whole process “automated” hides these distinctions. It also makes it difficult to explain who remains accountable when something goes wrong.
For a controller building an automation backlog, the immediate decision is how far to delegate each action and what evidence would justify expanding that boundary. Start with the consequence of an incorrect action, the quality of available evidence, and the ability to reverse or contain the result. Then consider speed and labor savings.
Split preparation from authority
Preparation assembles information or proposes a result. Authority permits the result to affect records, commitments, or money. Those functions can be separated even when the same application performs both.
For example, a service may gather reconciliation evidence, identify a likely match, and explain the fields used. It need not have permission to post an adjustment. A model may draft a variance commentary without publishing the management pack. An automated scheduler may prepare a recurring journal while the organization’s policy reserves approval for a reviewer.
This separation creates useful intermediate steps. The organization can learn whether preparation is accurate before granting broader execution rights. It also prevents a common mistake: assuming that because a system is good at reading a document, it can safely determine the accounting consequence.
Create an action register with the input, proposed output, affected system, required authority, failure consequence, and recovery method. Record whether the action follows an explicit rule or infers a result from patterns. Deterministic rules can be wrong, and probabilistic systems can be useful; the distinction matters because their testing and exception handling differ.
Use an automation boundary rather than a binary label
Working heuristic: preparation and execution authority can be assigned separately within one finance process.
Open full-size diagram
A practical working heuristic has three treatments. Automate execution where the case is well defined, the necessary evidence is complete, and the permitted action is bounded. Assist a human where preparation is valuable but ambiguity or consequence requires review. Keep the action manual where the decision cannot be adequately specified or controlled at a reasonable cost.
An exact match between two complete, independently sourced records may qualify for automated reconciliation under an approved rule. A likely match involving a short payment and a disputed deduction may be better presented as a proposal. A novel contract interpretation may require a specialist to gather facts and reach a documented conclusion.
The boundary should include explicit exclusions. Examples are changes to supplier payment details, transactions outside the trained or tested population, unusual related-party activity, incomplete entity identification, and cases exceeding the process’s approved exposure limits. These are illustrative operating choices; the organization must assess its own risks rather than adopt them as a universal policy.
Human review is not automatically effective. The reviewer needs the evidence, enough time, suitable expertise, and the ability to reject the proposal. A queue that permits only “approve” or makes investigation impractical turns nominal oversight into a formality.
NIST’s January 2023 AI Risk Management Framework addresses defined human oversight and operator proficiency. Its relevance is to the design of responsibility around AI; it does not certify a finance task as safe to automate. NIST AI RMF 1.0, MAP 3.4 and MAP 3.5
Test the difficult population before choosing a target rate
An automation rate measures the proportion of cases handled through a chosen path. It says little about whether important errors were accepted or whether the remaining cases became harder. A high target can encourage teams to widen tolerances until too many questionable cases pass.
Build an evaluation set covering ordinary cases and known failure mechanisms. Include incomplete inputs, duplicate messages, changed reference data, conflicting evidence, reversals, and conditions at decision thresholds. For inferred classifications, separate the data used to tune the system from the cases used to evaluate it. Otherwise the evaluation can exaggerate performance on new work.
Measure errors according to their business consequence. A missed match creates work; an incorrect accepted match can hide an unresolved balance. Those costs differ. Reviewers should examine both accepted and rejected populations so the system does not appear safe merely because it routes everything to a person.
Use a shadow period where the system proposes actions without executing them. Compare results with an independently reviewed reference, investigate disagreements, and record the circumstances under which the system should abstain. Only then decide whether a bounded execution path is justified. A pilot with no observed errors does not prove that rare consequential failures are impossible.
A hypothetical reconciliation business case
Suppose a finance team reconciles 6,000 items each month. In this hypothetical planning example, the current process requires 100 hours in total. A proposed design routes 4,800 items through approved exact-match rules, presents 900 possible matches for review, and sends 300 incomplete or conflicting cases for investigation.
Assume candidate review takes two minutes per item and investigation takes six minutes. The human queue would require 1,800 plus 1,800 minutes, or 60 hours. If rule maintenance, monitoring, and quality checks require another 20 hours, total operating effort is 80 hours. The modeled reduction is therefore 20 hours, rather than the 80 hours suggested by the 80 percent exact-match rate.
These assumptions are illustrative, not a performance benchmark. Actual effort depends on case complexity, control design, system reliability, and the distribution of exceptions. If the 300 investigations take ten minutes rather than six, they consume 50 hours by themselves. Total effort becomes 30 hours of candidate review plus 50 hours of investigation plus 20 hours of maintenance and checks: the original 100 hours.
The design could still be worthwhile if it improves traceability or allows staff to identify issues earlier, but those benefits should be assessed separately. Conversely, a labor reduction does not justify accepting errors outside the organization’s tolerance.
The controller’s next experiment is clear: measure exception effort on representative cases and test whether the exact-match rule incorrectly accepts any meaningful mismatch. That evidence is more useful than negotiating a higher straight-through-processing target.
Hypothetical operating model. At ten minutes per investigation, total effort returns to 100 hours. Exact-match proportion is not the same as effort reduction.
Open full-size diagram
Engineer an exception service
An exception queue needs more than a destination folder. Every item should show why it was routed, the evidence available, the action required, the owner, and the latest useful resolution time. Keep technical failures separate from business exceptions so an accountant is not asked to resolve an interface outage by manually correcting records.
Assign service expectations according to consequence. An exception blocking a time-sensitive reporting decision needs a different response from an incomplete descriptive field. Provide an escalation route when the assigned owner cannot resolve the facts. Monitor age and repeated handoffs, because neither disappears when data entry is automated.
Return resolved exceptions to the process carefully. A reviewer may correct one case without changing the general rule. Changes to rules require their own approval and tests. Otherwise the system can learn an inappropriate generalization from an exceptional workaround.
Retain the inputs and rule or model version used for consequential actions. Record overrides with reasons. For an automated action that fails partway through, the recovery procedure should determine what has already occurred before retrying. A blind retry can duplicate an entry or instruction even if the original interface reported an error.
Keep controls around the automation itself
Automated rules need change control, access management, monitoring, and an operational owner. A correct rule can become inappropriate when the business introduces a new entity, currency, product, or contractual arrangement. Review the conditions that invalidate the original design, rather than waiting for a calendar review to reveal them.
The 2014 GAO Green Book provides a historical design reference for considering automated and manual control activities and the information systems supporting them. It is a US federal framework, not a current universal requirement for businesses. GAO-14-704G, Principles 10 and 11
Test the fallback as a real operating process. If automation stops, who can access the necessary evidence, what work can continue, and which actions must pause? Maintain enough human capability to investigate the process, but do not assume staff can instantly absorb an unlimited failed queue. Capacity limits belong in the continuity plan.
There are sensible reasons to leave a task manual. Low frequency can make automation maintenance uneconomic. Rules may change faster than they can be validated. A decision may depend on facts that cannot be captured reliably. In those cases, structured preparation and better evidence can still improve the work without delegating execution.
Before production approval, ask the proposed service owner to demonstrate a deliberately broken case: incomplete evidence, a duplicate input, or a changed rule. The owner should show how the action stops, how the exception is assigned, and how the correct state is restored. This exercise makes the boundary observable. It also reveals whether the fallback depends on undocumented knowledge held by the implementation team.
Choose one recurring finance action and write its permitted boundary, abstention conditions, owner, and recovery steps. Test those conditions before expanding the boundary. The right automation plan progressively earns authority through evidence and gives people the information and capacity to handle the cases that still require them.