Accounting leaders should redesign work around verifiable evidence, review capacity, and professional development before changing headcount assumptions around AI. A system that produces a plausible draft can reduce preparation effort while increasing the need to inspect sources, challenge interpretations, and maintain the process. The team changes because the distribution of work changes, not because every task suddenly becomes autonomous.
The immediate decision for a controller is which responsibilities to create, strengthen, or preserve during a bounded AI pilot. That includes accountability for approved use cases, evaluation cases, source access, review, exception handling, and staff learning. These responsibilities need actual time and authority, rather than being added invisibly to an already full close calendar.
This is an operating-model argument, not a prediction that a fixed percentage of accounting jobs will disappear. Outcomes depend on the task, information quality, system behavior, implementation costs, and the organization’s ability to use any released capacity well.
Decompose the role before redesigning it
An accountant’s role contains several kinds of work: gathering evidence, applying an established procedure, investigating an exception, making or supporting a judgment, documenting a conclusion, and communicating implications. An AI tool may assist some of these tasks without being suitable for others.
Take variance commentary. Gathering approved source figures and drafting an initial explanation can be separated from establishing why the variance occurred. A text that repeats a number accurately may still invent a causal explanation. The reviewer must distinguish a supported observation from a hypothesis and a confirmed cause.
Map the current task with the people who perform it. Record what they inspect, which exceptions they recognize, and how they know a conclusion is adequately supported. This reveals tacit expertise that can disappear from a process description written only by project managers.
For each proposed AI-assisted task, define the output and the review burden. “Create monthly commentary” is too broad. “Draft a summary of approved variances, linking each factual statement to a specified source and leaving causal questions unresolved unless documented” gives the team something it can evaluate.
Create responsibility for evaluation
A finance team using AI needs someone accountable for deciding whether the system is useful and safe enough for its intended task. This may be a responsibility within an existing role rather than a new job title, but it should not be left entirely to the vendor or the employee most enthusiastic about the tool.
Maintain a set of representative cases with independently reviewed expected outcomes. Include ambiguous language, changed classifications, missing evidence, unusual transactions, and misleading but plausible explanations. Keep some cases separate from the examples used to improve prompts or configuration, so evaluation tests more than familiarity with the training exercise.
Measure the kinds of errors that matter to the workflow. For commentary, distinguish unsupported causal statements, wrong figures, omitted significant issues, and misleading certainty. For document extraction, distinguish a minor formatting problem from a wrong amount or counterparty. A single overall accuracy percentage can hide an unacceptable error class.
NIST’s January 2023 AI Risk Management Framework describes evaluation and monitoring within the MEASURE function, together with organizational responsibility under GOVERN. It is voluntary risk-management guidance, not an assurance opinion on a particular accounting use case. NIST AI RMF 1.0, Sections 5.1 and 5.3
Protect the review function from becoming a bottleneck
AI-assisted preparation can deliver more drafts to reviewers than they can examine responsibly. If the project counts the preparer’s saved minutes but ignores reviewer effort, it may relocate rather than remove the constraint.
Design the review interface around evidence. Show the source supporting a claim, the relevant period and version, any transformation, and the points requiring judgment. A reviewer should be able to reject or revise a proposal without losing the link to its source. The process should record significant overrides and recurring error patterns.
Allocate review according to competence and consequence. A new employee can verify that a number agrees to an approved table, but may not be qualified to assess a novel accounting conclusion. Review responsibilities should be explicit enough that an approval means something more than “someone looked at it.”
Watch for overreliance. A fluent explanation can feel complete even when it is unsupported. Periodically compare reviews of AI-assisted and independently prepared cases, and ask reviewers to explain what evidence would falsify the proposed conclusion. The purpose is to test the review method, not to shame people for trusting a convenient output.
A hypothetical commentary pilot
Suppose a team prepares twenty-four monthly commentary packs, each requiring two hours under its existing process. The baseline is forty-eight hours. In a hypothetical AI-assisted design, gathering approved inputs and producing the draft takes fifteen minutes per pack, or six hours. Review and correction take forty-five minutes per pack, or eighteen hours. Evaluation, maintenance, and issue analysis add eight hours. Total effort is thirty-two hours, releasing sixteen hours in the planning model.
This is an illustrative business case, not a performance claim. It assumes the output meets the agreed evidence and quality criteria. If review instead takes seventy-five minutes per pack, review consumes thirty hours, and total effort becomes forty-four hours. The modeled release falls to four hours.
The team should measure these components separately during the pilot. Counting only the six hours of preparation would misrepresent both the cost and the changed skill requirement. A rejected draft also consumes time, so include failures and difficult cases in the evaluation population.
Assume, for planning purposes, that sixteen hours are genuinely released. The controller might allocate eight to investigating recurring unexplained variances, four to staff development, and four to extending the evaluation cases. Those are proposed uses, not automatically realized benefits. The pilot should establish whether that work actually occurs and whether it produces useful results.
A headcount decision based on the initial preparation estimate would be premature. The organization first needs evidence of durable capacity, coverage during absences, the ongoing control burden, and the effect on the remaining work. Any employment decision also requires the organization’s appropriate human and legal processes.
Hypothetical commentary pilot with quality criteria held constant; preparation time alone is not the operating cost.
Open full-size diagram
Preserve the path by which accountants learn
Junior accountants often develop judgment through preparing work, encountering exceptions, and receiving review feedback. If automation removes that experience, the organization needs another deliberate way to build the same competence. Otherwise it can create a future shortage of people able to challenge the system.
Use supervised reconstruction exercises. Give staff a source population and ask them to derive the result independently before comparing it with an AI-assisted output. Discuss why a plausible answer is wrong, which facts matter, and how uncertainty should be expressed. Retain opportunities to investigate exceptions end to end.
Rotate staff through source-data and process work as well as reporting. Understanding why an operational event is incomplete or how a mapping changes a financial view is valuable preparation for reviewing automated outputs. Training should combine accounting knowledge, data interpretation, and practical system controls rather than treating prompt writing as the entire skill set.
Keep professional accountability clear. A generated conclusion should not be attributed to a person who has not performed the required review. The organization must decide who is qualified to approve each class of output and how that person can obtain the underlying evidence.
AI-assisted work still needs a deliberate route for people to acquire the expertise required to review it.
Open full-size diagram
Give the process an owner after the pilot
A pilot often has concentrated attention from finance, technology, and a vendor. Normal operations need named ownership of source permissions, configuration changes, evaluation, incidents, and the decision to suspend use. The process owner should know which changes require retesting, including new source formats, model changes, and expanded use cases.
Control access to financial and personal information according to the organization’s policies and applicable obligations. A useful tool does not justify sending all available records to it. Define the minimum inputs, approved destinations, retention expectations, and access rights before deployment, with appropriate security and privacy review.
Maintain a workable fallback. If the tool becomes unavailable or fails evaluation, the team should know which tasks can revert to a manual method and which deadlines need adjustment. Preserve the procedures and expertise needed for that fallback. A plan that assumes everyone can immediately absorb a large backlog is not a tested operating model.
The NIST framework also treats operator proficiency and human oversight as explicit considerations. For accounting leaders, that supports budgeting for competent supervision as part of the system, rather than assuming a person in the workflow automatically makes the result reliable. NIST AI RMF 1.0, MAP 3.4 and MAP 3.5
Build the next team from the work that remains
The useful future state combines dependable preparation, informed review, better investigation, and a sustainable learning path. It may involve different role proportions, but those proportions should follow observed work. Some tasks will remain manual because they are infrequent, ambiguous, or costly to evaluate reliably.
Choose one recurring deliverable and run a bounded pilot with explicit quality criteria, reviewer capacity, learning objectives, and operating ownership. At the end, decide what work has actually changed and what responsibility it creates. An accounting team becomes more capable when AI reduces avoidable effort while strengthening its ability to explain and challenge financial information.