When Should AI Make Decisions in Your ERP? Define what AI may suggest, prepare or carry out.
Before an AI-assisted ERP workflow can act, define the decision it is allowed to make, the evidence it must use, and the conditions that require a person to intervene. Keep recommendation, preparation, and execution as separate permissions.
A capable system can still make an unauthorized commitment. An authorized system can still make a poor decision. The pilot must examine both questions: whether the result is good enough for the business purpose and whether the action remains inside the approved boundary.
Start with one bounded process and a named business owner. The aim is to produce an authority record and a testable pilot decision before increasing autonomy.
“Improve purchasing” is too broad to govern. “Suggest replenishment quantities for specified noncritical items using approved demand and stock inputs” is a more useful starting point. It identifies a population, an output, and a basis for review.
Document what the workflow is meant to improve and the consequence of a wrong recommendation. Consider downstream effects as well as the immediate task: excess inventory, service disruption, inappropriate supplier communication, or commitments made on incomplete information.
Identify the accountable process owner, the technical owner, and the control or risk specialists needed for the use case. The business owner decides what outcomes and tradeoffs are acceptable. The technical team establishes how the approved boundary can be enforced and observed.
If the underlying business rule is unresolved, settle that question before asking automation to choose on the organization’s behalf. A model’s fluent answer is not a substitute for a decision the business has avoided making.
Use distinct permissions for distinct activities.
Observe. Read an approved data population and identify relevant facts or exceptions. Access itself needs a justified scope, particularly where information is sensitive.
Recommend. Suggest an action with supporting evidence and a concise explanation. The recommendation does not change the business record or communicate a commitment.
Prepare. Assemble a draft transaction or proposed update for review. Define where the draft is stored and ensure its status cannot be mistaken for an approved action.
Execute within limits. Perform a specifically authorized action after the required checks. Limit the action, population, timing, and exposure, and define how enforcement occurs outside the model’s own narrative instructions.
Escalation is available at every level. It is the route for missing evidence, conflicting instructions, unusual conditions, or an action outside the approved boundary. It should not be presented as a higher level of autonomy.
Permission to recommend does not imply permission to create a transaction. Permission to prepare one does not imply permission to release it. Make those boundaries visible in the workflow, user interface, credentials, and review record.
Create one row for each action the workflow might take:
Limits should reflect the real consequence. A per-transaction limit may be inadequate if many small actions create a large combined exposure. A permitted item range may be inadequate if the same action affects a restricted customer or location. Review how limits interact rather than treating each field independently.
Have the appropriate legal, privacy, security, and control owners determine requirements that apply to the use case. A generic authority matrix cannot establish compliance by itself.
Imagine a distributor piloting replenishment recommendations for a defined set of maintenance supplies. The workflow reads approved stock, open-order, demand, and lead-time information, then proposes quantities for a buyer to review.
At the first stage, it may recommend but cannot create or release purchase orders, change supplier records, alter payment details, or contact suppliers. Those exclusions make the pilot’s operational boundary clear.
A later stage might allow it to prepare draft requisitions after evidence supports that additional permission. The buyer would still review the current business context before approval. The scope expansion should identify what new failure becomes possible when a recommendation turns into a stored draft.
The pilot also needs exceptions. Missing lead-time information, a recent item-status change, conflicting stock records, or unusual demand may require human review rather than a confident numerical answer.
These are hypothetical design choices, not a claim that automated replenishment suits every organization or that the described controls exist in every product.
“Human in the loop” is incomplete unless the person can make a useful decision. Give the reviewer the proposed action, supporting source references, important assumptions, current record state, and reasons the workflow requested review.
The reviewer needs the authority and time to reject, amend, or investigate the proposal. Test how the queue behaves during absence and peak demand. If unresolved cases accumulate, the pilot may need a narrower scope or more capacity.
Define what happens while review is pending. The workflow should not quietly execute because the reviewer has not responded. A missed response is a state to manage according to approved rules, not implied authorization.
Measure review quality as well as speed. Examine whether reviewers identify flawed proposals and whether the information provided supports that judgment. Rapid approval can mean the process is efficient, but it can also mean the reviewer has too little context to challenge the output.
Build a reference set representing the approved use case. Include normal examples, important exceptions, poor or missing data, changed records, and cases where the correct response is to stop or escalate.
For each case, specify the acceptable outcome, prohibited actions, required evidence, and review route. Some tasks allow several reasonable answers; the evaluation should reflect that rather than pretending there is one exact phrase the system must produce.
Also test adversarial content in a controlled environment. A document, supplier message, or free-text field may contain instructions that conflict with the organization’s rules. Such content is evidence to interpret, not a source of permission to broaden access or take actions.
Have specialists test whether the workflow can cross its approved tool, data, or action boundaries. Technical controls should restrict what it can do even when its generated output is mistaken or manipulated. Successful performance on ordinary examples does not establish that these boundaries hold.
Treat a model-generated confidence statement cautiously. It may not be a calibrated estimate of correctness for the business task. If confidence influences routing, validate that use against relevant evidence and retain separate hard limits for authorization.
Retain enough information to reconstruct the business event: the request, relevant source identifiers and versions, proposed action, checks performed, approval record, executed action, result, and any correction.
A concise explanation of the factors supporting a recommendation can help reviewers. It should be evaluated against the source evidence and observed outcome rather than accepted as proof of how the system internally reasoned.
Protect the record itself. Avoid unnecessary personal data, credentials, or secrets in logs. Define access, retention, and permitted reuse with the appropriate owners. More logging is not automatically better if it creates a new information exposure.
Distinguish the system’s intended action from what the external service actually completed. A successful request submission may still leave a business process pending. Reconcile consequential actions to their authoritative result.
Name the person who can pause the workflow, the conditions that should trigger a pause, and the authorized way to do it. Test that the pause prevents further relevant actions and that in-flight work is visible.
Define the response to incorrect actions. Some effects can be reversed through supported procedures. Others need an approved correction, customer communication, or a separate business decision. Turning off the model does not undo a transaction already posted or a message already sent.
For the hypothetical replenishment pilot, a recommendation can be discarded before action. A created draft may need withdrawal. A released order may have consequences that require the buyer and relevant commercial owners to manage. Each authority level therefore needs its own recovery assessment.
The pilot review should examine more than average answer quality. Review consequential errors, unauthorized-action attempts, appropriate and unnecessary escalations, unresolved work, reviewer effort, and the completeness of execution evidence.
Compare results with the approved task definition and a credible baseline where available. Keep the evaluated population and limitations visible. A pilot consisting only of easy, clean cases does not support conclusions about the exceptions excluded from it.
Use a gate record covering data readiness, action scope, reviewer capacity, outcome evidence, monitoring, and recovery. The possible decisions include continuing within the same scope, narrowing the scope, correcting a weakness, stopping the use case, or authorizing a specific expansion.
Expand one meaningful boundary at a time when practical. A new data source, user population, action type, or operating condition can change risk even if the model is unchanged. Reassess after material changes to the model, prompts, tools, policies, or source processes.
The first useful deliverable is a signed-off statement of what the workflow may do and what it must hand back. Once that boundary is enforceable, observable, and tested, the business has a sound basis for deciding how much judgment it is ready to delegate.