An ERP system earns trust by preserving transactions, enforcing rules and making the business’s state explainable. Adding intelligence changes its role: the system may begin predicting a delay, recommending a purchase or identifying an unusual transaction before a person notices it. The opportunity is meaningful, but the transition introduces a new question. Under what conditions should the organization act on a prediction?
The most useful future is not defined by the number of AI features on a product roadmap. It is defined by better operational decisions with clear evidence, appropriate authority and a reliable way to recover when the recommendation is wrong. The underlying record remains essential because it supplies both the context for a decision and the evidence of what actually happened.
Leaders should evaluate intelligent capabilities as changes to decision-making, not merely as additions to the interface. That framing reveals the data, controls and operating work required to turn a promising model into a dependable business capability.
A prediction estimates an outcome, such as the likelihood of a supplier delivery arriving late. A recommendation proposes a response, such as ordering earlier or using an alternative supplier. An action changes the business, for example by releasing a purchase order.
These are different responsibilities. An accurate delay prediction does not establish which alternative is commercially sensible. A recommendation may ignore a contract, minimum order quantity or quality constraint. An action can create a commitment that is costly to reverse.
The architecture should preserve those boundaries. Record the prediction and relevant uncertainty, apply business constraints to generate options, and route the proposed action through the appropriate authority. Some low-consequence actions may eventually be automated within tight limits. Others should remain recommendations reviewed by a qualified person.
This separation also improves evaluation. If outcomes disappoint, the organization can investigate whether the forecast was poor, the response policy was unsuitable or execution failed. Combining all three into an opaque “AI decision” makes diagnosis much harder.
A useful candidate has a recurring decision, enough relevant information, an observable outcome and a feasible response. Detecting a problem creates little value if nobody can act in time. Predicting a delay after the last economical intervention point may improve awareness without improving performance.
Define the decision owner, frequency, current method and cost of errors. Compare the proposed capability with a simple baseline such as a fixed threshold, recent average or existing planner rule. Complexity needs to earn its place through better outcomes or a materially better trade-off.
Choose the evaluation measure from the decision. A forecast error statistic can be informative, but a planner may care more about avoiding stockouts without creating excessive inventory. A fraud or anomaly alert may look accurate while overwhelming reviewers with low-value cases. The operating capacity to use the output belongs in the design.
The January 2023 NIST AI Risk Management Framework provides voluntary guidance for managing AI risks across design, development, deployment and use. Its emphasis on context and lifecycle responsibility supports treating an intelligent ERP feature as an ongoing decision capability rather than a one-time technical installation.
Consider a fictional manufacturer that wants to anticipate late deliveries of packaging materials. Its existing rule flags purchase orders only after the promised date passes. The proposed model would identify risk earlier using order history, supplier performance and current transaction information. No performance results are asserted in this example.
The planning team selects a limited set of noncritical packaging items for a pilot. It first checks whether historical promised dates reflect the information available at the time or were overwritten after deliveries were rescheduled. Training on the final corrected dates could make retrospective performance look better than a live decision would be.
The team creates a time-based evaluation: earlier records support development, later records test performance, and inputs are restricted to information that would have been available when each prediction was made. It compares the model with the existing rule and a simple supplier-specific baseline. The purpose is to discover incremental value, not to prove that a sophisticated method must win.
Next, the team defines possible responses. A planner may confirm the supplier’s status, adjust an internal schedule or propose an alternative source. Supplier qualification, contractual conditions and available capacity constrain the options. The model cannot approve a new supplier or create a purchase commitment merely because it predicts a delay.
During an initial live period, the model operates in shadow mode: it records recommendations without changing orders. Reviewers compare the alerts with actual outcomes and record whether an intervention would have been possible. The team also checks whether certain suppliers or item types generate disproportionate false alarms.
If the evidence is favorable, planners begin reviewing recommendations within the ordinary approval process. Each recommendation displays the affected order, relevant source information, the suggested response and known limitations. The decision record captures what the planner chose and why, including overrides.
A later move to limited automation would require a separate decision. It might permit only a narrow, reversible internal scheduling adjustment, with defined limits and monitoring. It would not imply blanket authority to buy materials or change supplier terms. The pilot’s progression is governed by evidence and consequence rather than by a predetermined ambition to remove people.
ERP history reflects how the organization recorded work, which may differ from what physically happened. Missing timestamps, overwritten statuses and inconsistent exception codes can distort training and evaluation. A model can learn those distortions with apparent precision.
Document which fields are authoritative, when they become available and how they are corrected. Preserve the link between the original business event and the features used for prediction. Access should remain limited to the data needed for the approved purpose; intelligence is not a reason to expose every enterprise record to every user or service.
The Government Data Quality Framework emphasizes fitness for purpose and addressing quality at source. For predictive operations, that means judging whether historical data supports the particular decision being modeled, not simply whether the warehouse is populated.
Also examine feedback effects. If planners intervene on high-risk orders, those orders may arrive on time. A simplistic evaluation could then label the original alert incorrect. Record interventions so outcomes can be interpreted in context. Similarly, training only on previously approved actions can reproduce the limits of earlier decisions rather than reveal better alternatives.
A human approval button does not guarantee effective oversight. Reviewers need enough time, context and authority to challenge a recommendation. If the system produces hundreds of alerts and the operating plan provides a few minutes of attention, formal review can become habitual acceptance.
Design the review around the consequence of error. Show the underlying transaction, relevant constraints, source freshness and the reason the case needs attention. Avoid presenting a model score as a universal probability unless it has been appropriately calibrated and explained for the use case.
Give reviewers a meaningful alternative: accept, modify, defer, request more evidence or reject. Capture reasons in a usable form without turning every decision into a lengthy administrative task. Review patterns can reveal missing business constraints, data defects or a model that is poorly matched to actual work.
Do not assume human judgment is infallible. Compare the combined process with realistic alternatives, including the current method and simpler automation. The objective is better decisions under operational conditions, not merely the presence of a person in the workflow.
Performance can change when suppliers, products, demand or business policies change. Monitor both technical measures and business outcomes. Investigate shifts in input distributions, missing data, alert volumes, override patterns and the cost of errors.
NIST’s AI RMF 1.0 calls for testing before deployment and regularly during operation, with documented measures and attention to human-AI configurations. Its guidance does not provide a universal acceptable error rate. The organization must set thresholds that reflect its own context and risk tolerance.
Define an operational stop condition. A stale input feed, unexplained change in recommendations or a breached business constraint may require suspending the capability. Maintain a workable fallback, such as the prior planning rule or a manual queue with adequate staffing. A theoretical rollback is insufficient if the old process has been dismantled.
Changes to models, prompts, data pipelines or decision rules should be versioned and tested. Retain enough information to reconstruct consequential recommendations without storing unnecessary sensitive data. Assign responsibility for incident response and for deciding when the system is safe to resume.
A compact working heuristic can make the investment review more concrete. For each proposed capability, document the decision, accountable owner, allowed inputs, baseline method, success measures, prohibited actions, review requirements, monitoring and fallback.
Add the boundaries of authority. What may the system recommend? What may it execute? Which conditions force escalation? How does the organization detect that an action was completed, rejected or left in an uncertain state? A recommendation engine connected to transactional APIs needs the same discipline around permissions, duplicate prevention and reconciliation as other operational software.
Evaluate the whole cost: data preparation, integration, review effort, model operation, monitoring and continuing validation. A useful prediction may still be uneconomical if acting on it requires more effort than the benefit it creates. In stable, transparent processes, a well-designed rule may remain the better choice.
The future of ERP will be shaped by how responsibly organizations connect records to judgment and judgment to action. Intelligence should make decisions more timely and effective while preserving accountability. The strongest foundation remains an enterprise that knows what its data means, who can authorize change and how to learn from the results.