Recovering from an ERP Integration Failure. Check what happened before retrying and reconcile the result.
The most dangerous integration failure can be the one whose outcome is unknown. A sender times out, but the receiving system may already have created the order, posted the update, or initiated the next activity. Repeating the request without establishing what happened can turn a communication problem into a duplicate business action.
Recovery design should therefore begin before production use. Define how the team distinguishes a known failure from an uncertain result, which actions are safe for each state, who can authorize correction, and what evidence proves the business is consistent again.
The practical deliverables are a failure-mode matrix and a controlled recovery runbook. They should be specific enough for an on-call team to use under pressure, with business owners involved in the decisions that technical monitoring cannot make alone.
Choose an important interface and identify what would happen if its information were missing, repeated, delayed, out of order, or partly applied. A missing status update and a duplicated instruction may have very different consequences even when both appear as message errors.
Build a matrix containing the symptom, detection method, affected business process, safe initial action, recovery owner, required authority, and closure evidence. Start with these failure modes:
Add architecture-specific cases with the integration team. The list is a prompt for analysis, not an exhaustive technical model. For every row, identify how the business detects the problem if technical monitoring reports no obvious error.
A technically healthy connection can still carry stale or incomplete information. Monitoring should therefore consider expected business events and aging unresolved work, not only connectivity and response codes.
Use states with explicit evidence requirements. “Pending” means the work is known to be waiting or in progress. “Failed” should state what is known to have failed and whether any business effects occurred. “Unknown” means the evidence is insufficient to determine the outcome.
“Accepted” also needs a precise definition. It may mean that a message passed an initial checkpoint, not that every business effect completed. Where the process has later stages, record their completion evidence separately. A generic success label should not hide an outstanding dependency.
Assign ownership to state transitions. Who detects overdue pending work? Who investigates unknown outcomes? Who can declare that a failed request had no effect? What evidence changes the state? These questions make the monitoring view actionable.
An unknown result should enter an investigation path. Depending on the design, that may involve a supported status lookup, a receiving-system query, or reconciliation using business identifiers. Do not treat the absence of a response as proof that the action did not occur.
The recovery team needs a reliable way to connect the original business event, transmission attempts, receiving records, and later corrections. Document which identifiers provide that connection and how long the required evidence remains available.
Where the architecture supports an idempotent operation, repeated execution under its defined conditions has the same intended effect as one execution. Verify the actual conditions: the identity used, whether retries reuse it, the supported time window, and what happens if the same identity arrives with different content. A label in a design document is insufficient evidence.
Duplicate-prevention mechanisms must be evaluated across the relevant processing boundary. Detecting a repeated message at one component does not automatically prevent another downstream effect from repeating. Technical owners should test the implemented behavior, including concurrent attempts and interruption at important points.
Keep event identity distinct from a legitimate new request. Two orders with similar details may be intentional. Conversely, a retry can have a new transport identifier while representing the same original business action. Business meaning should guide the identification design, with technical mechanisms selected to enforce it.
Imagine an ERP process requesting a warehouse picking task. In this hypothetical scenario, the receiving system creates the task, but the acknowledgment is lost before the sender receives it. The sender displays a timeout.
The original support procedure says to resend failed requests. If an operator follows it without checking the outcome, a second picking task may be created. The communication fault has become an operational risk, even though no warehouse worker has yet acted on either task.
A revised runbook classifies the timeout as an unknown outcome. The operator uses the approved lookup route with the original event identifier. The receiving system shows that the task exists, so the operator records the verified task reference and resolves the sender's status through the supported procedure. The request is not blindly repeated.
A second test returns no conclusive result. Perhaps the receiving record is not yet searchable or the lookup service is unavailable. The operator keeps the outcome unknown and escalates according to the business deadline. The absence of immediate evidence does not authorize a duplicate-risk action.
If the architecture provides a tested replay mechanism that safely recognizes the original request, the runbook can permit that mechanism within its verified conditions. If it does not, the team must establish the prior outcome or use an explicitly approved recovery alternative. The example describes a design decision, not a capability every system provides.
Before closing the incident, the warehouse owner confirms the final task population and checks that no duplicate work entered execution. The integration owner reconciles the original request with the receiving outcome and documents any correction. Restoring communication alone would not establish either fact.
The runbook should identify the interface, business owner, technical owner, supported operating window, access requirements, and evidence locations. Then give the operator a bounded sequence:
The exact sequence depends on the process. Limiting processing can itself disrupt operations, so the runbook must specify who can make that decision and under which conditions. Avoid leaving an operator to choose between uncontrolled duplication and an unauthorized shutdown.
For a replay, record the selected population, original identifiers, reason, proof of safety, approver where required, execution result, and reconciliation evidence. Prevent the same incident from being recovered independently by two teams without coordination.
Use supported application procedures for corrections. Direct changes to underlying records can bypass controls or leave related data inconsistent. Any exceptional repair route requires appropriate technical review, business authorization, and verification for the actual system.
A delayed event may no longer be appropriate to apply in its original form. An order could have been cancelled while its earlier release instruction was in transit. The recovery procedure needs a rule for checking current business state before applying old work.
Out-of-order events require an agreed interpretation. The latest arrival is not necessarily the latest business event. Technical approaches vary, but the business must define which state transitions are valid and how conflicts are resolved. Do not allow a transport timestamp to make an unreviewed operating decision.
For partial processing, identify the smallest unit that can be safely recovered. Replaying an entire batch may repeat successful rows. Replaying only visibly failed rows may miss records with unknown outcomes. The procedure needs enough evidence to distinguish those populations or an architecture-tested method that safely handles them together.
Also examine multi-step business effects. A record may have been created while a related notification, allocation, or downstream transfer failed. Recovery should establish the intended complete outcome, including whether an adjustment or compensating business action is required. Such actions need their own authority and evidence; they are not automatically equivalent to undoing a technical call.
Automatic retries can be appropriate for some transient failures when the operation is safe to repeat under the verified design. Define which failures qualify, the retry limits, spacing, and escalation condition. Technical owners should validate the mechanism against the receiving service's behavior and capacity.
Repeated retries can consume resources while leaving the business deadline unchanged. The operating team needs to know when an unresolved transaction requires a different response, such as a controlled manual arrangement or a decision to hold dependent work. A retry counter alone does not express that consequence.
Keep credentials and unnecessary sensitive content out of incident logs and replay records. Preserve the evidence needed to investigate while following the organization's access and retention controls. The recovery process should not create a new information exposure while resolving the original failure.
In a controlled test environment, rehearse lost acknowledgments, repeated submissions, late events, partial batches, and unavailable lookup evidence. Include support staff and business owners who will respond in production. Observe where the runbook depends on knowledge that has not been written down.
The acceptance test should continue until the affected business population is reconciled, unresolved exceptions have owners, and the business agrees that normal processing can resume. Record what evidence was necessary and update the runbook accordingly.
For the next interface review, ask one question: “If the connection breaks immediately after the receiving system acts, how will we know whether it happened?” A tested answer provides a foundation for safe recovery. If the answer is unclear, resolve that design gap before relying on routine retries to keep the process running.