CURIOUSRUBIK
Let’s talk about your next move ↗View complete sitemap
Back to the blog

Integration Failures: How to Design for Recovery, Not Just Success

A recovery design should answer what happened before deciding what to retry. After a timeout or crash, some operations may be complete, some may be untouched, and others may have an uncertain outcome. Treating all failed messages as unfinished work can duplicate external actions or overwrite a newer business state.

For an operations and engineering lead responsible for a shipping integration, the immediate decision is how to classify interrupted work and restore the intended outcome safely. The recovery plan needs evidence, bounded authority, replay rules, and a clear endpoint. Restarting the service is only one technical step; it does not establish that the shipping process has recovered.

Design these capabilities while the system is being built. It is much harder to reconstruct operation identifiers, source versions, and external references after an incident if the normal flow never preserved them.

Define the states an operator must distinguish

Record enough durable information to distinguish work not yet attempted, work known to be applied, work explicitly rejected, and work with an uncertain outcome. These are suggested operational categories. The exact state machine should reflect the provider’s contract and the business effect.

An uncertain outcome is particularly important when an external service performs an action but the response is lost. The caller’s timeout says it did not receive a timely result. It does not establish that the provider did nothing.

Preserve a stable operation identifier before the external attempt where possible. Carry it through retries and retain any provider reference returned. Define whether the provider supports duplicate-safe submission or lookup by that identifier. If neither is available, the design may require an investigation route rather than automatic resubmission.

HTTP semantics explicitly limit automatic retries of non-idempotent requests unless the client knows retrying is safe or knows the original request was not applied. That constraint is directly relevant to recovery after ambiguous responses. RFC 9110, Section 9.2.2

A hypothetical interrupted shipping batch

Suppose a hypothetical shipping process has one thousand authorized shipment-creation operations when an incident stops progress. The recovery record shows 640 confirmed by the carrier, 200 not yet attempted, 100 with uncertain outcomes after lost responses, and 60 explicitly rejected for invalid address data. These categories are mutually exclusive and total one thousand.

The 640 confirmed operations should be reconciled with the expected shipment references, not submitted again simply because the local batch did not finish. The 200 unattempted operations can enter the normal authorized processing path once dependencies are healthy and the underlying shipment request remains valid.

The 100 uncertain operations require a lookup or other supported confirmation method. If the carrier confirms that a shipment exists, record its reference and continue from that state. If reliable evidence establishes nonapplication, retry under the agreed rules. If the outcome remains unknown, hold the item for an accountable investigation rather than choosing between duplicate creation and silent loss by guesswork.

The 60 invalid-address operations need source correction through the appropriate owner. Repeating the same invalid request at increasing frequency will not repair its meaning. After correction, preserve the relationship between the original operation, the revised data, and the authorized new attempt.

Some shipping requests may have changed during the outage. A customer may have canceled, an order may have been placed on hold, or a different dispatch may already have been arranged. Revalidate relevant business preconditions before replay. A durable old instruction is evidence of prior intent, not necessarily permission to execute it indefinitely.

The example is hypothetical and does not assume any particular carrier API supports the required lookup or idempotency behavior. Those capabilities must be verified before selecting an automatic recovery path.

Hypothetical mutually exclusive recovery populations: of 1,000 interrupted operations, 640 are confirmed and need reference reconciliation; 200 are unattempted and need revalidation before processing; 100 are uncertain and require lookup or investigation; 60 are invalid and need source correction. There is no direct uncertain-to-resend path. No blind resend. Confirm intent, provider behavior and existing authorization.
Hypothetical mutually exclusive recovery populations. Automatic actions depend on provider capabilities, current intent and existing authorization.
Open full-size diagram

Build a recovery record that survives the incident

For consequential operations, retain the source identifier and version, operation identifier, attempted destination, payload or reproducible evidence reference, attempt times, outcome, and provider reference where available. Protect the record according to its sensitivity. Broad-access logs should not become an uncontrolled copy of personal or confidential shipping information.

Record transitions durably enough that a restarted worker can determine its next step. Where a local state update and acknowledgment can be committed atomically, use that capability appropriately. Where they cannot, identify the gap and the evidence needed to resolve it.

The Idempotent Receiver pattern describes designing repeated messages so their repeated processing does not cause an additional unwanted effect. It is a useful goal, but the actual mechanism depends on the operation and storage boundary. Enterprise Integration Patterns, Idempotent Receiver

A local “already processed” flag cannot by itself establish what an external carrier did. Likewise, an in-memory cache of operation IDs does not survive every restart or support an arbitrary replay horizon. Match the persistence and retention design to the failures and recovery window the organization intends to support.

Separate transient faults from permanent or business failures

A temporary connection interruption may justify another attempt after a delay. Invalid input, missing authority, an unsupported version, or a business rejection generally needs a different response. Classification should follow the provider’s documented behavior and the actual operation, rather than a broad rule that all errors are retryable.

Bound retries by attempts, elapsed time, and business validity as appropriate. Introduce backoff and variation where repeated simultaneous attempts could overload a recovering dependency. Respect provider guidance and limits. A retry loop should eventually produce an actionable outcome for an owner rather than run forever without visibility.

Avoid nested retry multiplication. If a client, integration worker, gateway, and upstream scheduler each retry independently, one original operation can generate many attempts. Assign retry responsibility at explicit boundaries and ensure every layer preserves the operation identity where required.

Keep rejected work in an inspectable quarantine with a reason and owner. A dead-letter queue is storage, not a resolution process. The business obligation remains until it is corrected, canceled through an authorized route, confirmed already complete, or otherwise given a documented disposition.

Replay within a defined boundary

A replay plan should state which population is included, which source and rule versions apply, which side effects are permitted, and how completion will be verified. Replaying a historical event stream to rebuild an analytical projection differs from replaying it through a consumer that sends notifications or creates shipments.

Use a dry-run or comparison mode where feasible. Identify the intended destination effects and compare them with current state before executing consequential changes. Start with a bounded cohort and verify results before increasing the replay volume.

Do not assume that restoring a database backup restores the whole business process. External systems may retain actions performed after the backup point. A restored local database can therefore forget an action that still exists elsewhere. Reconciliation with external state is essential before resumed processing creates it again.

Compensation is also a business action, not a magical rollback. Canceling an unused shipping label may be possible under the provider’s rules; retrieving a dispatched parcel may not be. The organization must define which corrective actions are authorized and which require human approval or a customer-facing decision.

Reserve capacity for catching up

Recovery requires capacity beyond normal arrivals. Suppose, in a separate hypothetical queue example, new work arrives at forty items per minute and the consumer can sustainably process seventy. If a backlog contains three hundred eligible items and rates remain constant, the net drain rate is thirty per minute. Clearing that backlog takes ten minutes under those assumptions.

If processing capacity is only forty per minute, the backlog does not shrink while new work continues. A design sized only for normal throughput can remain permanently behind after an interruption.

The simple calculation omits variable item complexity, rate limits, shared dependencies, retries, and investigative holds. Measure these during testing. Do not include uncertain or invalid operations in an automatically replayable population merely to make the recovery forecast look faster.

Protect the recovering system from a surge. Control concurrency, prioritize work according to business consequence and approved rules, and observe both new and old populations. Catching up quickly is not useful if it creates another outage or starves time-sensitive current work.

Hypothetical constant rates: 40 new items arrive each minute and sustainable processing handles 70, leaving a net drain of 30 per minute. A 300-item eligible backlog therefore takes an idealized ten minutes to clear. If processing only equals arrivals at 40 per minute, there is no net drain. Investigative holds, rate limits and variable complexity are excluded.
Hypothetical constant-rate arithmetic excludes investigative holds, rate limits and variable complexity.
Open full-size diagram

Rehearse the stopping condition

An incident is not resolved merely because processes are running and queues are smaller. Define the evidence that ends recovery: expected operations have confirmed outcomes, unresolved items have accountable dispositions, destination state reconciles, and normal objectives are being met.

Run exercises that interrupt processing before an external action, after the action but before acknowledgment, and during local state recording. Test restoration from a checkpoint and a bounded replay. Include the operations team, because the runbook must be executable by the people who will use it under pressure.

Afterward, repair the design where the exercise required guesswork. Add missing identifiers, clearer provider contracts, safer duplicate handling, or better reconciliation. Do not rely solely on writing a longer runbook around an avoidable information gap.

Choose one consequential integration and ask its owner to demonstrate recovery from an ambiguous timeout. If the team cannot establish whether the external effect occurred, improving that evidence path should precede broader automation. Recovery becomes dependable when the system preserves enough truth to resume the business process without repeating or abandoning its obligations.

Further Reading

What’s on your mind?

A little context is all it takes to begin.

Please leave out passwords, payment details and confidential account data.