Systems integration
Cross-system workflows need a recovery decision at every step
A workflow that updates several business systems needs a recovery plan for partial completion. Record what each step actually changed, decide which changes can safely be reversed, and route irreversible or uncertain outcomes to a named owner. Restarting the entire process is rarely a complete recovery policy.
The practical issue
A workflow that updates several business systems needs a recovery plan for partial completion. Record what each step actually changed, decide which changes can safely be reversed, and route irreversible or uncertain outcomes to a named owner. Restarting the entire process is rarely a complete recovery policy.
Consider an illustrative equipment-rental workflow. A request reserves equipment, creates a delivery task and sends a confirmation. If delivery scheduling fails after the reservation succeeds, the systems now disagree. The customer may have no confirmed delivery, while the equipment remains unavailable to other customers. This is a design scenario, not a Quarro client result.
Why the usual retry button leaves a gap
A simple integration often treats the chain as one operation: run the steps, show success, or show an error. That is easy to understand until a response is lost or a later step fails. An error does not necessarily mean that nothing changed. A reservation may exist even when the requesting system never received its identifier.
A blanket retry can create another delivery task. A blanket rollback can remove a valid reservation that someone has already adjusted. The practical gap is knowledge of the business state after each step, including changes made by people while recovery is pending.
Borrow the pattern, keep the business decisions explicit
AWS describes saga orchestration as coordinated local transactions across services, with compensating transactions for failures. Its guidance also identifies maintenance complexity, eventual consistency and lack of transaction isolation as considerations. The pattern supplies a useful structure; it does not decide what a business should undo.
For the rental example, a recovery record could contain the request identifier, equipment reservation identifier, delivery-task identifier, last confirmed step and current recovery owner. Each external write should have its own recorded result. A timeout should produce an uncertain state until the destination can be checked, rather than being silently treated as a confirmed failure.
Describe the actions in operational language: release an unused reservation, cancel a draft delivery task, or ask dispatch to review. These descriptions help the person responsible for the workflow assess whether the proposed recovery is still safe.
Mark the point where automatic reversal stops
Some actions can be reversed cleanly. Others create commitments or real-world effects. A sent customer message cannot be unsent by deleting a database row. Equipment already loaded onto a truck cannot be made available simply by resetting a status.
Write an explicit boundary for each step. Before dispatch accepts the job, automatic cancellation might be permitted under the company's policy. After acceptance, recovery may require a coordinator. Preserve the original action and any correction in the history so the team can explain what happened.
Also check for intervening edits. If a dispatcher changed the delivery time, an automated compensation must not blindly restore an older version. Recovery should compare the current record with the version it expects and stop for review when the business situation has changed.
Give incomplete recovery a visible home
The recovery screen should answer four questions: what completed, what remains uncertain, what action is proposed, and who can authorize it? Provide a link to the relevant records and a plain-language reason for the stop. Restrict recovery controls according to their consequences.
Test failures after every external step, including a lost response, an unavailable destination and a compensation that fails. Check that a second recovery attempt resumes from the recorded state. A successful test should leave one explainable business outcome, not merely a green technical status.
A practical starting point
Choose one multi-system process and draw its commitments in order. For each, write the evidence of completion, the permitted correction and the human escalation point. If those answers are unclear, resolve them with operations before adding another connector.
Quarro's integration work can connect these business decisions to the software: durable step records, safe recovery controls and an exception view people can use. For the related mechanics of repeated jobs, read [Queue-Backed Workflows Need Retry Ownership](/blog/queue-backed-workflows-need-retry-ownership/).