Workflow improvement

Batch jobs need restart checkpoints before retry buttons

October 11, 2026 · Workflow improvement

A retry button is only safe when the job knows what already happened. Batch work needs checkpoints that persist progress, define which steps can rerun and explain what operators should do when a restart fails again.

Original batch job recovery diagram showing persisted checkpoints, repeatable steps, protected side effects and operator restart evidence.
Original explanatory illustration created for Quarro, October 11, 2026. · Original PNG diagram authored for this article. No third-party photos, logos, screenshots, fonts, or stock assets embedded.

The answer: checkpoint before you retry

A retry button is only safe when the job knows what already happened. Batch work needs checkpoints that persist progress, define which steps can rerun and explain what operators should do when a restart fails again. Otherwise retry may duplicate records, skip work or hide the original failure.

Imagine an illustrative nightly billing preparation job. It reads accounts, applies eligibility rules, writes draft charges and sends a summary. If it fails after writing half the draft charges, the next run cannot simply start from the top unless writes are idempotent or the job can identify and resume from a trustworthy checkpoint.

Define the unit of progress

A checkpoint is not just a timestamp. It is a recorded unit of completed work: file offset, source record key, batch number, step status, output marker or transaction boundary. The right unit depends on the process. Reading a file, posting to an external API and updating a database table may each need a different progress marker.

Persist checkpoint data somewhere the restarted job can read after a crash. Keeping progress only in memory helps a running process report status, but it does not support recovery. The checkpoint also needs to mean something operationally: staff should be able to tell whether the job is safe to restart, requires cleanup or needs engineering review.

Restart rules are business rules

Spring Batch documentation describes restart behavior where completed steps are normally skipped, while some steps can be configured to run again. It also notes that stateful writers need persisted state to pick up correctly after restart. Those framework details point to a broader principle: a restart is a designed path through the workflow, not a button that repeats everything.

Document which steps are repeatable, which are one-time, which require compensation and which require manual review. Validation may be safe to rerun. Sending customer notifications may not be. Loading a new file may be safe only if the file identity and row identities are preserved.

Give operators evidence, not mystery controls

Before enabling retry, show the job instance, input version, last completed step, checkpoint position, error summary, affected record count and recommended action. If the operator sees only Failed and Retry, they are being asked to make a decision without the evidence the system already has.

Use start limits for repeated failure. After a job fails the same step more than the allowed number of times, the system should stop automatic retries and require investigation. This prevents a bad input or dependency outage from being hammered repeatedly while downstream effects accumulate.

Test restart as a feature

Create authorized fixtures that fail after each major step. Verify that restart resumes exactly where intended, does not duplicate side effects and leaves an audit trail that explains the run. Test partial writes, downstream timeouts, repeated failures and a successful restart after correction. Keep real customer data out of fixtures.

The practical takeaway: retry is a user interface for recovery. Recovery is the system design underneath it. Build the checkpoints first, and the retry button becomes a controlled operation instead of a gamble.

Sources