Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Spring Batch orphaned STARTED execution: inspect before relaunch

Last updated: 1 Oct 20264 min read
tutorial
IntermediateBy AITrove Editorial

A process killed without recording failure can leave repository metadata in a running state even when no worker is alive.

Download Spring source kit

Do not invent another instance

A scheduler sees no live worker, but the JobRepository may still show STARTED because the process died before recording an ending status. Relaunching with the same identifying parameters can be refused as already running. Adding a timestamp creates new work, bypasses the old execution history and risks repeated writes. The checked two-JVM test deliberately starts from FAILED, not this state.

Fence recovery

First prove that the old worker cannot still write: check process ownership, scheduler lease and target database connections. Inspect the repository's JobExecution and StepExecution records, the last committed business key and source version. Only then apply the recovery procedure appropriate to the installed Batch version and repository store. Record the operator action and resume with the original manifest. There is no safe universal SQL update for every installation.

Test the real failure window

The checked integration test stops a worker after a committed chunk, confirms the process has exited, and verifies that an unchanged launch is refused before recovery. A second test should stop it between a business write and metadata update. Compare durable decisions by source ID. The source kit now force-kills one H2 worker between chunks, proves a same-manifest relaunch is blocked, and calls JobOperator.recover after the worker is confirmed dead. That checked path does not cover the in-commit race or a distributed lease.

Boundary sketch

Java
// Inspect the persisted execution and confirm the old worker is fenced.
// Apply the installed repository recovery procedure, then restart
// with the original identifying manifestId.

Cost and verification

Recovery needs an operator or controller with repository access and a fencing signal. A false death decision can run two workers on the same manifest.

Common Mistakes

  • Do not turn an orphaned execution into a new manifest silently.
  • Do not edit Batch metadata while the original worker may still be alive.
  • Do not claim a two-JVM graceful failure test covers kill -9.

Read next

Spring Batch restart across two JVMs after a recorded failed execution, Spring Batch crash recovery: test the repository and input across two processes, Spring Batch writers: use a stable source key when a chunk is replayed, Spring Batch restart input: pin the manifest before resuming a cursor.

spring
spring-boot
spring-batch
batch-running-execution-recovery
Storage details