Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Spring Batch killed worker: fence it, recover STARTED, then restart

Last updated: 1 Oct 20264 min read
tutorial
IntermediateBy AITrove Editorial

A hard-killed JVM leaves STARTED metadata; a new launcher must not run the same instance until recovery is deliberate.

Download Spring source kit

Kill after a known commit

The child worker pauses while processing source ID 3, after the first two-row chunk commits. The test forcibly terminates that process. Its H2 file retains two business rows and a STARTED JobExecution. A fresh JVM tries the same manifest and receives a running-execution rejection. The persisted state is not guessed from the old process's exit code: the test reads the repository table.

Make recovery conditional

Only after the test has waited for the child to die does its recovery path call JobOperator.recover on that STARTED execution. The same identifying parameters then complete IDs 1 through 4. The repository holds one JobInstance and two executions. In a deployed scheduler, a process check alone is insufficient if another host can still own the lease; fence the old worker and inspect the source version first.

Respect the limit

The fixture kills a process between chunks, not in the middle of a database commit. It uses H2 file storage and a key-based writer. It does not establish how a remote database, broker send or distributed lease behaves in the write/metadata race. The test matrix leaves those windows open.

Checked excerpt

Java
JobExecution stale = jobExplorer.getLastJobExecution(
    receiptJob.getName(), manifestParameters);
if (stale == null || !stale.getStatus().isRunning()) {
    throw new IllegalStateException("no running execution to recover");
}
jobOperator.recover(stale);

Cost and verification

A process-kill test starts multiple JVMs and persists H2 metadata, so it belongs in an integration tier. Recovery needs human or controller evidence that the prior writer is fenced.

Common Mistakes

  • Do not call recover while the old worker may still write.
  • Do not add a timestamp parameter to bypass STARTED metadata.
  • Do not describe a between-chunk kill as coverage of an in-commit crash.

Read next

Spring Boot 4 Batch JDBC starter: verify that job history is actually stored, Spring Batch orphaned STARTED execution: inspect before relaunch, Spring Batch source digest: reject changed rows before a restart, Spring Batch restart test matrix: separate failure, process loss and changed input.

spring
spring-boot
spring-batch
batch-killed-worker-recovery
Storage details