A file-backed JobRepository preserves job history across JVM exit. The source kit proves restart after a recorded failed execution and checks a forced kill between chunks with explicit operator recovery.
Spring Batch restart boundaries: recorded failure versus abrupt process loss
Use the checked path first
Process A commits receipt IDs 1 and 2, records FAILED and exits. Process B starts against the same H2 file with the same identifying manifest ID. It observes the same JobInstance and completes with IDs 1 through 4. The two-JVM lesson states the exact assertions. This is a meaningful durability check; the first process had already recorded its failure.
Keep abrupt death separate
The checked forced kill leaves STARTED metadata. The next same-manifest relaunch is refused until the test confirms the old worker is dead and calls JobOperator.recover. The recovery lesson gives the observed boundary. A kill between business and repository writes may replay input. The running-state guide maps that operational boundary, while the input guide pins the bytes being resumed. The H2 kit checks a between-chunk hard stop and a changed-input rejection; it does not prove an in-commit crash, production database or distributed fencing behavior.
Checked excerpt
String first = launch(databaseUrl, "fail", "MANIFEST-263");
String second = launch(databaseUrl, "resume", "MANIFEST-263");
assertEquals(instance(first), instance(second));Cost and verification
The checked test starts Spring Boot twice and uses file-backed H2. The separate H2 checks now cover a hard kill after a committed chunk and changed-source rejection. Target-database locking and an in-commit crash remain unproved.
Common Mistakes
- Do not call a recorded FAILED state an abrupt crash.
- Do not create a new manifest merely to bypass an orphaned running execution.
- Do not infer identical source input from a stable filename.
Read next
Two-process recorded restart, running-execution recovery, restart test matrix.
