A file-backed JobRepository retains a failed execution and its committed business rows after the first JVM exits.
Spring Batch restart across two JVMs after a recorded failed execution
Hold the job identity still
The source kit launches MANIFEST-263 twice against the same file-backed H2 database. Process A writes IDs 1 and 2, then fails while writing the next chunk. It exits after Spring records FAILED. Process B starts a fresh JVM with the same manifest parameter. It sees the same JobInstance ID, completes the job, and leaves exactly IDs 1, 2, 3 and 4 in the target. The test also reads JDBC metadata: one persisted JobInstance and two executions. Repository configuration matters because matching numeric IDs alone can mislead. The identifying parameter is the join between those executions.
Observe rows, not just status
The test checks both the Batch status and the target IDs after each process. A COMPLETED label alone could hide a missing record. The writer uses an H2 key-based MERGE, so a replay cannot create a duplicate source ID; it could still overwrite a changed payload. Input versioning and a payload comparison are separate requirements.
Mark the fault boundary
This is a graceful process exit after a recorded failure. It is not a kill during a database commit. The child JVM is gone before the second launch, so repository durability is checked, but The orphaned STARTED execution is now checked separately; the business-write/metadata-commit race is not. Test those against the intended database before advertising abrupt-crash recovery.
Checked excerpt
String databaseUrl = "jdbc:h2:file:" + temporaryDirectory.resolve("durable-receipts");
String failed = launch(databaseUrl, "fail", "MANIFEST-263");
String resumed = launch(databaseUrl, "resume", "MANIFEST-263");
assertEquals(instance(failed), instance(resumed));Cost and verification
Two Boot startups and disk-backed metadata make this slower than an in-memory step test. H2 file semantics cannot establish the same commit window as a production database.
Common Mistakes
- Do not call an orderly FAILED execution an abrupt process crash.
- Do not change the identifying manifest ID for the second launch.
- Do not infer exactly-once business effects from a final row count alone.
Read next
Spring Batch crash recovery: test the repository and input across two processes, Spring Batch job identity: a manifest ID defines restart versus a new run, Spring Batch restart input: pin the manifest before resuming a cursor, Spring Batch orphaned STARTED execution: inspect before relaunch.
