An upload format check decodes bounded bytes strictly and rejects content that does not match the import contract.
Spring CSV uploads: reject malformed UTF-8 and an unknown header
The two rejected bodies
One MockMvc request claims text/csv but contains HTML. Another starts with the expected header and ends with an invalid UTF-8 byte. Both receive 422 and leave no staged file. The controller uses a decoder configured to report malformed and unmappable input, then requires the exact receipt_id,amount_minor header followed by content. This avoids accepting replacement characters silently in the local import contract.
The header is only a first gate. It does not validate row quoting, numeric ranges, duplicates, line endings or a CSV injection policy. Those belong in an actual parser and domain validation stage. The byte limit bounds how much the decoder sees, while the path rule keeps a bad body from becoming a stored file.
Define the text contract explicitly
A deployed importer needs a documented encoding and line-ending policy. If it accepts a BOM or CRLF, add deliberate branches and tests; do not rely on platform defaults. After parsing, report row numbers without logging sensitive record values. The fixture's narrow header requirement is useful evidence for early rejection, not a complete CSV specification.
Checked source
String decoded = StandardCharsets.UTF_8.newDecoder()
.onMalformedInput(CodingErrorAction.REPORT)
.onUnmappableCharacter(CodingErrorAction.REPORT)
.decode(ByteBuffer.wrap(content)).toString();
if (!decoded.startsWith(HEADER))
throw new ResponseStatusException(HttpStatus.UNPROCESSABLE_CONTENT);Verification boundary
MultipartReceiptImportTest.rejectsFakeCsvMediaTypeWhenBytesHaveWrongHeader and MultipartReceiptImportTest.rejectsMalformedUtf8BeforeStorage in the downloadable Spring source kit. The excerpt is shortened; the kit contains the complete test.
Costs and limits
The fixture checks two malformed inputs and one valid short body. It does not test CSV grammar, spreadsheet formula injection, record counts or localized number formats. Strict decoding is O(n) in bounded bytes and stores one decoded string in memory.
Common Mistakes
- Do not trust a request Content-Type declaration as byte validation.
- Do not silently replace malformed UTF-8 if exact identifiers matter.
- Do not treat a matching header as proof that every row is valid.
Read next
Spring MVC multipart receipt import: validate before storing, Spring multipart size limits: enforce parser and application budgets, Spring MockMvc multipart tests: what the fixture proves.
