Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Java UTF-8 decoding: reject malformed bytes before parsing

Last updated: 30 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

A CharsetDecoder converts bytes to characters and can report malformed input instead of replacing it silently.

Download Java source kit

This complete program targets Java 8. Its displayed output is checked by the tutorial validation script.

Preserve the evidence

A receipt import accepts UTF-8 bytes, not arbitrary platform text. The input contract rejects an incomplete multibyte sequence. A convenience String constructor can replace bad bytes, leaving a character sequence that looks valid enough for a later parser while no longer matching what arrived. Decode first. Validate the resulting fields separately.

The fixture sends one ordinary amount record and one truncated sequence through the same decoder policy. It reports rejection without printing a payload that could contain personal data. This is a byte-level contract; it does not validate the record schema, numeric range or identity of the sender.

A decoder has state

Create a decoder per independent operation or manage its reset lifecycle explicitly. Decoder instances are mutable and should not be shared by concurrent requests. The one-shot decode method handles the end of the supplied buffer; incremental imports need an explicit final decode and flush while retaining incomplete bytes between chunks.

Cap the incoming bytes before decoding. Checking character count afterward does not prevent allocating an enormous byte buffer. A streaming reader can bound retained memory, but its chunk boundaries must not be mistaken for character boundaries. Encoding rules and HTTP field validation handle different layers.

Working program

Java
import java.nio.ByteBuffer;
import java.nio.charset.*;
public class ReceiptUtf8Boundary {
    static String decode(byte[] bytes) throws CharacterCodingException {
        if (bytes.length > 4096) throw new IllegalArgumentException("record too large");
        return StandardCharsets.UTF_8.newDecoder()
            .onMalformedInput(CodingErrorAction.REPORT)
            .onUnmappableCharacter(CodingErrorAction.REPORT)
            .decode(ByteBuffer.wrap(bytes)).toString();
    }
    public static void main(String[] args) throws Exception {
        System.out.println(decode("amount=125".getBytes(StandardCharsets.UTF_8)));
        try { decode(new byte[]{(byte)0xE2, (byte)0x82}); }
        catch (CharacterCodingException rejected) { System.out.println("malformed rejected"); }
    }
}

Output

Output
amount=125
malformed rejected

Costs and boundaries

For b input bytes, this one-shot conversion takes O(b) time and O(b) character storage in the worst case. The 4096-byte limit applies to this record format, not every import. Separate parsing and authentication costs remain outside the decoder.

Common Mistakes

  • Do not use replacement decoding for an import that must reject corrupted bytes.
  • Do not share one mutable decoder across threads.
  • Bound bytes before allocating the full payload.

Read next

Java Unicode: code units, code points and UTF-8 bytes, Java byte streams: partial reads and bounded copying, Spring MVC request validation: reject invalid commands before mutation.

Continue with checked upload and readiness

Continue with Spring CSV uploads: reject malformed UTF-8 and an unknown header.

java
strict-utf8
Storage details