Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Java Unicode: code units, code points and UTF-8 bytes

Last updated: 29 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

A Java String indexes UTF-16 code units; Unicode code points, displayed graphemes and encoded bytes are separate quantities.

Java 8+. The program uses only JDK classes and runs without a framework.

Choose the unit before enforcing a limit

A label containing a rocket and an accented letter can have a different length in every unit that matters to an application. A supplementary code point occupies two char positions. A base letter followed by a combining mark occupies two code points even when a reader sees one accented letter.

The program builds a label with one ASCII letter, one supplementary symbol and a decomposed accented letter. It reports UTF-16 length, code-point count and UTF-8 byte count independently. Those numbers answer different questions. A transport byte limit belongs after encoding; a cursor offset for String.substring belongs in UTF-16 units.

Do not call a code-point count a visible-character count. Grapheme segmentation needs a defined Unicode/text-processing policy. Normalizing canonically equivalent strings is also separate from changing case or removing accents. Choose normalization for the storage/search contract, then apply it consistently to both incoming data and comparisons.

Make decoding failures explicit

new String(bytes, charset) is convenient when replacement of malformed input is acceptable. A strict import often needs rejection instead. CharsetDecoder can report malformed sequences rather than produce replacement characters that later collide in an identifier or hide corrupted data.

The program intentionally feeds an incomplete UTF-8 sequence to a reporting decoder and catches the decoding exception at the import boundary. It does not pretend that UTF-8 verification makes a label safe for HTML, a filename or a database key. Each destination has additional rules. Use StringBuilder for assembly and byte streams for binary transport.

Working program

Java
import java.nio.ByteBuffer;
import java.nio.charset.CharacterCodingException;
import java.nio.charset.CodingErrorAction;
import java.nio.charset.StandardCharsets;
public class LabelEncoding {
    public static void main(String[] args) {
        String label = "A\uD83D\uDE80e\u0301";
        System.out.println("units=" + label.length());
        System.out.println("points=" + label.codePointCount(0, label.length()));
        System.out.println("bytes=" + label.getBytes(StandardCharsets.UTF_8).length);
        try {
            StandardCharsets.UTF_8.newDecoder()
                .onMalformedInput(CodingErrorAction.REPORT)
                .onUnmappableCharacter(CodingErrorAction.REPORT)
                .decode(ByteBuffer.wrap(new byte[]{(byte) 0xC3}));
        } catch (CharacterCodingException rejected) {
            System.out.println("Invalid UTF-8 rejected");
        }
    }
}

Output

Output
units=5
points=4
bytes=8
Invalid UTF-8 rejected

Cost and failure boundaries

Counting code points and encoding each require O(n) scans of the label. UTF-8 conversion allocates a byte array proportional to encoded size. Repeating conversion merely to recheck a limit creates unnecessary arrays; encode once at the boundary and reuse the result where ownership permits it.

A char-based reversal can split a surrogate pair. A code-point reversal preserves supplementary points but can still reorder combining marks away from their base. Neither is a universal user-visible reversal algorithm. Tests should include supplementary symbols, combining sequences, empty text and unpaired surrogates rather than only ASCII names.

Common Mistakes

  • Do not use String.length() as a UTF-8 payload size.
  • Do not cut a string at an arbitrary char boundary and assume every symbol survives.
  • Do not silently repair malformed identifiers when rejection is required.

Extend this boundary

Continue with Java UTF-8 decoding: reject malformed bytes before parsing.

java
unicode-encoding
Storage details