Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Java grapheme boundaries: truncate display text without splitting a cluster

Last updated: 29 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

An extended grapheme cluster is a text segment used for character-boundary analysis; its UTF-16 length can be greater than one and its display width is not fixed.

Download Java source kit

Java 25+. The program uses JDK classes and requires no preview flags.

Choose the unit before truncating

A receipt label can contain a base character followed by a combining accent. Cutting after one UTF-16 unit can leave the accent behind. A supplementary symbol uses a surrogate pair, so even counting code points does not settle every displayed-character boundary.

This Java 25 fixture asks the character BreakIterator for boundaries, then keeps a declared number of clusters. The returned index is still a UTF-16 offset suitable for substring. The program uses a composed-looking sequence with a combining mark and checks its first segment. It also checks a supplementary symbol followed by a letter.

Cluster segmentation is not text shaping. A cluster can occupy several columns, and its appearance depends on font and rendering. A label limited by pixels needs measurement in the UI, not just a server-side cluster count. Java and Unicode data versions must also be part of reproducible text-processing tests.

Own the iterator state

BreakIterator holds the text being analyzed and a current position. Give an operation its own instance rather than sharing a single cursor among requests. Preserve the original label when building a shortened display string; do not overwrite the canonical stored value with a presentation limit.

Working program

Java
import java.text.BreakIterator;
import java.util.Locale;
public class ReceiptLabelClusters {
    static String prefix(String text,int limit){
        if(limit<0)throw new IllegalArgumentException("Negative cluster limit");
        java.util.Objects.requireNonNull(text);
        BreakIterator boundaries=BreakIterator.getCharacterInstance(Locale.ROOT);
        boundaries.setText(text);int end=boundaries.first();
        for(int count=0;count<limit;count++){
            int next=boundaries.next();if(next==BreakIterator.DONE)return text;end=next;
        }
        return text.substring(0,end);
    }
    public static void main(String[] args){
        String label="e\u0301clair";
        System.out.println(prefix(label,1).length());
        System.out.println(prefix("\ud83d\ude00X",1).length());
        System.out.println(prefix(label,0).isEmpty());
        System.out.println(prefix(label,20).equals(label));
    }
}

Output

Output
2
2
true
true

Costs and boundaries

Boundary analysis examines text; budget for work proportional to the analyzed input rather than treating a character limit as a free operation. The returned substring allocates its result. This fixture targets the Java 25 character-boundary contract; it does not establish identical segmentation on every earlier runtime or font.

Common Mistakes

  • String.length counts UTF-16 units, not display clusters.
  • A cluster count is not a pixel-width limit.
  • Do not reuse a mutable iterator concurrently.

Read next

UTF-16 and encoding, Normalization.

java
grapheme-boundaries
Storage details