Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Java HashSet: deduplication and equality

Last updated: 28 Sept 20263 min read
tutorial
IntermediateBy AITrove Editorial

HashSet stores distinct elements using hashing and equality, with no promised iteration order.

Choose what counts as the same value

Deduplication is a policy. Are two receipt IDs equal when their letter case differs? Are leading spaces meaningful? Decide before inserting values. A set can enforce your chosen equality rule; it cannot infer the business meaning of an identifier.

This import treats trimmed receipt strings as case-sensitive identifiers. It rejects an empty identifier and counts duplicates. It does not normalise letter case, because changing a case-sensitive provider ID could incorrectly merge two receipts.

A HashSet permits a null element, but this pipeline does not. Restricting inputs makes downstream work easier: a reader knows every stored value is a usable identifier. For original encounter order, choose LinkedHashSet instead of sorting after each insertion.

Use the result of add

add returns false when an equal element already exists. Checking contains and then calling add repeats the search and creates a wider race window if the code is later shared incorrectly between threads.

If the imported record is a custom type, its equality and hash contract determine deduplication. Object identity is often the wrong rule for a record reconstructed from a file.

A process-local set is not a permanent idempotency store. Separate service instances have separate sets, and restarting a process clears its history. Use a durable uniqueness constraint when accepting an ID twice would create an external side effect.

Working program

Java
import java.util.HashSet;
import java.util.Set;

public class ReceiptDeduplicator {
    public static void main(String[] args) {
        String[] importedIds = {" rcpt-41 ", "rcpt-82", "rcpt-41", ""};
        Set<String> acceptedIds = new HashSet<>();
        int duplicates = 0;
        int rejected = 0;
        for (String rawId : importedIds) {
            String receiptId = rawId.trim();
            if (receiptId.isEmpty()) {
                rejected++;
            } else if (!acceptedIds.add(receiptId)) {
                duplicates++;
            }
        }
        System.out.println("accepted=" + acceptedIds.size());
        System.out.println("duplicates=" + duplicates);
        System.out.println("rejected=" + rejected);
    }
}

Output

Output
accepted=2
duplicates=1
rejected=1

Cost and design choices

Processing n bounded-length IDs takes expected O(n) time with well-distributed hashes and O(k) entry storage for k distinct IDs. Trimming can allocate new strings; account for input text length when estimating a large import.

HashSet uses hashing storage rather than a sorted index. If the main operation is checking membership, that is a sensible fit. If the application needs nearest values, ordered ranges, or a floor entry, study ordered maps instead.

An unbounded stream of distinct IDs creates an unbounded memory requirement. A maximum batch size, retention window, or external store is a design decision the collection will not make for you.

Common Mistakes

  • Do not assert the order of HashSet.toString().
  • Do not use mutable fields in an element’s hash while it is stored.
  • Do not claim process-local deduplication guarantees exactly-once delivery.
java
hashset
Storage details