TTL expires data logically, while tombstones and compaction shape the physical cleanup cost.
Spring Data Cassandra TTL: expiry is a storage and query cost
Expiry does not erase work instantly
A delivery telemetry row may be useful for 47 days. Assigning a TTL bounds its logical lifetime, but expired cells become tombstones until compaction can safely discard them. A wide range query across expired data can spend time scanning tombstones. Partition design limits the range that query must examine.
Choose one retention contract
Do not mix short TTLs, manual deletes and long unbounded partitions without a deletion plan. If a legal hold can extend retention, put held records in a separately modeled table; do not promise that changing a TTL later restores expired values. Use the table's default TTL or per-write options according to the actual retention rule, then inspect the resulting CQL.
Test over time
Write a row with a short test TTL, verify it is readable before expiry and absent afterward, then run a representative range query on a data set containing expired rows. Compare read latency and tombstone warnings; a unit test of the Java entity cannot prove compaction behavior.
Implementation sketch
CREATE TABLE parcel_telemetry (
tenant_id text, parcel_id text, recorded_at timestamp, reading text,
PRIMARY KEY ((tenant_id, parcel_id), recorded_at)
) WITH default_time_to_live = 4060800;Cost and verification
TTL saves explicit delete traffic but creates tombstones. The query cost depends on partition shape, expiry density and compaction policy, not just row count.
Common Mistakes
- Do not equate TTL expiry with immediate disk reclamation.
- Do not assume a later TTL increase revives an expired record.
- Do not test only point reads when production uses range reads.
Read next
Spring Data Cassandra partition keys: model the read before the row, Spring Data Cassandra paging state: continue the same bounded query, Spring Batch restart input: pin the manifest before resuming a cursor.
