Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Python CSV headers: reject duplicate names before row mapping

Last updated: 30 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

CSV header names become mapping keys in DictReader, so repeated names can hide an earlier column.

Download Python source kit

Operation contract

The fixture reads a header row first and checks that it is exactly the expected two-column set with no duplicate spelling. It then hands the remaining text to DictReader using those names. A repeated amount column is rejected before any row is accepted. The normal rows retain text values so a later step can parse amounts under a separate numeric rule.

Failure and ownership boundary

Opening a real file needs an encoding, newline policy, byte cap, row cap and field-size bound. Checking after converting a row to a dict is too late. Column names that differ only by case are distinct here; choose a canonicalization rule before changing that contract. Python CSV imports: parse quoted fields before validating rows, Python CSV ingestion: cap bytes, rows and fields before publication and Python JSON duplicate keys: reject ambiguous object fields cover adjacent formats.

Working program

python
import csv
import io

def receipt_rows(text):
    if type(text) is not str or len(text) > 200:
        raise ValueError("input size")
    stream = io.StringIO(text, newline="")
    header = next(csv.reader(stream), None)
    if header is None or len(header) != 2 or set(header) != {"id", "minor"}:
        raise ValueError("header contract")
    rows = list(csv.DictReader(stream, fieldnames=header))
    if len(rows) > 5 or any(None in row or any(value is None for value in row.values()) for row in rows):
        raise ValueError("row width")
    return rows

print(receipt_rows("id,minor\nR41,125\nR42,75\n"))
try:
    receipt_rows("id,minor,minor\nR41,125,0\n")
except ValueError:
    print("duplicate header rejected")

Output

Output
[{'id': 'R41', 'minor': '125'}, {'id': 'R42', 'minor': '75'}]
duplicate header rejected

Costs and limits

Parsing is O(n) in bounded text length and materializes O(n) row text here. DictReader does not grant a streaming memory guarantee once its result is converted to a list.

Common Mistakes

  • Validate header names before row mapping.
  • A parsed numeric string is not yet a validated amount.

Connected lessons

Python CSV imports: parse quoted fields before validating rows, Python CSV ingestion: cap bytes, rows and fields before publication, Python JSON duplicate keys: reject ambiguous object fields.

python
csv-duplicate-headers
Storage details