Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Python JSON Lines: cap line bytes and validate the whole batch before returning it

Last updated: 1 Oct 20264 min read
tutorial
IntermediateBy AITrove Editorial

JSON Lines represents one JSON value per line, but a line separator does not impose byte, row or schema limits.

Download Python source kit

Operation contract

The receipt reader accepts a caller-owned binary stream, at most four rows and at most 128 received bytes. Each line is capped at 64 bytes and must end in a newline. Duplicate object keys are rejected during decoding; accepted rows require one ASCII receipt ID and an exact bounded integer amount. The reader returns its private list only after every consumed row passes.

Failure and ownership boundary

A rejected later line produces no returned batch, but bytes already consumed are not put back into the source. This operation does not close the stream or promise a network deadline. Decoded JSON types still need validation, and an accepted ID still needs existence/authorization policy before database access. Python JSON validation: reject duplicate members and non-integer amounts, Python input and output: separate parsing from terminal I/O and Pandas chunked CSV aggregation: validate every chunk before returning totals make these distinctions explicit.

Working program

python
import io
import json
import re

def unique_object(pairs):
    result = {}
    for key, value in pairs:
        if key in result: raise ValueError("duplicate JSON key")
        result[key] = value
    return result

def read_receipt_lines(source):
    staged = []; received = 0
    while True:
        line = source.readline(65)
        if not line: return staged
        received += len(line)
        if len(line) > 64 or not line.endswith(b"\n") or received > 128 or len(staged) >= 4:
            raise ValueError("stream budget rejected")
        record = json.loads(line.decode("utf-8"), object_pairs_hook=unique_object)
        if not isinstance(record, dict) or set(record) != {"id", "amount"}:
            raise ValueError("record fields rejected")
        if not isinstance(record["id"], str) or re.fullmatch(r"R-[0-9]{4}", record["id"]) is None:
            raise ValueError("identifier rejected")
        if type(record["amount"]) is not int or not 0 <= record["amount"] <= 1000000:
            raise ValueError("amount rejected")
        staged.append(record)

source = io.BytesIO(b'{"id":"R-0041","amount":125}\n{"id":"R-0042","amount":75}\n')
print(read_receipt_lines(source))
print("caller stream open:", not source.closed)
try:
    read_receipt_lines(io.BytesIO(b'{"id":"R-0041","amount":125}\n{"id":"R-0042","amount":true}\n'))
except ValueError:
    print("later row rejected")

Output

Output
[{'id': 'R-0041', 'amount': 125}, {'id': 'R-0042', 'amount': 75}]
caller stream open: True
later row rejected

Costs and limits

The declared byte and row budgets cap retained parsing work for this fixture. Readline may still block on a real transport. Private staging retains O(n) accepted rows until completion; publication into a database requires its own transaction.

Common Mistakes

  • A line-oriented format does not cap received bytes.
  • Reject duplicate JSON keys before accepting a schema.

Connected lessons

Python JSON validation: reject duplicate members and non-integer amounts, Python input and output: separate parsing from terminal I/O, Pandas chunked CSV aggregation: validate every chunk before returning totals.

Follow the service contract

Python bounded gzip decoding: cap expanded output before publishing it, Python strict Base64: reject alternate spellings before decoding a record.

Continue with Python SQLite SQL length limits: cap statement text separately from bound values.

python
bounded-json-lines
Storage details