Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Python strict Base64: reject alternate spellings before decoding a record

Last updated: 30 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

Base64 represents bytes with printable symbols; it does not encrypt or authenticate those bytes.

Download Python source kit

Operation contract

The receipt decoder accepts an ASCII string no longer than 64 characters, uses validate=True, and then requires re-encoding the decoded bytes to reproduce the received spelling. The second check enforces this application’s canonical padded standard alphabet, including unused pad bits. It also limits the decoded payload to 32 bytes. A valid envelope must still satisfy the binary record schema before it becomes a receipt.

Failure and ownership boundary

Whitespace, URL-safe alphabets and omitted padding are deliberately rejected by this protocol. Other protocols may accept them, but silently mixing conventions creates signature and identity disagreements. Compare Python struct: a fixed-endian receipt record with exact byte length, Python HMAC envelopes: authenticate exact bytes and separate replay policy and Python regular expressions: use fullmatch for a complete field contract. A size check here cannot reclaim a large string already buffered by a transport.

Working program

python
import base64
import binascii

def canonical_bytes(encoded):
    if type(encoded) is not str or not encoded.isascii() or len(encoded) > 64:
        raise ValueError("bounded ASCII encoding required")
    try:
        decoded = base64.b64decode(encoded, validate=True)
    except binascii.Error as failure:
        raise ValueError("invalid Base64") from failure
    if len(decoded) > 32 or base64.b64encode(decoded).decode("ascii") != encoded:
        raise ValueError("canonical bounded encoding required")
    return decoded

print("decoded:", canonical_bytes("fQ==").hex())
for received in ("fQ==\n", "fQ", "fR==", "__8="):
    try:
        canonical_bytes(received)
    except ValueError:
        print("spelling rejected")

Output

Output
decoded: 7d
spelling rejected
spelling rejected
spelling rejected
spelling rejected

Costs and limits

Decoding and canonical re-encoding take O(n) time and storage for n encoded characters, capped here. Empty Base64 decodes to empty bytes; whether an empty payload is valid belongs to the next record schema. Do not treat decoding success as schema acceptance or trust.

Common Mistakes

  • Base64 is reversible encoding, not secrecy.
  • Canonical spelling and decoded record validation are separate checks.

Connected lessons

Python struct: a fixed-endian receipt record with exact byte length, Python HMAC envelopes: authenticate exact bytes and separate replay policy, Python strings and bytes: reject decoding errors before parsing records.

python
base64-records
Storage details