Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Python urlsplit: parsing a URL does not authorize a request target

Last updated: 30 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

Urlsplit separates URL components without establishing that the result belongs to an application’s accepted target set.

Download Python source kit

Operation contract

The policy accepts only an ASCII HTTPS URL at learn.aitrove.in, an absent or 443 port, exactly one allowed path, and no user information, query or fragment. Controls and spaces are rejected before parsing because a parser may discard some leading or embedded characters. The function returns a canonical owned target string instead of forwarding received text. No URL is requested by this program; the host is a teaching identifier.

Failure and ownership boundary

This narrow policy is not a general SSRF defense. A real request still needs DNS/address restrictions, redirect checks, TLS verification and network policy at every connection. Percent-encoded path variants are rejected rather than decoded into the accepted path. Python regular expressions: use fullmatch for a complete field contract, Python JSON validation: reject duplicate members and non-integer amounts and Python subprocess: argument vectors, exit codes and bounded fixtures use the same separation between parsing and permission.

Working program

python
from urllib.parse import urlsplit

def accepted_target(text):
    if type(text) is not str or not 1 <= len(text) <= 128 or any(not 33 <= ord(char) <= 126 for char in text):
        raise ValueError("ASCII target bound")
    parsed = urlsplit(text)
    if (parsed.scheme != "https" or parsed.hostname != "learn.aitrove.in" or
        parsed.username is not None or parsed.password is not None or
        parsed.port not in (None, 443) or parsed.path != "/receipts" or
        parsed.query or parsed.fragment):
        raise ValueError("target policy")
    return "https://learn.aitrove.in/receipts"

print(accepted_target("https://learn.aitrove.in:443/receipts"))
for target in ("https://learn.aitrove.in.evil/receipts",
               "https://user@learn.aitrove.in/receipts", "\nhttps://learn.aitrove.in/receipts"):
    try:
        accepted_target(target)
    except ValueError:
        print("target rejected")

Output

Output
https://learn.aitrove.in/receipts
target rejected
target rejected
target rejected

Costs and limits

The 128-character gate bounds parser work and returned storage. Address resolution and network costs are absent. Reusing a hostname allowlist as a production network-boundary claim would omit rebinding, redirects and private-address routing.

Common Mistakes

  • A parsed hostname is not proof that a request is safe or permitted.
  • Reject controls before a parser normalizes the received text.

Connected lessons

Python regular expressions: use fullmatch for a complete field contract, Python JSON validation: reject duplicate members and non-integer amounts, Flask routing: application factories, converters and test clients.

python
url-target-policy
Storage details