Playground / Anatomy of a URL (urllib.parse)

Take a URL apart

Anatomy of a URL (urllib.parse)

Interactive lab

Try it: Anatomy of a URL (urllib.parse)

How Python's urllib.parse splits a URL into scheme, netloc, path, params, query and fragment, decodes a query string into a dictionary, and percent-encodes values with urlencode — following CPython's rules exactly, including the surprising cases.

How it works

  1. urlparse strips leading spaces, then takes the text before the first ':' as the scheme if it looks like one.
  2. A '//' starts the netloc, which runs to the next '/', '?' or '#'; then the fragment is cut at '#', the query at '?', and params at ';' in the last path segment.
  3. The netloc gives username and password (before the last '@'), hostname (lower-cased) and port (must be 0–65535).
  4. parse_qsl splits the query at '&' and '=', turns '+' into a space, decodes %XX as UTF-8, and drops blank values unless keep_blank_values=True; parse_qs groups the values into lists.
  5. urlencode applies quote_plus (or quote) to every key and value: safe characters stay, every other UTF-8 byte becomes %XX.

Default run (13 steps): urlparse() splits this 95-character URL into six parts: scheme://netloc/path;params?query#fragment. It only cuts the text; % escapes are not expanded. … Done: ParseResult(scheme='https', netloc='ana:pw@Docs.Example.com:8080', path='/guide/a%20b', params='v=2', query='q=python+urllib&lang=en&lang=fr&empty=', fragment='top').

Simplified: URLs are limited to 120 printable ASCII characters and bracketed IPv6 hosts are not modelled. The logic is re-implemented from CPython 3.12's urllib/parse.py and checked against the real module; nothing is fetched.

Educational simulation

Loading the simulation…