Playground / Parsing HTML with BeautifulSoup

Turn raw HTML into a soup and pull out the links

Parsing HTML with BeautifulSoup

Interactive lab

Try it: Parsing HTML with BeautifulSoup

How BeautifulSoup(html, 'html.parser') turns raw HTML text into a tree of Tag objects, tolerating unclosed and mis-nested tags, and how soup('a'), find() and find_all() with class_ / attrs filters pick tags out of it so .get('href', None), tag['href'], .attrs, .contents, .get_text() and .string can extract their data.

How it works

  1. html.parser reads the text left to right and reports start tags (with their attributes), end tags, text, comments and declarations.
  2. Each start tag becomes a Tag inside the tag that is currently open; void elements such as <br> and <img> are closed at once; class and rel values are split into lists.
  3. An end tag closes the most recently opened tag of that name, and every tag still open inside it; an end tag with no matching open tag is ignored. Tags still open at the end are closed then. Whitespace-only text becomes a single '\n' or ' ' string.
  4. soup('a') is the same call as soup.find_all('a'): it walks every tag in document order and keeps those whose name matches and whose attributes pass the filter (a class_ value matches any one class). find() stops at the first match and returns None when nothing matches.
  5. Each result is a Tag: .get('href', None) returns the value or None, tag['href'] raises KeyError when it is missing, .attrs is the attribute dict, .contents the list of children, .get_text() all the text inside, .string the single string child (or None).

Default run (69 steps): Raw HTML: 356 characters of plain text. BeautifulSoup(html, 'html.parser') will turn it into a tree. … [tag.get('href', None) for tag in soup('a')] → ['https://www.py4e.com/', '/lessons', '#top', None]

Simplified: A from-scratch port of Python 3.12's html.parser and BeautifulSoup 4.15's html.parser tree builder, checked against the real libraries on 330+ documents and 12,000+ queries. Documents are limited to 1,500 characters; named character references are limited to about 50 common names (others stay literal text); searches take one tag name and at most one attribute filter (no regular expressions, lists, functions, string= or CSS select()). Other parsers (lxml, html5lib) build different trees for broken HTML.

Educational simulation

Loading the simulation…