Parsing HTML with BeautifulSoup
Try it: Parsing HTML with BeautifulSoup
How BeautifulSoup(html, 'html.parser') turns raw HTML text into a tree of Tag objects, tolerating unclosed and mis-nested tags, and how soup('a'), find() and find_all() with class_ / attrs filters pick tags out of it so .get('href', None), tag['href'], .attrs, .contents, .get_text() and .string can extract their data.
How it works
- html.parser reads the text left to right and reports start tags (with their attributes), end tags, text, comments and declarations.
- Each start tag becomes a Tag inside the tag that is currently open; void elements such as <br> and <img> are closed at once; class and rel values are split into lists.
- An end tag closes the most recently opened tag of that name, and every tag still open inside it; an end tag with no matching open tag is ignored. Tags still open at the end are closed then. Whitespace-only text becomes a single '\n' or ' ' string.
- soup('a') is the same call as soup.find_all('a'): it walks every tag in document order and keeps those whose name matches and whose attributes pass the filter (a class_ value matches any one class). find() stops at the first match and returns None when nothing matches.
- Each result is a Tag: .get('href', None) returns the value or None, tag['href'] raises KeyError when it is missing, .attrs is the attribute dict, .contents the list of children, .get_text() all the text inside, .string the single string child (or None).
Default run (69 steps): Raw HTML: 356 characters of plain text. BeautifulSoup(html, 'html.parser') will turn it into a tree. … [tag.get('href', None) for tag in soup('a')] → ['https://www.py4e.com/', '/lessons', '#top', None]
Simplified: A from-scratch port of Python 3.12's html.parser and BeautifulSoup 4.15's html.parser tree builder, checked against the real libraries on 330+ documents and 12,000+ queries. Documents are limited to 1,500 characters; named character references are limited to about 50 common names (others stay literal text); searches take one tag name and at most one attribute filter (no regular expressions, lists, functions, string= or CSS select()). Other parsers (lxml, html5lib) build different trees for broken HTML.
Loading the simulation…