XML Element Trees
Try it: XML Element Trees
How ET.fromstring() turns XML text into a tree of Elements (tag, attributes, .text), and how find(), findall() and iter() walk that tree along a path such as users/user.
How it works
- The parser reads start tags, text and end tags left to right, keeping a stack of open elements; each start tag becomes an Element that is a child of the element on top of the stack.
- Text before an element's first child becomes its .text (whitespace included); an end tag must match the open element or parsing stops with ParseError.
- findall(path) starts at the element it is called on (the root is not part of the path) and follows each path segment down one level of children.
- find(path) returns the first match or None; iter(tag) visits every element in document order, including the root.
- Reading .text, .tag or .get('attr') on each result gives the extracted values; reading an attribute from None raises AttributeError.
Default run (29 steps): ET.fromstring() receives 183 characters. The parser reads tags left to right and keeps a stack of open elements. … findall() returned a list of 2 elements; reading e.get('x') from each gives ['2', '7'].
Simplified: A small, hand-written XML parser for a subset: elements, attributes, text, self-closing tags, comments, CDATA, processing instructions and the five predefined entities plus numeric character references. DOCTYPE/DTDs and namespaces are not supported. Error messages follow Python's expat parser for these cases (column numbers can differ for non-ASCII documents). Paths support tags, *, ., //, [@attr], [@attr='v'], [tag] and [n].
Loading the simulation…