Understanding XML Structure and Hierarchy
The findall method retrieves a list of XML nodes using a path that must include all parent-level elements (except the root) separated by forward slashes.
A Path That Finds Nothing
A findall statement can fail quietly. If user elements are nested inside a users element, findall('user') does not search everywhere for user elements. It looks for user elements directly under the current node. Because that location does not exist in this structure, the result is an empty list, and a loop over that list produces no output.
What do you think happens?
Suppose the current root contains users, and users contains user elements. What does findall('user') return?
Reveal answer
Answer: The direct user children of the current node
In this structure, user is not a direct child of the current node. The result is therefore an empty list, and a for loop over it never executes.
Following Parent Elements
findall retrieves a list of XML nodes by following a path. Each parent-level element between the current node and the target must appear in the path, separated by forward slashes. The root is the starting point for the search, so its name is omitted from the path. For user elements nested inside users, the path is users/user. This means: begin at the current node, move into users, and collect its user children.
The result stored in lst is a list of matching user nodes. The list contains the user nodes in the order in which they appear in the XML document. The path selects the nodes; it does not yet extract the text inside their children or the attributes attached to them.
Reading Each Matched Node
Extracting Names, IDs, and Attributes
Process the user nodes returned by findall('users/user'). For each node, retrieve the name text, the id text, and the x attribute.
Retrieve the list: findall('users/user') returns the matching user nodes as a list.
Visit one node: The for loop assigns one user node at a time to the variable item.
Read child text: item.find('name').text retrieves the text inside the name child element. item.find('id').text does the same for id.
Read the attribute: item.get('x') retrieves the value of the x attribute attached directly to the current user node.
The loop processes each user once, extracting child-element text and the current node's attribute value.
lst = stuff.findall('users/user') for item in lst: name = item.find('name').text user_id = item.find('id').text attribute_x = item.get('x')
Chuck 001 2
Brent 009 7The loop visits the first user node, extracting Chuck, 001, and 2. It then visits the second user node, extracting Brent, 009, and 7. After the second iteration, the list has no remaining nodes, so the loop ends.
Text and Attributes
A child element's text and a node's attribute are stored in different parts of the XML structure. Text such as Chuck appears inside a child element such as name, so it is accessed with find('name').text. An attribute such as x="2" appears on the opening tag of the user element, so it is accessed with get('x'). Use find for a child element and its text; use get for an attribute on the current node.
| XML location | Access pattern | Value in the source example |
|---|---|---|
| Child element content | item.find('name').text | Chuck |
| Child element content | item.find('id').text | 001 |
| Attribute on the current node | item.get('x') | 2 or 7 |
find followed by text reads child-element content; get reads an attribute on the current node.
Path Mistakes and Repairs
Leaving out the parent element
The search looks for user elements directly under the current node, but the user elements are children of users.
Fix:
Use findall('users/user').Assuming findall searches the entire tree from any shortened name
The path describes the required hierarchy from the current node; it is not an unrestricted search.
Fix:
Write every parent-level element between the current node and the target, separated by forward slashes.Using get for child-element content
name is a child element in the source structure, not an attribute on the user node.
Fix:
Use item.find('name').text.Using find to read an attribute
x is an attribute attached to the user node, not a child element.
Fix:
Use item.get('x').
Check Your Path
An XML tree has a users element containing user elements. Write the findall path that retrieves the user nodes, then state which method reads a child name element and which method reads the x attribute on the current user node.
Hints
- Include the parent element before the target element.
- Use find followed by text for the name child.
- Use get with the attribute name for x.
Practice Solution
Retrieve user nodes nested inside users, then access each user's name text and x attribute.
Build the path: The parent users and target user are written as users/user.
Iterate: A for loop processes each node in the list one at a time.
Read the two locations: Use item.find('name').text for the child text and item.get('x') for the node attribute.
lst = stuff.findall('users/user'); then use a for loop with item.find('name').text and item.get('x') inside it.
The Complete Mental Model
- findall returns a list of nodes selected by a parent-to-child path.
- The path includes all parent-level elements between the current node and the target, but not the root itself.
- A for loop visits each matching node once and preserves the order in which the nodes appear.
- Use find followed by text for child-element content and get for attributes on the current node.
- An omitted parent can produce an empty list, causing the loop to run without output.
Key Takeaways
- Describe the XML hierarchy from the current node to the target when writing a findall path.
- Use users/user rather than user when user elements are nested inside users.
- Process the returned list with a for loop, one node at a time.
- Read child text with find(...).text and attributes with get(...).
- Treat an empty list and a loop with no output as possible signs of an incomplete path.