Accessing XML Attributes and Elements
The findall method retrieves a list of XML nodes using a path that must include all parent-level elements (except the root) separated by forward slashes.
The Path Determines the Result
When you retrieve repeated XML elements, the important question is not only what element you want, but also where that element is located. The findall method returns a list of matching nodes, but its path must describe the parent levels between the current node and the target. If user elements are inside a users element, the path is users/user. The root element is usually omitted because the search begins from the current node, which is the root in this situation.
Finding the User Nodes
lst = stuff.findall('users/user')
The path users/user does not mean that there is one single user. It describes the route to the target level: first users, then each user beneath it. The returned list contains the matching user nodes. A path such as user asks for user elements that are direct children of the current node. When users is the direct child and user is nested inside users, that shortened path does not reach the target.
Moving Through the Node List
The for loop visits each node in the list one at a time and in the order in which the nodes appear in the XML document. During one iteration, item refers to one user node. The find method searches inside that node for a child element. Adding .text retrieves the text inside that child element. The get method reads an attribute attached directly to the current user node.
| Current node | Child-element text | Attribute value |
|---|---|---|
| first user | name Chuck; id 001 | x is 2 |
| second user | name Brent; id 009 | x is 7 |
Values extracted as the loop processes the two user nodes
Separating Text from Attributes
An XML node can contain both child elements and attributes, but they are accessed differently. In the user node with x="2", x is an attribute on the user opening tag, so item.get('x') retrieves its value. The name and id values are inside child elements such as name and id, so item.find('name').text and item.find('id').text retrieve their text content. The choice between find and get depends on where the data is stored.
| XML location | Access pattern | Value in the first user node |
|---|---|---|
| Child element | item.find('name').text | Chuck |
| Child element | item.find('id').text | 001 |
| Attribute | item.get('x') | 2 |
Diagnosing Path Mistakes
Leaving out the parent level
The search looks for user elements directly under the current node, but the user elements are inside users.
Fix:
Use stuff.findall('users/user').Using the parent name as the target path
This searches for users nodes rather than the user nodes nested inside them.
Fix:
Include the target level with stuff.findall('users/user').Using the wrong element name
The path does not name the user elements that exist in the structure.
Fix:
Check each element name and use the actual path users/user.Confusing an attribute with a child element
x is an attribute on the user node, not a child element.
Fix:
Use item.get('x') for the attribute.
A Complete Extraction Pass
Read Each User Record
Retrieve the name, id, and x attribute from every user node beneath users.
Locate the repeated nodes: Use findall('users/user') so the path includes the users parent and reaches the user children.
Process the first node: The loop assigns the first user node to item. Its child text values are Chuck and 001, and its x attribute is 2.
Process the second node: The loop then assigns the second user node to item. Its child text values are Brent and 009, and its x attribute is 7.
Finish the loop: After the second node, the list has no remaining nodes, so iteration ends.
The two user nodes are visited once each, in document order, and each node supplies its name text, id text, and x attribute.
What do you think happens?
If findall returns an empty list, what happens when the for loop runs?
Reveal answer
Answer: The loop never executes
A for loop processes the items that are present in the list. An empty list has no items, so there are no iterations and no output from the loop.
Practice and Verification
Suppose a tree has a users element containing user elements. Write the findall statement that retrieves the user nodes, then write the two expressions needed to obtain the name text and the x attribute from the current loop variable item.
Hints
- Include both users and user in the path, separated by a forward slash.
- Use find with the child-element name and follow it with .text.
- Use get with the attribute name.
- To extract repeated XML records, start at the current node and write every parent-level element needed to reach the target, leaving out the root itself. Use findall('users/user') to receive a list of user nodes. A for loop visits those nodes one at a time and in document order. Within each iteration, use item.find('child').text for text inside a child element and item.get('attribute') for an attribute attached directly to the current node. If the path is incomplete or names the wrong level, findall may return an empty list and the loop may produce no output.
Key Takeaways
- findall returns a list of nodes reached through the specified parent-to-child path.
- The path normally omits the root because the search starts from the current node.
- A for loop processes each returned node once and in document order.
- Use find followed by .text for child-element content and get for attributes.
- An incomplete or incorrect path can return an empty list without an error.