Extracting Attributes and Text from XML Elements
findall retrieves a Python list of all XML nodes matching a specified path, enabling batch processing of similar structures.
From Repeated XML to Python Data
When XML contains several similar structures, such as multiple user elements, extracting only one node is not enough. The useful pattern is to find every matching node, receive those nodes as a Python list, and process the list with a for loop. This avoids writing separate extraction code for each user or record.
Selecting Matching Elements
findall retrieves a Python list containing all XML nodes that match the path supplied to it. In the example, root.findall("user") searches the parsed XML tree below the users container for user elements. The path determines which elements are selected from the document hierarchy.
Checking the Matching User Nodes
Determine what lst contains after root.findall("user") runs on the two-user XML document.
Search: The path user is applied to the parsed XML tree below root.
Match: Both user elements match the supplied path.
Collect: findall places the two matching XML nodes into the Python list named lst.
Confirm: The list length is two, so the query located both users before the loop begins.
lst is a Python list containing two Element objects, one for Chuck and one for Brent.
2Walking Through the List
After findall returns, a for loop visits each Element in the list one at a time. The loop variable represents the current XML node. During the first iteration, item represents the first user. During the second iteration, item represents the second user. The loop body runs once for each item and stops after all list items have been processed.
for item in lst: print(item.find("name").text) print(item.get("x"))
Chuck
2
Brent
7Reading Child Text and Attributes
XML data inside a matched node can appear as child-element text or as an attribute on the element itself. Use item.find("name").text to locate the name child element and read the text between its tags. Use item.get("x") to retrieve the x attribute directly from the current user element. Attributes belong to the element being processed, so find is not needed before get.
| Expression | What it accesses | Value for Chuck |
|---|---|---|
| item.find("name") | The name child element | The name Element |
| item.find("name").text | Text inside the name child element | Chuck |
| item.get("x") | The x attribute on the user element | 2 |
The child element and attribute are accessed through different methods.
Tracing One Complete Pass
Processing Both Users
Explain what the loop extracts from each user element.
First iteration: item is bound to the first user Element. item.find("name").text gives Chuck, and item.get("x") gives 2.
Second iteration: item is rebound to the second user Element. item.find("name").text gives Brent, and item.get("x") gives 7.
Completion: The loop has visited both Elements in the list, so it stops.
The extraction produces the name and x attribute for each user: Chuck with 2, followed by Brent with 7.
What do you think happens?
What will the loop print when it processes the two user elements in order?
Reveal answer
Answer: Chuck, 2, Brent, 7
The loop visits the first user before the second user, and the same extraction statements run on each current Element.
Diagnosing Empty or Failing Loops
When iteration does not produce the expected result, inspect findall before changing the loop. First check the list length and inspect the first item. If len(lst) is 0, the path may be wrong or the XML structure may not match the expected structure. If the list length is correct but extraction fails inside the loop, inspect each item and verify the child-element and attribute names exactly. XML names are case-sensitive.
Treating the result of findall as one XML node instead of a list.
findall returns a Python list containing matching Element objects. The individual Element is available inside the for loop through its loop variable.
Fix:
Loop over the list and call find or get on the current item.Assuming an empty list means the loop is broken.
A zero-length list means no elements matched the supplied path, so there is no item for the loop to process.
Fix:
Check the path and compare it with the XML hierarchy.Using the wrong child-element or attribute name.
XML is case-sensitive, so names must match exactly.
Fix:
Inspect the current item and verify the child and attribute names before extracting.
Debug in two stages: first confirm that findall selected the expected number of nodes, then inspect the current item inside the loop before attempting to extract child text or attributes.
Practice the Extraction Pattern
Given the two-user XML structure from this article, write a loop that processes every element in lst and retrieves each user's id text, name text, and x attribute. Before the loop, include a check of the list length.
Hints
- Use for item in lst to visit each current Element.
- Use item.find("id").text and item.find("name").text for child text.
- Use item.get("x") for the attribute on the user element.
- Parse the XML into a tree.
- Call findall with the path for the repeating elements.
- Check the returned list length and inspect an item if needed.
- Use a for loop to bind each Element to the loop variable.
- Use find followed by text for child-element content.
- Use get for an attribute on the current Element.
Key Takeaways
- findall returns a Python list of every XML Element matching a path.
- A for loop visits the returned Elements one at a time, with the loop variable holding the current node.
- Use find(child_name).text to read text from a child element.
- Use get(attribute_name) to read an attribute from the current Element.
- When extraction fails, check the list length, inspect an item, and verify names and paths exactly.
Key Takeaways
- findall gathers all matching XML nodes into a Python list.
- The for loop processes each matching Element separately.
- Child text is accessed with find followed by text.
- Attributes are accessed with get on the current Element.
- Debugging begins by checking what findall returned and whether names match the XML exactly.