Extracting Text and Attributes from XML
findall() returns a Python list of all Element objects matching a given path in the XML tree.
From One Record to Many
Real XML data commonly contains groups of related records rather than one isolated element. A users list may contain many user records, a product catalog may contain many products, and a weather feed may contain forecasts for multiple days. Processing this kind of data requires a way to collect every matching element and then examine the elements one at a time.
The findall() method is the collection step. It follows a path through the XML tree, finds every matching element, and returns those matches as a Python list. Each item in that list is a complete XML Element object, not merely a text value. This means every item can still be queried for child elements, attributes, and text.
Following the XML Path
Before using findall(), picture the hierarchy of the XML. In the teaching example, user elements are nested inside a users container. Each user contains child elements such as id and name, and also has an x attribute. The path users/user means: move into the users element, then collect the user elements found inside it.
Collecting User Elements
An XML hierarchy contains a users container with several user elements inside it. Determine what the path users/user is intended to collect.
Read the path: The path names users first and user second, so the search moves through the users container to its user elements.
Collect every match: findall() gathers every user element matching that path rather than stopping after the first one.
Store the matches: The result is a Python list. Each list item is one complete user Element object.
The list represents all matching user subtrees and can be processed with a for loop.
Reading Each Element
findall() gives you a collection, but the useful information is inside each Element in that collection. A for loop processes the list one item at a time. During one loop iteration, the loop variable refers to a single Element object. From that object, you can use find() to locate a child element, get() to access an attribute, or .text to access text content.
Comparing find() and findall()
| Method | Result | Use when |
|---|---|---|
| find() | One matching Element, or None if no match exists | You need one matching element |
| findall() | A Python list of all matching Element objects | You need to process multiple matching elements |
This difference determines what you do next. A result from find() can be queried as one Element. A result from findall() is a list, so you generally loop through it before querying each individual Element. Treating the list as though it were one Element confuses the collection with the objects inside the collection.
Checking the Extraction
When an extraction does not produce the expected records, first inspect the XML hierarchy and compare it with the path you supplied. The path syntax must match the structure. It is also useful to check the length of the list returned by findall() so you can confirm whether matches were found.
Expecting findall() to return one Element
findall() returns a Python list containing the matching Element objects.
Fix:
Iterate through the list with a for loop, then query each individual Element.Using find() when all matching records are needed
find() retrieves only the first matching element, or None if no match exists.
Fix:
Use findall() when the task requires multiple matching elements.Using a path that does not match the XML hierarchy
findall() follows the path provided, so an incorrect path will not identify the intended structure.
Fix:
Picture the XML hierarchy first and verify the path against it.Assuming each loop item is already text
Each item is a complete Element object that may contain child elements, attributes, and text.
Fix:
Use find(), get(), or .text on the current Element.
A users container contains several user elements. Explain what happens in sequence when findall() is used with the path users/user and the returned list is processed by a for loop. Include what the loop variable represents and name two ways to retrieve information from that Element.
Hints
- Start with the result type returned by findall().
- Describe how many times the loop runs in relation to the list.
- Remember that the loop variable holds one complete Element during each iteration.
- The source identifies find(), get(), and .text as ways to query an Element.
The Extraction Pattern
- findall() follows an XML path and returns a Python list of every matching Element object.
- A for loop processes the returned list one Element at a time.
- Each loop item is a complete Element that can be queried with find(), get(), or .text.
- find() returns one matching Element or None, while findall() returns a list of all matches.
- If the result is unexpected, verify the path against the XML hierarchy and check the length of the returned list.
Key Takeaways
- Use findall() when you need all XML elements matching a path.
- Expect findall() to return a list, not a single Element.
- Use a for loop to process each matched Element individually.
- Read child elements, attributes, and text from the current Element with find(), get(), and .text.
- Use find() for one match and findall() for multiple matches.