Concepts / Parsing XML Documents with ElementTree

Parsing XML Documents with ElementTree

findall retrieves a Python list of all XML nodes matching a specified path, enabling batch processing of similar structures.

  • Programming

From Repeated XML to Python Data

XML often contains several elements with the same structure, such as multiple user, product, or record elements. The practical task is usually to process every matching element rather than extract only one. ElementTree's findall method supports this pattern by locating all XML nodes that match a path and returning them as a Python list.

What findall Returns

Suppose a parsed XML tree has a users container with two user elements. Calling findall with the path for user elements does not return one user element. It returns a Python list containing two Element objects. Each Element represents one matching user subtree. This distinction matters: the result is a collection, while each item inside the collection is an individual XML node.

searchfindallcontainscontainsusers XMLuser, useruser pathmatching pathPython listtwo Element objectsfirst userElementsecond userElement
What XML nodes match the path, and how are those nodes collected into the list returned by findall?
python

After this call, lst is the list. It is not one individual user element. The list can contain zero, one, or multiple matching Element objects, depending on the XML structure and the path.

Tracing the Loop Variable

A for loop processes the list sequentially. On each iteration, the loop variable is bound to the current Element object. With two matching user elements, the loop body runs twice: first with the Element for Chuck, whose id is 001 and whose x attribute is 2, and then with the Element for Brent, whose id is 009 and whose x attribute is 7. The loop stops automatically after all items have been processed.

iterate over listfirst iterationsecond iterationnext itemno items remainfindall resultPython listfor itemone ElementChuckid 001, x 2loop completeall items processedBrentid 009, x 7
What does findall return, what does the loop variable contain on each iteration, and why can using the wrong level cause an error?
python
Output
Chuck 2
Brent 7

The loop variable item is an individual Element during the loop body. That is why item can be used to search for a child element or retrieve an attribute.

Reading Child Text and Attributes

An XML element can hold text between its opening and closing tags, and it can also have attributes on its opening tag. These are accessed differently. To read child-element text, first use find with the child name and then access the resulting element's text property. For a user element, item.find("name").text retrieves Chuck or Brent. To read an attribute belonging directly to the current user element, use item.get("x").

XML dataAccess patternWhat it retrieves
name child elementitem.find("name").textThe text inside the name element
x attribute on the current user elementitem.get("x")The value of the x attribute

Child text and attributes require different access methods.

Extracting One User During an Iteration

A loop variable item currently represents the user whose name is Chuck, id is 001, and x attribute is 2. Determine which access expression reads the name and which reads x.

Locate the child: Use item.find("name") to locate the name child element within the current user element.

Read child text: Add .text to obtain the text content of that child element, which is Chuck for this iteration.

Read the attribute: Use item.get("x") because x is an attribute on the current user element, not a child element.

item.find("name").text reads Chuck, while item.get("x") reads 2.

Debugging the Wrong Iteration Level

When iteration fails, first verify what findall returned before changing the loop body. Check the list length and inspect the first item. A length of zero means that the path or the expected XML structure does not match. If the length is correct but extraction fails inside the loop, inspect each item and check that the child-element and attribute names match exactly. XML names are case-sensitive.

print(len(lst)) print(lst[0]) for item in lst: print(item.find("name").text) print(item.get("x"))

  • Treating the result of findall as one XML Element.

    findall returns a Python list, while the loop variable holds an individual Element.

    Fix: Iterate with for item in lst and call find or get on item.

  • Ignoring an empty findall result.

    A list length of zero indicates that the path or expected XML structure did not match.

    Fix: Check the path and the XML structure, then inspect the returned list.

  • Using the wrong child or attribute name.

    XML is case-sensitive, so names must match exactly.

    Fix: Compare the names in the extraction expressions with the XML names.

Practice the Extraction Pattern

MEDIUM

A findall call has returned a list named records. Write the loop body that retrieves the text of each record's name child and the value of its x attribute. Then state what you would check first if the loop processes no records.

Hints
  • The loop variable should represent one Element from records.
  • Use find followed by .text for the name child.
  • Use get for the x attribute.
  • Check the length of records before investigating the loop body.

What do you think happens?

If findall returns two matching user Elements, how many times does the for-loop body run?

  • Zero times
  • One time
  • Two times
  • Once for every child element inside both users
Reveal answer

Answer: Two times

The loop visits each Element in the returned list once. With two matching user Elements, the body runs once for the first user and once for the second.

The Complete Mental Model

  1. findall searches a parsed XML tree and returns a Python list of matching Element objects.
  2. The list is the collection; the for-loop variable is one individual Element during each iteration.
  3. Use item.find(child_name).text to retrieve text from a child element.
  4. Use item.get(attribute_name) to retrieve an attribute from the current Element.
  5. When something fails, check the list length, inspect the first item, and verify XML names exactly.

Key Takeaways

  • findall converts a matching XML path into a Python list of Element objects.
  • A for loop processes each returned Element one at a time.
  • Child text is accessed with find followed by text, while attributes are accessed with get.
  • Debugging begins by checking the number of returned elements and inspecting an item.
  • The most important level distinction is list versus individual Element.