Concepts / Extracting Attributes and Text from XML Elements

Extracting Attributes and Text from XML Elements

findall retrieves a Python list of all XML nodes matching a specified path, enabling batch processing of similar structures.

  • Programming

From Repeated XML to Python Data

When XML contains several similar structures, such as multiple user elements, extracting only one node is not enough. The useful pattern is to find every matching node, receive those nodes as a Python list, and process the list with a for loop. This avoids writing separate extraction code for each user or record.

python
matchesmatchesincluded inincluded inuserpathuserChucklsttwo Element objectsuserBrent
What does findall return, and how does one XML path become a Python list containing multiple matching XML nodes?

Selecting Matching Elements

findall retrieves a Python list containing all XML nodes that match the path supplied to it. In the example, root.findall("user") searches the parsed XML tree below the users container for user elements. The path determines which elements are selected from the document hierarchy.

Checking the Matching User Nodes

Determine what lst contains after root.findall("user") runs on the two-user XML document.

Search: The path user is applied to the parsed XML tree below root.

Match: Both user elements match the supplied path.

Collect: findall places the two matching XML nodes into the Python list named lst.

Confirm: The list length is two, so the query located both users before the loop begins.

lst is a Python list containing two Element objects, one for Chuck and one for Brent.

Output
2
containscontainscontainscontainsusersuserChuckid001userBrentnameChuck
How does the path supplied to findall determine which XML elements are selected from the document hierarchy?

Walking Through the List

After findall returns, a for loop visits each Element in the list one at a time. The loop variable represents the current XML node. During the first iteration, item represents the first user. During the second iteration, item represents the second user. The loop body runs once for each item and stops after all list items have been processed.

for item in lst: print(item.find("name").text) print(item.get("x"))

Output
Chuck
2
Brent
7
first itemnext itemno items remainlsttwo Element objectsitemChuckloop completeitemBrent
How does a for loop move from one XML node in the findall result to the next, and what does the loop variable represent at each step?

Reading Child Text and Attributes

XML data inside a matched node can appear as child-element text or as an attribute on the element itself. Use item.find("name").text to locate the name child element and read the text between its tags. Use item.get("x") to retrieve the x attribute directly from the current user element. Attributes belong to the element being processed, so find is not needed before get.

ExpressionWhat it accessesValue for Chuck
item.find("name")The name child elementThe name Element
item.find("name").textText inside the name child elementChuck
item.get("x")The x attribute on the user element2

The child element and attribute are accessed through different methods.

findtextusercurrent Elementnamechild ElementChuck
How does a loop move from a matched parent node to a child element and then to that child element's text value?
getfinduserx2nameChuck
Where are attribute values stored on an XML node, and how does get retrieve a named attribute while processing each matching node?

Tracing One Complete Pass

Processing Both Users

Explain what the loop extracts from each user element.

First iteration: item is bound to the first user Element. item.find("name").text gives Chuck, and item.get("x") gives 2.

Second iteration: item is rebound to the second user Element. item.find("name").text gives Brent, and item.get("x") gives 7.

Completion: The loop has visited both Elements in the list, so it stops.

The extraction produces the name and x attribute for each user: Chuck with 2, followed by Brent with 7.

What do you think happens?

What will the loop print when it processes the two user elements in order?

  • Chuck, 2, Brent, 7
  • The list object only
  • Brent, 7, Chuck, 2
Reveal answer

Answer: Chuck, 2, Brent, 7

The loop visits the first user before the second user, and the same extraction statements run on each current Element.

Diagnosing Empty or Failing Loops

When iteration does not produce the expected result, inspect findall before changing the loop. First check the list length and inspect the first item. If len(lst) is 0, the path may be wrong or the XML structure may not match the expected structure. If the list length is correct but extraction fails inside the loop, inspect each item and verify the child-element and attribute names exactly. XML names are case-sensitive.

  • Treating the result of findall as one XML node instead of a list.

    findall returns a Python list containing matching Element objects. The individual Element is available inside the for loop through its loop variable.

    Fix: Loop over the list and call find or get on the current item.

  • Assuming an empty list means the loop is broken.

    A zero-length list means no elements matched the supplied path, so there is no item for the loop to process.

    Fix: Check the path and compare it with the XML hierarchy.

  • Using the wrong child-element or attribute name.

    XML is case-sensitive, so names must match exactly.

    Fix: Inspect the current item and verify the child and attribute names before extracting.

for loop selects onechild elementattributelstlist of Element objectsfinditemone Elementget
What is the difference between the list returned by findall and a single XML node, and why can using the wrong one cause iteration or extraction failures?

Debug in two stages: first confirm that findall selected the expected number of nodes, then inspect the current item inside the loop before attempting to extract child text or attributes.

Practice the Extraction Pattern

MEDIUM

Given the two-user XML structure from this article, write a loop that processes every element in lst and retrieves each user's id text, name text, and x attribute. Before the loop, include a check of the list length.

Hints
  • Use for item in lst to visit each current Element.
  • Use item.find("id").text and item.find("name").text for child text.
  • Use item.get("x") for the attribute on the user element.
  1. Parse the XML into a tree.
  2. Call findall with the path for the repeating elements.
  3. Check the returned list length and inspect an item if needed.
  4. Use a for loop to bind each Element to the loop variable.
  5. Use find followed by text for child-element content.
  6. Use get for an attribute on the current Element.

Key Takeaways

  1. findall returns a Python list of every XML Element matching a path.
  2. A for loop visits the returned Elements one at a time, with the loop variable holding the current node.
  3. Use find(child_name).text to read text from a child element.
  4. Use get(attribute_name) to read an attribute from the current Element.
  5. When extraction fails, check the list length, inspect an item, and verify names and paths exactly.

Key Takeaways

  • findall gathers all matching XML nodes into a Python list.
  • The for loop processes each matching Element separately.
  • Child text is accessed with find followed by text.
  • Attributes are accessed with get on the current Element.
  • Debugging begins by checking what findall returned and whether names match the XML exactly.