Concepts / Working with Python Lists and For Loops

Working with Python Lists and For Loops

The findall method retrieves a list of XML nodes using a path that must include all parent-level elements (except the root) separated by forward slashes.

  • Programming

A List of XML Nodes

When XML data is nested, finding one kind of element often begins by retrieving several matching nodes at once. Python's findall method returns those matching nodes as a list. A for loop can then visit each node in order, allowing you to extract the child-element text and attributes belonging to that node.

Tracing Each Loop Iteration

The for loop takes one item from the list at a time. On the first iteration, the loop variable item refers to the first user node. After that iteration finishes, item refers to the second user node. When every node in the list has been processed, the loop stops.

first itemnext itemno items remainlsttwo user nodesuserChuck, 001, x=2userBrent, 009, x=7endlist exhausted
How does the loop move through the list, and when does it stop?

for item in lst: name = item.find('name').text user_id = item.find('id').text attribute_x = item.get('x') print(name, user_id, attribute_x)

Reading Children and Attributes

A current XML node can contain both child elements and attributes, but they are accessed differently. For child-element content, use find with the child name and then read that element's .text property. For an attribute attached directly to the current node, use get with the attribute name.

find then .textfind then .textgetuserx=2nameChuckid001x2
Which parts of the current user node are read with find and which part is read with get?
XML dataAccess patternExtracted value
<name>Chuck</name>item.find('name').textChuck
<id>001</id>item.find('id').text001
<user x="2">item.get('x')2

Child-element text and node attributes use different access methods.

Diagnosing Empty Results

A path mistake in findall can be difficult to notice because the method can return an empty list without producing an error. If the list is empty, the for loop has no items to process, so its body never runs and no output appears.

include usersuse target nameuse required separatorsuserparent omittedusers/userparent includedusers/memberelement name differsusers/usertarget name matchesusers//userextra slashusers/usersingle slash
What changes when the path omits the parent, names the wrong element, or does not follow the required slash-separated structure?
  • Omitting the parent element from the path

    The user elements are inside users, so user is not a direct child of the current root node. The result is an empty list and the loop does not execute.

    Fix: Use stuff.findall('users/user').

  • Using the parent name without the target element

    This path identifies the users element rather than the user elements that the loop is intended to process.

    Fix: Include the target level with users/user.

  • Using an element name that does not match the XML structure

    The path must describe the names of the parent and target elements in the XML tree.

    Fix: Check the XML names and use users/user.

  • Adding an extra slash to the path

    The described path format uses the parent-level element names separated by forward slashes, as in users/user.

    Fix: Write the path with the required single separator between the listed levels.

Practice the Extraction

Trace the Two Users

Given a list created with stuff.findall('users/user'), determine the values extracted during each loop iteration.

First node: The first user node contains name text Chuck, id text 001, and attribute x with value 2.

Second node: The second user node contains name text Brent, id text 009, and attribute x with value 7.

Loop completion: After the second node, every item in the list has been visited, so the loop ends.

The loop extracts Chuck, 001, 2 and then Brent, 009, 7.

EASY

A root node contains a users element, and users contains user elements. Write the findall path that retrieves the user nodes, then state which method retrieves the name text and which method retrieves the x attribute.

Hints
  • Include each parent-level element between the current node and user.
  • Child-element content uses find followed by .text.
  • An attribute attached to the current node uses get.

Key Takeaways

  1. findall returns a list of matching XML nodes.
  2. The path must include the parent-level element names, such as users/user, while excluding the root name.
  3. A for loop visits each returned node once and in document order.
  4. Use find followed by .text for child-element content.
  5. Use get for attributes attached directly to the current node.
  6. An incorrect path can produce an empty list silently, causing the loop body not to run.

Key Takeaways

  • findall retrieves a list of XML nodes by following a slash-separated path through the tree.
  • A for loop processes those nodes one at a time in document order.
  • Child-element text is extracted with find and .text, while attributes are retrieved with get.
  • A missing parent in the path can silently create an empty list and prevent the loop from running.