Concepts / Accessing XML Attributes and Elements

Accessing XML Attributes and Elements

The findall method retrieves a list of XML nodes using a path that must include all parent-level elements (except the root) separated by forward slashes.

  • Programming

The Path Determines the Result

When you retrieve repeated XML elements, the important question is not only what element you want, but also where that element is located. The findall method returns a list of matching nodes, but its path must describe the parent levels between the current node and the target. If user elements are inside a users element, the path is users/user. The root element is usually omitted because the search begins from the current node, which is the root in this situation.

containscontainscontainsrootuserspath segment: usersusertarget nodeusertarget node
Which parent and child elements must the path include for findall to reach the target user nodes, and why is the root omitted?

Finding the User Nodes

lst = stuff.findall('users/user')

The path users/user does not mean that there is one single user. It describes the route to the target level: first users, then each user beneath it. The returned list contains the matching user nodes. A path such as user asks for user elements that are direct children of the current node. When users is the direct child and user is nested inside users, that shortened path does not reach the target.

Moving Through the Node List

python

The for loop visits each node in the list one at a time and in the order in which the nodes appear in the XML document. During one iteration, item refers to one user node. The find method searches inside that node for a child element. Adding .text retrieves the text inside that child element. The get method reads an attribute attached directly to the current user node.

first iterationnext iterationlist exhausteduser nodestwo nodesChuckfirst user nodeloop endsBrentsecond user node
What nodes does findall return, and how does the for loop visit each node one at a time?
Current nodeChild-element textAttribute value
first username Chuck; id 001x is 2
second username Brent; id 009x is 7

Values extracted as the loop processes the two user nodes

Separating Text from Attributes

An XML node can contain both child elements and attributes, but they are accessed differently. In the user node with x="2", x is an attribute on the user opening tag, so item.get('x') retrieves its value. The name and id values are inside child elements such as name and id, so item.find('name').text and item.find('id').text retrieve their text content. The choice between find and get depends on where the data is stored.

child elementchild elementattributeusernameChuckx2id001
Where are a user node's child-element text and attribute value located while the loop processes that node?
XML locationAccess patternValue in the first user node
Child elementitem.find('name').textChuck
Child elementitem.find('id').text001
Attributeitem.get('x')2

Diagnosing Path Mistakes

omit parentchange element nameadd path segmentusers/usertwo user nodesuserempty listusers/memberno matching nodesusers/user/namedifferent target level
How does the set of matched nodes change when the parent level is missing, the element name is wrong, or an extra path segment is added?
  • Leaving out the parent level

    The search looks for user elements directly under the current node, but the user elements are inside users.

    Fix: Use stuff.findall('users/user').

  • Using the parent name as the target path

    This searches for users nodes rather than the user nodes nested inside them.

    Fix: Include the target level with stuff.findall('users/user').

  • Using the wrong element name

    The path does not name the user elements that exist in the structure.

    Fix: Check each element name and use the actual path users/user.

  • Confusing an attribute with a child element

    x is an attribute on the user node, not a child element.

    Fix: Use item.get('x') for the attribute.

A Complete Extraction Pass

Read Each User Record

Retrieve the name, id, and x attribute from every user node beneath users.

Locate the repeated nodes: Use findall('users/user') so the path includes the users parent and reaches the user children.

Process the first node: The loop assigns the first user node to item. Its child text values are Chuck and 001, and its x attribute is 2.

Process the second node: The loop then assigns the second user node to item. Its child text values are Brent and 009, and its x attribute is 7.

Finish the loop: After the second node, the list has no remaining nodes, so iteration ends.

The two user nodes are visited once each, in document order, and each node supplies its name text, id text, and x attribute.

What do you think happens?

If findall returns an empty list, what happens when the for loop runs?

  • The loop executes once with no current node
  • The loop never executes
  • The loop automatically searches deeper in the tree
  • The loop raises an error because the list is empty
Reveal answer

Answer: The loop never executes

A for loop processes the items that are present in the list. An empty list has no items, so there are no iterations and no output from the loop.

Practice and Verification

MEDIUM

Suppose a tree has a users element containing user elements. Write the findall statement that retrieves the user nodes, then write the two expressions needed to obtain the name text and the x attribute from the current loop variable item.

Hints
  • Include both users and user in the path, separated by a forward slash.
  • Use find with the child-element name and follow it with .text.
  • Use get with the attribute name.
  1. To extract repeated XML records, start at the current node and write every parent-level element needed to reach the target, leaving out the root itself. Use findall('users/user') to receive a list of user nodes. A for loop visits those nodes one at a time and in document order. Within each iteration, use item.find('child').text for text inside a child element and item.get('attribute') for an attribute attached directly to the current node. If the path is incomplete or names the wrong level, findall may return an empty list and the loop may produce no output.

Key Takeaways

  • findall returns a list of nodes reached through the specified parent-to-child path.
  • The path normally omits the root because the search starts from the current node.
  • A for loop processes each returned node once and in document order.
  • Use find followed by .text for child-element content and get for attributes.
  • An incomplete or incorrect path can return an empty list without an error.