Concepts / Navigating XML Trees with findall() and XPath

Navigating XML Trees with findall() and XPath

ET.fromstring() converts an XML string into a tree structure of Element objects, making the hierarchical data queryable.

  • Programming

From Text to Tree

XML begins as text, but its nested tags describe a hierarchy. ET.fromstring() converts that XML string into a tree structure made of Element objects. After conversion, you work with the hierarchy rather than treating the data as one undivided string.

ET.fromstring()containscontainscontainsperson XMLlinear textnamechild Elementpersonroot Elementphonechild Elementemailchild Element
How does nested XML text become a root Element with child Elements after ET.fromstring()?

Suppose the XML describes a person with name, phone, and email elements. After tree = ET.fromstring(data), tree is an Element object representing the person root. The tree contains the complete hierarchy, including the three child Elements. It is not simply the original string printed in another form; it is a structured object that you query with methods.

Tracing the Root Element

tree = ET.fromstring(data)

inputreturnsdataXML stringET.fromstring()conversionpersonroot Element
What path does the XML string follow before you can retrieve an element?

Selecting an Element

The find() method searches the tree for the first element matching a given tag name. Starting with the root Element, you provide the tag name you want to retrieve. For example, tree.find('name') searches for the name element in the person hierarchy.

searchesreturnspersonroot Elementfind('name')tag searchnamefirst matching Element
What path does find() follow, and which single element does it return?
python
Output
Chuck

The two lines separate retrieval from extraction. First, find() produces the name Element. Then .text reads the text inside that Element. In the source example, tree.find('name').text returns Chuck.

Reading Text and Attributes

An Element can hold text content and attributes. Use the .text property to retrieve the text inside an element. Use .get(attribute_name) to retrieve the value of a named attribute. These are two different kinds of information: text belongs to the element's content, while an attribute value is retrieved by its attribute name.

containshas attributereturnsemailElementemail text.textyesattribute valuehide.get('hide')
Where are an element's text value and attribute key-value pairs located, and how are they accessed?
python
ExpressionWhat it accessesSource result
tree.find('name').textText content of the name ElementChuck
tree.find('email').get('hide')Value of the hide attribute on the email Elementyes

The source example separates element text from attribute retrieval.

XPath and Multiple Matches

The topic of this article also includes findall() and XPath as ways of navigating XML trees. The supplied reference facts explain the conversion step, find(), .text, and .get() in detail, but do not specify the exact findall() or XPath syntax or matching rules. Keep the verified distinction clear: find() is documented here as returning the first element matching a tag name.

path targetpath targetpath targetpersonrootnamechildphonechildemailchild
How can a navigation query be understood as moving through the XML hierarchy toward named elements?

When reading XML code, trace it in this order: identify the Element that represents the current tree, identify the tag or navigation expression being used, identify the selected Element, and only then read .text or .get() to determine the extracted value.

Mistakes Beginners Make

  • Treating the result of ET.fromstring() as an ordinary string

    The conversion produces an Element object representing the root and its hierarchy.

    Fix: Use Element methods and properties such as find(), .text, and .get().

  • Confusing the Element with its text content

    find() returns an element, while .text retrieves the content inside that element.

    Fix: Use tree.find('name').text when you need the text value.

  • Using .text to retrieve an attribute

    The source example retrieves the hide attribute with .get('hide').

    Fix: Use tree.find('email').get('hide') for the attribute value.

  • Ignoring preserved whitespace

    Whitespace in XML text is preserved.

    Fix: Use .strip() when surrounding whitespace needs to be removed.

  • Assuming find() represents every possible match

    The documented behavior here is that find() searches for the first matching element.

    Fix: Interpret the result of find() as one selected Element, and consult the specific API documentation when using findall() or XPath.

Practice the Trace

EASY

Using the source example's person tree, trace each expression and state what kind of result it produces: tree = ET.fromstring(data), tree.find('name'), tree.find('name').text, and tree.find('email').get('hide').

Hints
  • The first expression performs the conversion.
  • The second expression selects an Element.
  • The third expression reads text from the selected Element.
  • The fourth expression reads an attribute value from the email Element.

Tracing the Source Example

Determine what each operation represents when the XML root is person with name, phone, and email children.

Convert: ET.fromstring(data) turns the XML string into an Element tree whose root is person.

Select: tree.find('name') searches the tree for the first element matching the name tag.

Read text: Adding .text retrieves the text inside the name Element, which is Chuck in the source example.

Read attribute: tree.find('email').get('hide') retrieves the hide attribute value from the email Element, which is yes in the source example.

The conversion creates the structure, find() selects an Element, .text reads element content, and .get() reads an attribute value.

Key Takeaways

  1. ET.fromstring() converts an XML string into a hierarchy of Element objects.
  2. The resulting root Element represents the complete XML tree and its descendants.
  3. find() searches for the first element matching a tag name.
  4. .text retrieves an element's text content, while .get(attribute_name) retrieves an attribute value.
  5. XML whitespace is preserved, so use .strip() when surrounding whitespace must be cleaned.

Key Takeaways

  • ET.fromstring() changes linear XML text into a queryable Element tree.
  • The root Element contains the hierarchy represented by the XML.
  • find() returns the first Element matching the requested tag name.
  • Use .text for element content and .get() for attribute values.
  • Remember that XML whitespace is preserved and may require .strip().