Concepts / Parsing XML Files with ElementTree

Parsing XML Files with ElementTree

ET.fromstring() converts an XML string into a tree structure of Element objects, making the hierarchical data queryable.

  • Programming

From Text to Structure

XML begins as text, but its opening and closing tags describe a hierarchy. ElementTree makes that hierarchy queryable in Python. The central operation is ET.fromstring(): it converts an XML string into a tree of Element objects. After conversion, you work with structured elements rather than treating the XML as one long string.

ET.fromstring()childchildchildXML string<person>...</person>nameChuckpersonroot Elementphonetext contentemailhide=yes
How does one linear XML string become nested Element objects with parent-child relationships?

Conversion Step

What do you think happens?

What does ET.fromstring(data) return when data contains XML describing a person?

Reveal answer

Answer: It returns an Element object representing the root element of the XML tree. In the described example, the root is the person element.

The conversion changes the way the data is represented: tree refers to a structured Element object rather than the original XML string. That object represents the entire hierarchy beneath the root.

data = '<person><name>Chuck</name><phone>...</phone><email hide="yes">...</email></person>' tree = ET.fromstring(data)

The root Element represents the outermost XML element. In the example described by the source, the root is person, and the tree contains three child elements: name, phone, and email. The resulting tree object does not behave like a printed copy of the original string. Instead, you interact with it through ElementTree operations such as find(), .text, and .get().

Finding Nested Elements

Once the XML has become a tree, find() searches that tree for the first element matching a given tag name. Calling tree.find('name') searches from the tree represented by tree and retrieves the name element. Calling tree.find('email') retrieves the email element described in the example.

childchildchildfirst matching tagpersonrootnameChucknamefind('name')phonetext contentemailhide=yes
How does find() move through the tree to locate a specific nested element?
python

find() retrieves an element, not the element's text directly. After finding an element, use .text to read the content inside it or .get(attribute_name) to read one of its attributes.

Reading Element Data

read textread attributeemail<emailhide="yes">...</email>.textelement content.get('hide')attribute value
Where are an element's text value and attributes located, and how are they accessed?
python
Output (expected)
name_text is 'Chuck'; hide_value is 'yes'.

The .text property retrieves the text content inside an element. In the source example, tree.find('name').text returns Chuck. The .get(attribute_name) method retrieves the value of a named attribute. In that same example, tree.find('email').get('hide') returns yes.

python

Complete Extraction

Parse and extract person data

Convert an XML string describing a person into an Element tree, then retrieve the name text and the email hide attribute.

Create the XML string: The XML string contains a person root with name, phone, and email child elements. The email element has a hide attribute.

Parse the string: Pass the XML string to ET.fromstring(). The returned Element represents the person root and its hierarchy.

Find the name element: Use tree.find('name') to retrieve the first element matching the name tag.

Read the name text: Use .text on the found name element. The source example gives Chuck.

Find the email element: Use tree.find('email') to retrieve the email element.

Read the hide attribute: Use .get('hide') on the email element. The source example gives yes.

The XML string has become a structured tree, and the requested values are retrieved as the text Chuck and the attribute value yes.

data = '<person><name>Chuck</name><phone>...</phone><email hide="yes">...</email></person>' tree = ET.fromstring(data) name = tree.find('name').text hidden = tree.find('email').get('hide')

Common Parsing Mistakes

  • Treating the result of ET.fromstring() as if it were still a plain string.

    The conversion produces an Element object representing the root and its hierarchy.

    Fix: Use tree methods and element access such as find(), .text, and .get().

  • Expecting find() to return the text value directly.

    find() returns the matching element.

    Fix: Use tree.find('name').text to retrieve the element's text content.

  • Using .text to retrieve an attribute.

    The text property reads content inside the element, while attributes are retrieved with .get(attribute_name).

    Fix: Use tree.find('email').get('hide') for the hide attribute.

  • Assuming whitespace is automatically removed.

    Whitespace in XML text is preserved.

    Fix: Use .strip() when the extracted text needs whitespace cleaned.

Practice Check

EASY

Suppose an XML string has a person root, a name child containing Ada, and an email child with hide="no". Write the three operations that parse the string, retrieve the name text, and retrieve the hide attribute.

Hints
  • Use ET.fromstring() for the conversion.
  • Use find('name') and then .text for the name.
  • Use find('email') and then .get('hide') for the attribute.
  1. A reliable sequence is: parse the XML string with ET.fromstring(), use find() to retrieve the first matching element, use .text for content, and use .get(attribute_name) for attributes. Remember that the result of parsing is a hierarchical Element object, and that XML whitespace remains present unless you clean it with .strip().

Key Takeaways

  • ET.fromstring() converts an XML string into a tree of Element objects.
  • The returned root Element represents the outermost XML element and its hierarchy.
  • find() searches for the first element matching a tag name.
  • Use .text for element content and .get(attribute_name) for attribute values.
  • XML whitespace is preserved, so use .strip() when cleaned text is needed.