Concepts / Error Handling in Python

Error Handling in Python

ET.fromstring converts an XML string into a tree structure, transforming flat text into a navigable, queryable object.

  • Programming

From Text to Structure

XML data often arrives from a web service, file, or another system as a flat string of characters. Although that string contains tags, nesting, and attributes, your Python code cannot conveniently search it as a tree until it has been parsed. ElementTree provides fromstring for converting the XML string into a hierarchical tree object.

The transformation has two important stages. First, ET.fromstring reads the XML text and builds an in-memory tree. The outermost tag becomes the root element, and nested tags become child elements. Second, methods such as find let your code search that hierarchy instead of scanning the original characters. This is the basic parse-then-search process for working with XML.

fromstringchildchildXML stringperson with name and emailpersonroot elementnametext: Adaemailtext: ada@example.com
How does the nested structure written in the XML string become parent and child Element objects after ET.fromstring runs?

Tracing the Parse

python

In this example, xml_data is still only a string before parsing. After ET.fromstring(xml_data) runs, root refers to the tree's outermost person element. The nested name and email tags are represented as child elements. The XML's hierarchy has therefore become a structure that Python code can query.

ET.fromstringXML stringflat charactersElement treenested objects
What changes when the XML string is passed to ET.fromstring?

fromstring does not merely return part of the original text. It builds a tree object representing the entire XML document, including its root, child elements, text content, and attributes.

Finding and Reading Elements

Once the tree exists, call find on an element to search for the first element matching a tag name. The returned value is an element object, not the text itself. To obtain the characters inside that element, use .text. To obtain an attribute, use .get('attribute_name').

childchildfind('name')personnameAdaname elementfirst matching elementemailada@example.com
How does find locate the specific element requested by a tag name?

Reading a Name and an Attribute

Given a parsed person tree, retrieve the text inside the name element and the hide attribute on the email element.

Parse: Pass the XML string to ET.fromstring. The result is the root person element and its hierarchy.

Find the name: Call root.find('name'). The first matching name element is returned.

Read text: Access the returned element's .text property to obtain Ada.

Find the email: Call root.find('email') to obtain the email element.

Read the attribute: Call email.get('hide') to retrieve the value associated with the hide attribute.

The extracted values are Ada and yes.

import xml.etree.ElementTree as ET xml_data = "<person><name>Ada</name><email hide=\"yes\">ada@example.com</email></person>" root = ET.fromstring(xml_data) name_element = root.find("name") email_element = root.find("email") print(name_element.text) print(email_element.get("hide"))

Output
Ada
yes

Text and Attribute Data

An XML element can carry different kinds of information. The characters between an opening and closing tag are available through .text. Attribute values are stored separately and are retrieved by naming the attribute in .get(). In the generated example, Ada is the text of name, while yes is the value of the hide attribute on email.

.text.get('hide')emailada@example.comtexthideyes
What information belongs to an element's text content, and what information is stored separately as attribute key-value pairs?
XML informationHow to access itGenerated example
Characters inside an element.textAda
Named attribute value.get('attribute_name')yes
The element itselffind('tag_name')The name or email element object

Missing Matches and Exact Names

  • Treating find as though it directly returns text.

    find returns an element object or None, not the characters inside the element.

    Fix: Use name.text after confirming that the element was found.

  • Using .get() to retrieve element text.

    The source distinguishes text content from attributes. Text is accessed with .text.

    Fix: Use name.text for the characters inside the element, and .get('attribute_name') for an attribute.

  • Changing the capitalization of a tag or attribute name.

    XML tag and attribute names are case-sensitive.

    Fix: Match the exact spelling used in the XML.

  • Using a missing element without handling None.

    If no matching element exists, find returns None.

    Fix: Handle the None case before attempting to access element data.

Guided Practice

EASY

Given the XML string <person><name>Ravi</name><email hide="no">ravi@example.com</email></person>, describe the result of parsing it with ET.fromstring. Then identify which expression retrieves Ravi and which expression retrieves no from the hide attribute.

Hints
  • The outermost person tag becomes the root element.
  • Use find with the exact tag name.
  • Use .text for characters inside a tag and .get('hide') for the named attribute.

What do you think happens?

What does root.find('email').get('hide') retrieve from <email hide="no">ravi@example.com</email>?

  • ravi@example.com
  • email
  • no
  • None
Reveal answer

Answer: no

find returns the email element, and get('hide') retrieves the value of its hide attribute. The address between the tags is the element's text and would be accessed with .text.

Key Takeaways

  1. ET.fromstring converts a flat XML string into a hierarchical Element tree.
  2. The outermost XML tag becomes the root element, with nested tags represented as child elements.
  3. find searches for the first element matching a tag name and returns an element object or None.
  4. Use .text for an element's text content and .get('attribute_name') for an attribute value.
  5. XML tag and attribute names are case-sensitive, and code must handle the None result when no match is found.

Key Takeaways

  • Parsing comes before searching: ET.fromstring turns XML text into a navigable tree.
  • find locates the first matching element, but it returns an element object or None rather than a string.
  • Read characters inside an element with .text and read named attributes with .get().
  • Exact capitalization matters for XML tag and attribute names.
  • Handle the None case when a requested element may not exist.