Parsing XML with Python's ElementTree
fromstring converts an XML string into a tree structure, enabling programmatic access to its data.
From Text to Structure
When XML data arrives as a plain string, it is only a sequence of characters. You cannot conveniently navigate to a particular piece of information until the string has been transformed into a structure. Python's ElementTree provides the fromstring function for this conversion. It reads the XML string and builds a hierarchical tree whose elements mirror the nesting in the XML.
The overall extraction process has three steps: parse the XML string with ET.fromstring, locate an element with find, and extract either its text with .text or an attribute value with .get().
Building the Element Tree
The fromstring function returns an Element object representing the root element. The root is the entry point to the complete tree. For example, if the outermost tag is person, the returned Element represents person and gives your code access to the nested elements inside it.
import xml.etree.ElementTree as ET xml_data = "<person><name>Ada</name><phone>555-0100</phone></person>" root = ET.fromstring(xml_data) print(root.tag)
personFinding a Nested Element
After parsing, use the find method to search the tree for a tag name. find returns an Element object for the first matching element. That returned object becomes the specific part of the tree from which you can extract data.
AdaThe check is important. If no element has the requested tag, find returns None. None is not an Element, so attempting to access .text or .get() on it causes the code to fail. Check the result before reading its properties.
Reading Text and Attributes
XML stores information in two distinct places. Text content appears between an element's opening and closing tags and is accessed through .text. Attributes are key-value pairs in the opening tag and are accessed with .get(). An element may have text, attributes, both, or neither.
| XML location | Python access | What it retrieves |
|---|---|---|
| Between opening and closing tags | .text | Element content |
| Key-value pair in the opening tag | .get() | Attribute value |
Text content and attributes use different access mechanisms.
Ada
p1In this example, name_element.text reads the content inside the name tags. In contrast, root.get("id") reads the id attribute from the person opening tag. The two expressions retrieve different kinds of XML data.
Cleaning Extracted Text
The value returned by .text can include whitespace from the original XML. This is especially noticeable when the XML contains newlines and spaces around the content. When clean text is needed, use the string .strip() method after confirming that you have a valid element.
Mistakes in XML Extraction
Using .get() to read element text
Text content and attributes are separate. .get() accesses an attribute value, not the content between tags.
Fix:
Use name_element.text for content between the opening and closing tags.Using .text to read an attribute
The id value is stored as an attribute in the opening tag, not as the element's text content.
Fix:
Use root.get("id") to retrieve the attribute value.Accessing properties without checking find
If email does not exist, find returns None, and None does not provide a .text property.
Fix:
Check that email_element is not None before accessing .text or .get().Treating whitespace as part of the desired value
The original XML may contribute newlines and spaces to the returned text.
Fix:
Use phone_element.text.strip() when the extracted text needs whitespace removed.
Practice the Extraction Sequence
Given the XML string "<book code='b7'><title>Python Basics</title></book>", describe the three operations needed to parse it, locate title, and retrieve both the title text and the book code attribute.
Hints
- Start by converting the string with ET.fromstring.
- Use find with the title tag.
- Use .text for the title content and .get("code") for the attribute.
- Check the result of find before reading .text.
Parsing and extracting from one XML string
Extract the title text and code attribute from the book element.
Parse: Pass the XML string to ET.fromstring so it becomes an Element tree with book as the root.
Locate: Call root.find("title") to obtain the first matching title element.
Read text: After checking that the returned element is valid, read title_element.text to obtain Python Basics.
Read attribute: Call root.get("code") to obtain the code attribute value b7.
The extraction uses .text for title content and .get() for the code attribute.
Key Takeaways
- ET.fromstring converts an XML string into an Element tree whose root represents the outermost element.
- find locates the first element with a matching tag name and returns an Element object.
- Use .text for content between an element's opening and closing tags.
- Use .get() for attribute values stored in an element's opening tag.
- Check the result of find before accessing properties, and strip whitespace from text when necessary.
Key Takeaways
- Convert XML text into a navigable tree with ET.fromstring.
- Use find to locate the first matching element.
- Use .text for element content and .get() for attribute values.
- Check for None after find and use .strip() when extracted text contains unwanted whitespace.