Navigating Nested XML Structures
fromstring converts an XML string into a tree structure, enabling programmatic access to its data.
From Text to a Navigable Tree
An XML document held in a plain string is only a sequence of characters. Before your program can locate specific tags or extract values, it needs a structure that represents the XML nesting. ElementTree provides fromstring for this conversion. It reads the XML string and returns an Element object representing the root element, which becomes the entry point to the rest of the tree.
The central workflow is convert, locate, then read: convert the XML string with fromstring, locate an element with find, and read either its text or one of its attributes.
The Tree Created by fromstring
After conversion, the returned Element represents the XML root. Its tag identifies the root element, and the element provides access to the hierarchical structure formed by the nested XML elements. This is the important change in state: the data is no longer being treated as unstructured text; it is represented as an XML tree that program code can navigate.
Identifying the Root Element
An XML string has a root element named person and nested elements named name and phone. What does fromstring provide for the next navigation step?
Convert: Pass the XML string to fromstring. The function parses the string and returns an Element object for the root element.
Identify: The returned Element has the root tag person. This object is now the entry point to the complete XML tree.
Navigate: Use the returned root Element when searching for nested tags with find.
fromstring changes the XML from plain text into a hierarchical Element structure that code can navigate.
Finding a Nested Element
Once you have the root Element, use find to locate an element by tag name. The method returns the first element with a matching tag name within the tree. The result is itself an Element object, so it can provide access to the selected element's text and attributes.
Think of find as producing a selected Element rather than producing the final value immediately. After locating a tag, you perform a second operation to extract what that element contains. For content between the opening and closing tags, use .text. For a value written as an attribute in the opening tag, use .get().
Reading Text and Attributes
Extracting Two Kinds of XML Data
Consider the XML string <book code="B7"><title>Learning XML</title></book>. After converting it into a tree, locate the title element and obtain the title text and the book code.
Convert the string: Use fromstring to create the root Element for the book tree.
Locate the title: Use find with the tag name title. The result is the Element containing the text Learning XML.
Read element content: Use the selected title Element's .text property to obtain Learning XML.
Read the attribute: Use the root book Element's .get() method with the attribute name code to obtain B7.
The title value comes from .text, while the book code comes from .get(). These are separate extraction operations because XML text and XML attributes are stored differently.
Text may include whitespace from the original XML formatting, including newlines and spaces. When the extracted value should not retain that surrounding whitespace, apply the .strip() method to the text value.
Missing Elements and Unclean Text
Trying to navigate the original XML string as though it were already a tree.
A plain string is only a sequence of characters. ElementTree navigation requires the structure returned by fromstring.
Fix:
Convert the XML string first and use the returned root Element as the entry point.Using .text to retrieve an attribute.
Text content and attributes are separate kinds of XML data.
Fix:
Use .text for content between tags and .get() for a named attribute.Accessing .text or .get() without checking find's result.
find returns None when it cannot locate a matching tag. Accessing properties on that missing result can cause an error.
Fix:
Store the result, check that it is a valid element, and only then access .text or .get().Assuming extracted text is free of formatting whitespace.
Whitespace from the original XML can be included in the element's text.
Fix:
Use .strip() when surrounding whitespace should be removed.
Treat the result of find as something that must be validated before use. A missing tag is a normal possibility in real XML documents, so check that the result is present before reading its text or attributes.
Apply the Extraction Pattern
For the XML string <person><name>Ada</name><phone> 555-0100 </phone></person>, describe the operations needed to convert the string into a tree, locate the phone element, and obtain clean phone text. Then explain how the process would differ if the phone value were stored as an attribute instead of between the tags.
Hints
- Begin with fromstring.
- Use find with the phone tag.
- Read the selected element with .text.
- Use .strip() if the surrounding spaces should be removed.
- Use .get() when the value is an attribute rather than element text.
- The dependable pattern is to convert, locate, and read. Convert the XML string into a tree with fromstring. Locate the first matching tag with find. Read content between tags with .text or retrieve an opening-tag attribute with .get(). Check the result of find before using it, and strip surrounding whitespace when necessary.
Key Takeaways
- fromstring converts XML text into a hierarchical tree and returns an Element representing the root.
- find locates the first element with a matching tag name and returns an Element or None.
- The .text property reads content between an element's opening and closing tags.
- The .get() method retrieves a named attribute value from an element.
- Check find results before accessing properties, and use .strip() when extracted text contains unwanted whitespace.