Concepts / Introduction to XML Structure

Introduction to XML Structure

ET.fromstring converts an XML string into a tree structure, transforming flat text into a navigable, queryable object.

  • Programming

From Flat Text to a Tree

XML data often arrives from a web service, a file, or another system as a flat string of characters. Although that string contains tags, nesting, text, and attributes, Python cannot conveniently search it as a hierarchy until it has been parsed. ElementTree, Python's built-in XML library, provides fromstring for this transformation.

ET.fromstring reads the XML string, interprets its tags, nesting, and attributes, and builds an in-memory tree object. The outermost tag becomes the root element. Elements inside it become child elements, so the parsed result represents the XML document as a nested hierarchy that Python code can query.

fromstringcontainscontainsperson XML<person>...</person>nametext contentpersonroot elementemailhide attribute
How does the nested XML text become a tree of parent and child elements that can be navigated?

Reading Parent and Child Relationships

After parsing, the XML is represented as a hierarchy. The root is the outermost element. An element may contain child elements, text content, and attributes. In a document describing a person, the person element can contain children such as name and email. Elements at the same level are siblings because they share the same parent.

parent ofparent ofpersonrootnametextemailhide
What contains what in the XML structure, and how are elements connected to their parent and sibling elements?

Parsing is the first step and searching is the second. fromstring creates the structure; methods such as find work on that parsed structure.

Finding an Element by Tag

The find method searches the tree for the first element matching a supplied tag name. It returns an element object when a match is found. If no matching element exists, it returns None. This means find does not directly return the element's text; it returns the element so that your code can then inspect its text or attributes.

find("name")siblingpersonrootnamefirst matching elementemaildifferent tag
How does find move through the XML hierarchy to locate a specific element?
python

In this example, fromstring parses the XML string into a tree represented by root. Calling root.find("name") searches that tree for the first name element. Calling root.find("email") searches for the first email element. The variables name_element and email_element refer to the located element objects, not yet to their text values.

Reading Text and Attributes

Extracting Person Details

Parse a person XML string, locate the name and email elements, and extract the name text and the email hide attribute.

Parse: Pass the XML string to ET.fromstring so ElementTree builds a hierarchical, queryable representation.

Locate: Use find("name") and find("email") to locate the first elements with those tag names.

Read text: Use the name element's .text property to access the text inside the name tag.

Read attribute: Use the email element's .get("hide") method to retrieve the value of the hide attribute.

The name text is Ada, and the email hide attribute is yes.

name_text = name_element.text hide_value = email_element.get("hide") print(name_text) print(hide_value)

textattributeemailelement objectada@example.com.textyes.get("hide")
Where are an element's text value and attributes located in the parsed tree, and how are they accessed?

Checking Search Results

  • Treating find as though it returns text directly.

    find returns an element object or None, not a string.

    Fix: Use name.text after locating the element.

  • Ignoring the possibility of None.

    find returns None when no element matches the requested tag.

    Fix: Handle the no-match case before trying to read the result.

  • Changing the spelling or capitalization of a tag or attribute.

    XML tag and attribute names are case-sensitive.

    Fix: Match the exact spelling used in the XML.

Practice the Parse-and-Search Process

EASY

Given the XML string <book><title>XML Basics</title><format type="digital">guide</format></book>, describe the tree created by ET.fromstring. Then identify which expression would retrieve the title text and which expression would retrieve the format element's type attribute.

Hints
  • The outermost book tag is the root element.
  • Use find with the exact tag name to locate an element.
  • Use .text for text content and .get("type") for the type attribute.
  1. The XML string is the starting text form, not yet a convenient searchable hierarchy. ET.fromstring parses it into a tree with a root and nested child elements. find searches that tree for the first matching tag and returns an element object or None. After locating an element, .text reads its text content and .get("attribute_name") retrieves an attribute value.

Key Takeaways

  • ET.fromstring transforms an XML string into a hierarchical tree structure.
  • The root is the outermost element, and nested elements form parent-child relationships.
  • find searches for the first element matching an exact tag name and returns an element object or None.
  • Use .text for element text and .get("attribute_name") for an attribute value.
  • XML tag and attribute names are case-sensitive.