Concepts / Working with XML Attributes and Namespaces

Working with XML Attributes and Namespaces

ET.fromstring() converts an XML string into a tree structure of Element objects, making the hierarchical data queryable.

  • Programming

From Text to Tree

XML often begins as a sequence of characters, but useful XML processing depends on its structure. ET.fromstring() converts an XML string into a tree structure made of Element objects. After conversion, the XML is no longer handled only as plain text: the resulting Element object represents the root and provides a structured way to query the hierarchy.

ET.fromstring()containscontainscontainsperson XML<person>...</person>personroot ElementnameChuckphonetext contentemailhide=yes
How does a linear XML string become nested Element objects, and which elements contain which child elements?

The Conversion Step

The conversion begins with an XML string. Calling ET.fromstring(data) parses that string and returns an Element object representing the root of the XML tree. In the source example, the root is person. The returned object contains the complete hierarchy, including the person element and its name, phone, and email child elements.

What do you think happens?

Before reading further, predict what tree represents after ET.fromstring(data) runs.

  • The original XML characters only
  • An Element object representing the root and its hierarchy
  • The text inside the first child element
  • The value of the first attribute
Reveal answer

Answer: An Element object representing the root and its hierarchy

ET.fromstring() converts the XML string into a tree structure of Element objects. The returned tree object represents the root element and contains the hierarchy beneath it.

python

After the assignment, tree is an Element object for person rather than a string. You interact with this structured object through methods and properties. The XML hierarchy is now available for retrieval instead of requiring you to manually search the original character sequence.

Finding an Element

Once the XML has been converted, find() searches the tree for the first element matching a given tag name. For example, tree.find('name') retrieves the name element from the person tree. The result is another Element object, so you can then inspect its text or attributes.

childchildchildfind('name')personroot ElementnameChucknamefirst matching Elementphonephone textemailhide=yes
How does find() move through the element hierarchy to locate a specific element?

name_element = tree.find('name') name_text = name_element.text email_element = tree.find('email') hidden_value = email_element.get('hide')

Output
name_text: Chuck
hidden_value: yes

Text and Attribute Values

An Element can hold text content and attribute values, and the two are retrieved differently. Use the .text property for the characters inside an element. Use .get(attribute_name) for the value of a named attribute. In the source example, tree.find('name').text produces Chuck, while tree.find('email').get('hide') produces yes.

text content.get('hide')email Element.textemail contenthideyes
Where are an element's text content and attribute key-value pairs accessed in the resulting Element object?

Namespace-Aware Names

Namespaces are part of the name information attached to XML elements. A namespace declaration connects a prefix or URI with the elements it qualifies, and that qualification affects element lookup. When an XML tree includes qualified names, treat the namespace-qualified name as part of what must be matched during retrieval rather than assuming that the visible local tag name alone is sufficient.

qualifiesmust be matchednamespace declarationprefix or URIqualified elementnamespace + tagfind() lookupmatching qualified name
How does a namespace declaration connect a prefix or URI to the elements it qualifies, and how does that affect element lookup?

A Complete Retrieval Path

Read a Name and an Attribute

Starting with an XML string containing person, name, phone, and email elements, retrieve the name text and the email hide attribute.

Convert: Pass the XML string to ET.fromstring(). The result is an Element object representing the person root and its hierarchy.

Locate: Call tree.find('name') and tree.find('email') to retrieve the first matching elements.

Read text: Access the name Element's .text property to retrieve Chuck.

Read attribute: Call the email Element's .get('hide') method to retrieve yes.

The conversion, lookup, and extraction steps work together: the name text is Chuck and the hide attribute value is yes.

Keep the stages conceptually separate: first convert the string, then find the required Element, then choose .text for element content or .get(attribute_name) for an attribute value. This makes it easier to see whether a problem occurred during conversion, lookup, or extraction.

Mistakes to Avoid

  • Treating the result of ET.fromstring() as if it were still the original string.

    The result is an Element object representing the root and its hierarchy.

    Fix: Use Element methods such as find(), and properties or methods such as .text and .get().

  • Using .text to retrieve an attribute.

    .text retrieves content inside the element, while .get(attribute_name) retrieves an attribute value.

    Fix: Use email_element.get('hide') for the hide attribute.

  • Using .get() when the desired data is the element's text.

    Chuck is text content inside the name element, not the value of a named attribute.

    Fix: Use name_element.text.

  • Ignoring preserved whitespace.

    Whitespace in XML text is preserved.

    Fix: Use .strip() when surrounding whitespace should be removed.

  • Searching for only an unqualified local tag when the XML uses a qualified name.

    Namespace qualification affects element lookup.

    Fix: Inspect the XML's namespace declaration and match the qualified name during retrieval.

Practice the Sequence

EASY

Explain the retrieval sequence for this XML data: a person root contains name, phone, and email elements, and the email element has a hide attribute. Identify which operation converts the XML, which operation locates name, which expression retrieves the name content, and which expression retrieves the hide value.

Hints
  • The conversion operation is ET.fromstring().
  • Use find() with the tag name to retrieve an Element.
  • Use .text for content inside an element.
  • Use .get(attribute_name) for an attribute value.
  1. ET.fromstring() converts an XML string into a tree of Element objects.
  2. The returned root Element represents the XML hierarchy beneath it.
  3. find() searches for the first element matching a tag name.
  4. .text retrieves element text, while .get(attribute_name) retrieves an attribute value.
  5. Whitespace is preserved, and namespace qualification can affect how an element must be located.

Key Takeaways

  • An XML string becomes queryable after ET.fromstring() converts it into an Element tree.
  • The root Element contains the XML hierarchy, including its child elements.
  • Use find() to retrieve the first matching element.
  • Use .text for element content and .get() for attribute values.
  • Account for preserved whitespace and namespace qualification when interpreting or locating XML data.