Working with XML Attributes and Namespaces
ET.fromstring() converts an XML string into a tree structure of Element objects, making the hierarchical data queryable.
From Text to Tree
XML often begins as a sequence of characters, but useful XML processing depends on its structure. ET.fromstring() converts an XML string into a tree structure made of Element objects. After conversion, the XML is no longer handled only as plain text: the resulting Element object represents the root and provides a structured way to query the hierarchy.
The Conversion Step
The conversion begins with an XML string. Calling ET.fromstring(data) parses that string and returns an Element object representing the root of the XML tree. In the source example, the root is person. The returned object contains the complete hierarchy, including the person element and its name, phone, and email child elements.
What do you think happens?
Before reading further, predict what tree represents after ET.fromstring(data) runs.
Reveal answer
Answer: An Element object representing the root and its hierarchy
ET.fromstring() converts the XML string into a tree structure of Element objects. The returned tree object represents the root element and contains the hierarchy beneath it.
After the assignment, tree is an Element object for person rather than a string. You interact with this structured object through methods and properties. The XML hierarchy is now available for retrieval instead of requiring you to manually search the original character sequence.
Finding an Element
Once the XML has been converted, find() searches the tree for the first element matching a given tag name. For example, tree.find('name') retrieves the name element from the person tree. The result is another Element object, so you can then inspect its text or attributes.
name_element = tree.find('name') name_text = name_element.text email_element = tree.find('email') hidden_value = email_element.get('hide')
name_text: Chuck
hidden_value: yesText and Attribute Values
An Element can hold text content and attribute values, and the two are retrieved differently. Use the .text property for the characters inside an element. Use .get(attribute_name) for the value of a named attribute. In the source example, tree.find('name').text produces Chuck, while tree.find('email').get('hide') produces yes.
Namespace-Aware Names
Namespaces are part of the name information attached to XML elements. A namespace declaration connects a prefix or URI with the elements it qualifies, and that qualification affects element lookup. When an XML tree includes qualified names, treat the namespace-qualified name as part of what must be matched during retrieval rather than assuming that the visible local tag name alone is sufficient.
A Complete Retrieval Path
Read a Name and an Attribute
Starting with an XML string containing person, name, phone, and email elements, retrieve the name text and the email hide attribute.
Convert: Pass the XML string to ET.fromstring(). The result is an Element object representing the person root and its hierarchy.
Locate: Call tree.find('name') and tree.find('email') to retrieve the first matching elements.
Read text: Access the name Element's .text property to retrieve Chuck.
Read attribute: Call the email Element's .get('hide') method to retrieve yes.
The conversion, lookup, and extraction steps work together: the name text is Chuck and the hide attribute value is yes.
Keep the stages conceptually separate: first convert the string, then find the required Element, then choose .text for element content or .get(attribute_name) for an attribute value. This makes it easier to see whether a problem occurred during conversion, lookup, or extraction.
Mistakes to Avoid
Treating the result of ET.fromstring() as if it were still the original string.
The result is an Element object representing the root and its hierarchy.
Fix:
Use Element methods such as find(), and properties or methods such as .text and .get().Using .text to retrieve an attribute.
.text retrieves content inside the element, while .get(attribute_name) retrieves an attribute value.
Fix:
Use email_element.get('hide') for the hide attribute.Using .get() when the desired data is the element's text.
Chuck is text content inside the name element, not the value of a named attribute.
Fix:
Use name_element.text.Ignoring preserved whitespace.
Whitespace in XML text is preserved.
Fix:
Use .strip() when surrounding whitespace should be removed.Searching for only an unqualified local tag when the XML uses a qualified name.
Namespace qualification affects element lookup.
Fix:
Inspect the XML's namespace declaration and match the qualified name during retrieval.
Practice the Sequence
Explain the retrieval sequence for this XML data: a person root contains name, phone, and email elements, and the email element has a hide attribute. Identify which operation converts the XML, which operation locates name, which expression retrieves the name content, and which expression retrieves the hide value.
Hints
- The conversion operation is ET.fromstring().
- Use find() with the tag name to retrieve an Element.
- Use .text for content inside an element.
- Use .get(attribute_name) for an attribute value.
- ET.fromstring() converts an XML string into a tree of Element objects.
- The returned root Element represents the XML hierarchy beneath it.
- find() searches for the first element matching a tag name.
- .text retrieves element text, while .get(attribute_name) retrieves an attribute value.
- Whitespace is preserved, and namespace qualification can affect how an element must be located.
Key Takeaways
- An XML string becomes queryable after ET.fromstring() converts it into an Element tree.
- The root Element contains the XML hierarchy, including its child elements.
- Use find() to retrieve the first matching element.
- Use .text for element content and .get() for attribute values.
- Account for preserved whitespace and namespace qualification when interpreting or locating XML data.