Working with XML Namespaces
fromstring converts an XML string into a tree structure, enabling programmatic access to its data.
From Text to Tree
An XML document may begin as a plain string: a sequence of characters that is difficult to navigate directly. ElementTree provides fromstring to read that string and build a hierarchical tree of elements. The tree mirrors the nesting structure of the XML, giving your program an entry point for locating elements and extracting data.
The Root Element
Calling fromstring returns an Element object representing the root element. The root is the entry point to the entire tree. Once you have it, you can use its tag information and search methods to work with the XML structure.
Building a Small XML Tree
Suppose an XML string has a root element named person with child elements named name and phone. What does fromstring provide?
Start with text: The XML is initially just a plain string. At this stage, the program does not yet have a navigable element tree.
Parse the string: Pass the XML string to fromstring. The function reads the string and constructs a hierarchical tree that follows the XML nesting.
Receive the root: The result is an Element object for person. That object is the entry point for accessing the child elements.
The XML string has become a tree whose root Element is person, with the nested elements available through that tree.
Namespaces and Lookup Labels
Namespace-related XML can make element lookup depend on more than the short visible tag label. A prefix and a namespace URI are the two labels that learners commonly need to keep connected when reasoning about namespaced elements. The source material establishes that find locates the first element with a matching tag name, but it does not specify a namespace-specific lookup syntax. Therefore, the reliable rule in this lesson is to treat the exact tag used for lookup as important and to verify the returned Element before reading from it.
Finding an Element
The find method searches the XML tree for a matching tag name and returns the first matching Element object. That returned object becomes the next point of access: you can inspect its text or retrieve one of its attributes. If no matching tag exists, find returns None instead of an Element.
Searching for a Child
An XML tree has a person root with name and phone children. What happens when find searches for name?
Search: Call find with the tag name you want to locate.
Match: The method looks for an element with that matching tag name.
Use the result: The returned value is the matching Element object, which can then provide text or attribute data.
find returns the first matching Element. A search for a tag that is not present returns None.
Text and Attributes
XML stores element content and attribute data in different places. Text content appears between an element's opening and closing tags and is accessed through the .text property. Attribute data appears as key-value information in the opening tag and is accessed with .get(). An element can have text, attributes, both, or neither.
| Data | Where it lives | How to access it |
|---|---|---|
| Element text | Between the opening and closing tags | .text |
| Attribute value | As a key-value pair in the opening tag | .get() |
Choosing the Right Extraction Method
A located Element contains the text value Ada and has an attribute named role with the value developer. Which access method should be used for each value?
Read the element content: The value Ada is text between the element's opening and closing tags, so read it through .text.
Read the attribute: The value developer belongs to the role attribute in the opening tag, so retrieve it with .get().
Keep the sources separate: Do not use .text to retrieve an attribute, and do not use .get() as a replacement for reading element text.
Use .text for Ada and .get() for the role attribute value developer.
Safe Extraction
Assuming find always returns an Element
find returns None when no matching tag exists, and None does not provide the expected Element properties.
Fix:
Check that find returned a valid element before accessing .text or calling .get().Using .text to retrieve an attribute
Element text and attribute values are stored separately.
Fix:
Use .get() for an attribute value and .text for content between opening and closing tags.Returning unclean whitespace as meaningful text
The original formatting can become part of the returned text.
Fix:
Use .strip() when the application needs cleaned text.
Treat XML extraction as a guarded sequence: first create the tree with fromstring, then search with find, then verify that an Element was returned, and only then read .text or call .get(). This order prevents a missing element from causing an error during extraction.
Practice Check
Describe the extraction sequence for an XML string containing a namespaced element. State what fromstring provides, what find is expected to return when the tag matches, which property reads element text, which method reads an attribute, and what must be checked before reading from the result.
Hints
- Begin with the conversion from a plain string to an Element tree.
- Distinguish the result of a successful find from the None result of an unsuccessful search.
- Separate text between tags from key-value data in an opening tag.
Key Takeaways
- fromstring converts an XML string into a hierarchical tree and returns an Element for the root.
- find locates the first element with a matching tag name and returns None when no match exists.
- Use .text for content between an element's opening and closing tags.
- Use .get() for attribute values stored in an element's opening tag.
- Check the result of find before accessing properties, and use .strip() when extracted text contains unwanted whitespace.
Key Takeaways
- Convert XML text into a navigable Element tree with fromstring.
- Search the tree with find and handle the possibility that it returns None.
- Read element content with .text and attribute values with .get().
- Keep namespace labels and lookup requirements precise, because the supplied source does not define namespace-specific search syntax.
- Clean whitespace from extracted text with .strip() when necessary.