Handling XML Parsing Errors
fromstring converts an XML string into a tree structure, enabling programmatic access to its data.
From Text to Navigable Data
An XML document may begin as a plain string: a sequence of characters containing opening tags, closing tags, text, and attributes. That string is not yet a structure your program can conveniently navigate. The fromstring function reads the XML string and builds a hierarchical tree that mirrors the nesting of the XML elements.
The Root Element After Parsing
Inspecting a Parsed Person Document
A program receives the XML string <person><name>Ada</name><phone>555-0100</phone></person>. What does fromstring provide?
Parse the string: Applying fromstring to the XML string creates an Element object representing the root element.
Identify the entry point: The root Element is person, so its tag property is person. This Element is the entry point to the entire tree.
Use the tree: From the person element, the program can search for nested elements such as name and phone.
fromstring changes the XML from plain text into a tree whose root Element is person.
The result of fromstring is an Element object, not merely another string. Its tag identifies the element's name, and the root Element provides access to the rest of the parsed tree.
When Parsing Cannot Build a Tree
The useful result of fromstring is a tree. If the supplied XML is invalid or malformed, that successful tree-building step cannot be completed. This is the first point to investigate when XML processing fails: determine whether the input is valid enough for fromstring to produce a root Element before attempting to search it.
Finding and Reading Element Data
Once fromstring has produced a tree, use find with a tag name to locate the first element with that matching tag name. The returned value is an Element object when the tag exists. You can then read the element's content through .text or retrieve one of its attributes with .get().
| XML data location | Python access | What it provides |
|---|---|---|
| Between an element's opening and closing tags | .text | The element's text content |
| As a key-value pair in the opening tag | .get() | The value of a named attribute |
Text content and attributes are separate kinds of XML data.
Extracting Text and an Attribute
Consider the XML element <person status="active"><name>Ada</name></person>. How would the tree's data be interpreted?
Locate name: find with the tag name name returns the name Element.
Read text: The name Element's .text value is Ada because Ada appears between the opening and closing name tags.
Read the attribute: The person Element's .get() method can retrieve the value active for the status attribute.
Keep the locations distinct: Ada is element text, while active is an attribute value. The two values come from different parts of the XML.
Use .text for content between tags and .get() for a named attribute in an opening tag.
Missing Elements and Safe Checks
Assuming every find call returns an Element
find returns None when the requested tag does not exist. None does not provide the Element properties you intended to read.
Fix:
Check whether find returned a valid element before accessing .text or calling .get().Using .text to retrieve an attribute
Attributes and text content are stored in different places in XML.
Fix:
Use .text for content between tags and .get() for a named attribute.Treating whitespace as part of the meaningful value
The .text property can include whitespace present in the source XML.
Fix:
Use .strip() when the extracted text should have surrounding whitespace removed.Searching before confirming that parsing produced a tree
The navigation methods depend on a successfully created Element tree.
Fix:
Handle the parsing problem first, then search the tree only when a usable root Element is available.
Treat XML extraction as a guarded sequence: parse the string, confirm that a tree is available, use find, check the returned Element, and only then access .text or .get(). This prevents a missing tag from turning into an attempt to read properties from None.
Practice the Extraction Sequence
A document contains a root element named person, a nested element named phone, and a status attribute on person. Describe the sequence you would use to parse the XML, locate phone, obtain its text, and retrieve status. Then explain what you would check before reading phone's properties.
Hints
- Begin with fromstring because the input starts as a plain XML string.
- Use find with the tag name phone.
- Use .text for the phone content and .get() for the status attribute.
- find may return None, so check the result before reading its properties.
What do you think happens?
If find searches for a tag that does not exist, what should you expect before trying to read .text?
Reveal answer
Answer: None, which must be checked
The find method returns None when the requested tag does not exist. Accessing .text or .get() without checking can cause an error.
The Reliable Workflow
- Start with the XML string and pass it to fromstring to build a hierarchical tree.
- Use the returned root Element as the entry point to the tree.
- Call find with a tag name to locate the first matching element.
- Check that find returned a valid Element rather than None.
- Read element content with .text, cleaning surrounding whitespace with .strip() when needed.
- Read attribute values with .get(), remembering that attributes and text are separate.
The central idea is to move carefully from representation to navigation to extraction. fromstring turns plain XML text into an Element tree. find locates an element in that tree. .text reads content between tags, while .get() reads an attribute value. Checking each stage makes missing or malformed data easier to handle.
Key Takeaways
- fromstring converts an XML string into a hierarchical tree and returns an Element representing the root.
- find locates the first element with a matching tag name.
- The .text property reads content between an element's opening and closing tags.
- The .get() method retrieves a named attribute value.
- Always check both the parsing stage and the result of find before accessing element properties.