Concepts / Converting XML to JSON

Converting XML to JSON

fromstring converts an XML string into a tree structure, enabling programmatic access to its data.

  • Programming

From Text to Structure

An XML document stored in a variable begins as plain text: a sequence of characters. That text is difficult to navigate directly because your program has not yet represented the XML nesting as objects. ElementTree provides fromstring for this first step. It reads the XML string and builds a hierarchical tree of elements that mirrors the XML nesting structure.

Building the Element Tree

fromstringcontainscontainsXML string<person>...</person>personroot Elementnamechild Elementphonechild Element
What does the XML string become after fromstring, and how are its parent and child elements arranged?

The result of fromstring is an Element object representing the root element. If the outermost tag is person, the returned Element represents person and becomes the entry point to the entire tree. Its child elements remain arranged beneath it according to the original XML nesting. This arrangement is what allows later operations such as find to search the data programmatically.

python
Output
root represents the person Element, which is the root of the tree.

Finding and Reading Values

Once the XML string has become a tree, use find to locate an element by tag name. The find method returns the first element with a matching tag name within the tree. The returned Element is the object from which you read the data stored in that element.

childchildmatching tagpersonsearch begins herenameAdaname Elementfind("name")phone555-0100
How does find navigate the XML tree to locate the requested element?

Reading element text and an attribute

Extract the name text and the identifier attribute from a person XML tree.

Create the tree: Call ET.fromstring with the XML string. The returned Element represents the root element.

Locate the name: Call root.find("name"). The result is the first matching name Element.

Read text: Use name_element.text to retrieve the content between the name opening and closing tags.

Locate the profile: Call root.find("profile") to retrieve the profile Element.

Read the attribute: Use profile_element.get("id") to retrieve the value associated with the id attribute.

Element text is accessed with .text, while an attribute value is accessed with .get(). These are separate locations in an XML element.

python
Output
name is "Ada"
identifier is "p1"
XML locationPython accessPurpose
Content between opening and closing tags.textRetrieves element text
Key-value pair in the opening tag.get("attribute_name")Retrieves an attribute value

Safe Extraction Practice

A tag may be absent from a real XML document. When find cannot locate a matching tag, it returns None. None is not an Element, so it does not have .text or .get() properties. Check the result of find before accessing either property.

  • Reading text directly from the result of find without checking it.

    find returns None when the requested tag does not exist, and None does not have a .text property.

    Fix: Store the result, check that it is not None, and then read .text.

  • Using .text to retrieve an attribute.

    The .text property retrieves content between tags, while attributes are key-value pairs in the opening tag.

    Fix: Use profile_element.get("id") for the id attribute.

  • Assuming extracted text is free of formatting whitespace.

    Text can include newlines and spaces from the original XML.

    Fix: Use phone_element.text.strip() when surrounding whitespace should be removed.

MEDIUM

Given a tree whose root element is person, decide which operation is appropriate for each value: the text inside a name element, the id attribute on a profile element, and a possibly absent email element.

Hints
  • Use find to locate each element by tag name.
  • Use .text for content between tags.
  • Use .get() for an attribute.
  • Check the email result before reading a property.

Extraction Sequence

  1. Start with the XML data as a plain string.
  2. Call ET.fromstring to create an Element tree and obtain the root Element.
  3. Use find with a tag name to locate the first matching element.
  4. Check that the result is not None before reading its properties.
  5. Use .text for element content and .get() for attribute values.
  6. Use .strip() on text when whitespace from the original XML should be removed.

The central transformation is from unstructured XML text into a navigable hierarchical tree. After that transformation, find locates elements, .text reads element content, and .get() reads attributes.

Key Takeaways

  • ET.fromstring converts an XML string into an Element tree whose root represents the outermost element.
  • find locates the first element with a matching tag name.
  • The .text property reads content between an element's opening and closing tags.
  • The .get() method retrieves an attribute value from an element.
  • Always check for None after find, and strip extracted text when source whitespace should be removed.