Concepts / Fetching Data from Web APIs

Fetching Data from Web APIs

ET.fromstring converts an XML string into a tree structure, transforming flat text into a navigable, queryable object.

  • Programming

From Response Text to Usable Data

When XML data arrives from a web service, a file, or another system, it arrives as a flat string of characters. The string contains tags, nesting, attributes, and values, but Python cannot conveniently search that raw text as a tree. ElementTree provides the parsing step that turns the string into a structured object.

The central workflow is parse first, then search: pass the XML string to ET.fromstring, use find to locate an element, and then read its text or an attribute.

ET.fromstringfind.textXML responseflat stringElementTreehierarchical objectname elementlocated by findAlicetext value
How does XML data travel from a web API response through parsing and element lookup to the extracted value?

Building the XML Tree

ET.fromstring converts an XML string into a tree structure. ElementTree parses the XML tags, nesting, and attributes, then builds an in-memory object representing the document. The result is no longer merely a sequence of characters: it is a nested hierarchy that Python code can navigate and query.

ET.fromstringcontainscontainsperson XML<person>...</person>personroot elementnameAliceemailhide="true"
How does ET.fromstring change a linear XML string into nested elements that can be navigated and queried?

The outermost XML tag becomes the root element. Elements inside it become child elements. Each element can have child elements, text content, and attributes. This structure is what makes later operations such as locating name or email meaningful.

containscontainscontainspersonrootnameAliceemailhide="true"contactchild
What contains what in the parsed XML tree, and how are the root, child, and nested elements connected?

Finding and Reading Elements

After parsing, find searches the tree for the first element matching a supplied tag name. It returns an element object when a match is found, or None when no matching element is found. The returned object is not the text itself, so you read its text separately with .text.

find("name")find("email")personsearch begins herenameAliceemailhide="true"
How does find move from the root to locate a specific element in the tree?
python
Output
Alice
true

The first lookup returns the name element object. Accessing name_element.text reads the content between the name tags. The second lookup returns the email element object. Accessing email_element.get("hide") reads the value of the hide attribute. Text content and attributes are separate parts of an XML element and are accessed separately.

text contenthide attributeemailelement objectalice@example.com.texttrue.get("hide")
Where are an element's text content and attributes stored, and how are they accessed separately?

Tracing the Complete Operation

What do you think happens?

What will the two print statements display after the XML is parsed and searched?

  • The name element object and the email element object
  • Alice and true
  • name and hide
Reveal answer

Answer: Alice and true

find returns element objects. The .text access extracts Alice from the name element, and .get("hide") extracts the true value from the email element's attribute.

Following the XML Through Four States

Explain how an XML string containing person data becomes the printed values Alice and true.

1. Receive the string: The XML begins as flat text containing the person, name, and email tags, along with the hide attribute.

2. Build the tree: ET.fromstring parses the flat XML string and returns a tree object whose root is the person element.

3. Locate elements: root.find("name") locates the name element, while root.find("email") locates the email element.

4. Extract values: The name element's .text is Alice. The email element's .get("hide") retrieves the attribute value true.

The parsed tree makes the XML queryable: text is obtained with .text, and the attribute is obtained with .get("hide").

Mistakes in XML Queries

  • Treating the result of find as though it were already a string

    find returns an element object or None, not the element's text value.

    Fix: Use name_element.text to read the text content.

  • Reading an attribute with .text

    The email text and the hide attribute are separate parts of the element.

    Fix: Use email_element.get("hide") to retrieve the hide attribute.

  • Changing the capitalization of a tag or attribute name

    XML tag names and attribute names are case-sensitive.

    Fix: Match the exact spelling used in the XML, such as "name" and "hide".

  • Assuming every search succeeds

    find returns None when no matching element is found.

    Fix: Handle the None case before attempting to access element data.

Keep the two stages visible in your code: first create the tree with ET.fromstring, then search the tree with find, and finally choose the correct extraction operation. Use .text for content between tags and .get with the exact attribute name for an attribute.

Practice the Parse-Search-Extract Pattern

MEDIUM

Suppose an XML response contains a root element named person, a child element named name with the text Jordan, and a child element named email with the attribute hide set to false. Describe the three operations needed to obtain Jordan and false.

Hints
  • Use ET.fromstring for the initial XML string.
  • Use find with the exact tag names.
  • Use .text for name and .get("hide") for the attribute.
  1. Start with the XML string received from the web service or another source.
  2. Pass the string to ET.fromstring to create the tree and obtain its root element.
  3. Call find with the exact tag name to locate the desired element.
  4. Read element text with .text or read an attribute with .get("attribute_name").
  5. Check for None when a requested element might not exist.

Key Takeaways

  1. An XML response begins as a flat string of characters.
  2. ET.fromstring parses that string into a hierarchical ElementTree object.
  3. The root element represents the outermost tag, and nested elements form the parent-child structure.
  4. find locates the first matching element and returns an element object or None.
  5. Use .text for element content and .get("attribute_name") for attributes, while preserving exact case.

Key Takeaways

  • ET.fromstring transforms flat XML text into a navigable tree.
  • The parsed tree contains a root element and nested child elements.
  • find searches for the first element with a matching tag name.
  • Element text is accessed with .text, while attributes are accessed with .get('attribute_name').
  • XML names are case-sensitive, and find may return None when no match exists.