Parsing XML: Reading and Extracting Data
An XML element is a complete unit consisting of an opening tag, content, and a closing tag; it can also be represented as a self-closing tag if empty.
Reading XML as Structured Data
When you encounter XML in a configuration file, an API response, or a data export, you are reading a structured arrangement of elements, attributes, and nested relationships. XML focuses on what information means rather than how it should look on a screen. Its strict, predictable structure helps both humans and machines understand the data.
The Three-Part Element
An XML element is a complete unit consisting of an opening tag, content, and a closing tag. The opening tag begins with a less-than symbol, contains the element name, and ends with a greater-than symbol. The closing tag uses the same element name but places a forward slash before that name. Everything between the two tags is the element's content.
Identifying an element's parts
Read the element <email>contact@example.com</email> and identify its three components.
Opening tag: The element begins with <email>, which names the element.
Content: The text contact@example.com appears between the two tags, so it is the element's content.
Closing tag: The element ends with </email>, which matches the opening name and includes the required forward slash.
The opening tag is <email>, the content is contact@example.com, and the closing tag is </email>.
Metadata in Opening Tags
An attribute is optional metadata attached to an opening tag. It has the form name="value": a name, an equals sign, and a quoted value. Multiple attributes can appear in one opening tag, separated by spaces. Attributes describe or modify the element; they are not the element's nested content.
Use attributes for metadata, properties, or classifications such as type, id, status, or version. Use nested elements for primary content or for data that may have its own structure or attributes.
Empty Elements
An element with no text and no child elements is empty. It can be written with a traditional opening tag and closing tag, or with a self-closing tag that combines both parts and ends with a forward slash before the closing bracket. Both forms are valid and mean the same thing.
Parent and Child Elements
XML elements can be nested inside other elements, creating a tree-like hierarchy. An element that contains other elements is a parent, and the elements inside it are children. The outermost element containing all the others is the root element. A document has a single root element containing the rest of its elements.
Tracing the document structure
Read the person document and identify the root, its children, and the content of each child.
Find the outermost element: The person opening and closing tags surround the entire document, so person is the root.
List direct children: The elements directly inside person are name, phone, and email.
Separate content from metadata: Chuck is the content of name. The phone number is the content of phone, while type="intl" is metadata. email has no content and is self-closing.
Interpret the relationship: The hierarchy shows that the name, phone, and email belong to the person element.
The root is person. Its children are name, phone, and email; name and phone hold text, phone also has an attribute, and email is empty.
Common Reading Mistakes
Treating an attribute as the element's main content
type="intl" describes the phone element, while the phone number between the tags is the primary content.
Fix:
Read the attribute as metadata and the text between the tags as content.Assuming an empty element cannot have attributes
An empty element has no text or child elements, but it can still carry metadata.
Fix:
Check the opening or self-closing tag for optional name="value" attributes.Confusing a child element with ordinary text
name is a nested element and therefore a child of person, not merely text belonging to the parent.
Fix:
Look for a complete opening-and-closing pair inside the parent.Overlooking the root element
The outermost element establishes the top of the XML hierarchy and contains the other elements.
Fix:
Find the element whose opening tag begins the structure and whose matching closing tag surrounds the others.
Practice: Classify the Parts
Consider this XML structure: <person><name>Chuck</name><phone type="intl">+1-555-0100</phone><email /></person>. Identify the root element, the three child elements, the attribute, the text content, and the empty element.
Hints
- Start with the outermost matching pair of tags.
- Then inspect the complete elements directly inside it.
- Separate values inside tags from name="value" metadata in an opening tag.
What do you think happens?
Which element is empty in the structure: person, name, phone, or email?
Reveal answer
Answer: email
email is written as <email />, so it has no text content and no child elements. It is still a valid child of person.
Key Takeaways
- A standard XML element has an opening tag, content, and a matching closing tag.
- Attributes are optional name="value" metadata attached to an opening tag.
- An empty element can use paired empty tags or a compact self-closing tag; both forms mean the same thing.
- Nested elements create parent-child relationships in a tree-like hierarchy with one root element.
- Use attributes for metadata and nested elements for primary content or structured child data.
Key Takeaways
- Read every XML element by locating its opening tag, content, and closing tag.
- Interpret attributes as metadata attached to opening tags, not as nested content.
- Recognize both paired empty tags and self-closing tags as valid representations of empty elements.
- Use nesting to understand which elements are parents, which are children, and which element is the root.
- Distinguish descriptive metadata from primary data when deciding between attributes and nested elements.