Concepts / Parsing XML: Reading and Extracting Data

Parsing XML: Reading and Extracting Data

An XML element is a complete unit consisting of an opening tag, content, and a closing tag; it can also be represented as a self-closing tag if empty.

  • Programming

Reading XML as Structured Data

When you encounter XML in a configuration file, an API response, or a data export, you are reading a structured arrangement of elements, attributes, and nested relationships. XML focuses on what information means rather than how it should look on a screen. Its strict, predictable structure helps both humans and machines understand the data.

The Three-Part Element

An XML element is a complete unit consisting of an opening tag, content, and a closing tag. The opening tag begins with a less-than symbol, contains the element name, and ends with a greater-than symbol. The closing tag uses the same element name but places a forward slash before that name. Everything between the two tags is the element's content.

containsends at<name>opening tagChuckcontent</name>closing tag
Which text is the opening tag, which is the content, and which is the closing tag?
xml

Identifying an element's parts

Read the element <email>contact@example.com</email> and identify its three components.

Opening tag: The element begins with <email>, which names the element.

Content: The text contact@example.com appears between the two tags, so it is the element's content.

Closing tag: The element ends with </email>, which matches the opening name and includes the required forward slash.

The opening tag is <email>, the content is contact@example.com, and the closing tag is </email>.

Metadata in Opening Tags

An attribute is optional metadata attached to an opening tag. It has the form name="value": a name, an equals sign, and a quoted value. Multiple attributes can appear in one opening tag, separated by spaces. Attributes describe or modify the element; they are not the element's nested content.

opening tag containsuseselement may containphoneelement nametypeattribute name"intl"quoted valuenumberelement content
Where does an attribute attach, and how is its metadata kept separate from the element's content?
xml

Use attributes for metadata, properties, or classifications such as type, id, status, or version. Use nested elements for primary content or for data that may have its own structure or attributes.

Empty Elements

An element with no text and no child elements is empty. It can be written with a traditional opening tag and closing tag, or with a self-closing tag that combines both parts and ends with a forward slash before the closing bracket. Both forms are valid and mean the same thing.

same meaning<email></email>opening and closing tags<email />one self-closing tag
How do the two valid forms represent an element with no content?
xml

Parent and Child Elements

XML elements can be nested inside other elements, creating a tree-like hierarchy. An element that contains other elements is a parent, and the elements inside it are children. The outermost element containing all the others is the root element. A document has a single root element containing the rest of its elements.

containscontainscontainspersonrootnamechildphonechildemailchild
How do nested elements fit inside one another to form the document's hierarchy?
xml

Tracing the document structure

Read the person document and identify the root, its children, and the content of each child.

Find the outermost element: The person opening and closing tags surround the entire document, so person is the root.

List direct children: The elements directly inside person are name, phone, and email.

Separate content from metadata: Chuck is the content of name. The phone number is the content of phone, while type="intl" is metadata. email has no content and is self-closing.

Interpret the relationship: The hierarchy shows that the name, phone, and email belong to the person element.

The root is person. Its children are name, phone, and email; name and phone hold text, phone also has an attribute, and email is empty.

Common Reading Mistakes

  • Treating an attribute as the element's main content

    type="intl" describes the phone element, while the phone number between the tags is the primary content.

    Fix: Read the attribute as metadata and the text between the tags as content.

  • Assuming an empty element cannot have attributes

    An empty element has no text or child elements, but it can still carry metadata.

    Fix: Check the opening or self-closing tag for optional name="value" attributes.

  • Confusing a child element with ordinary text

    name is a nested element and therefore a child of person, not merely text belonging to the parent.

    Fix: Look for a complete opening-and-closing pair inside the parent.

  • Overlooking the root element

    The outermost element establishes the top of the XML hierarchy and contains the other elements.

    Fix: Find the element whose opening tag begins the structure and whose matching closing tag surrounds the others.

Practice: Classify the Parts

EASY

Consider this XML structure: <person><name>Chuck</name><phone type="intl">+1-555-0100</phone><email /></person>. Identify the root element, the three child elements, the attribute, the text content, and the empty element.

Hints
  • Start with the outermost matching pair of tags.
  • Then inspect the complete elements directly inside it.
  • Separate values inside tags from name="value" metadata in an opening tag.

What do you think happens?

Which element is empty in the structure: person, name, phone, or email?

  • person
  • name
  • phone
  • email
Reveal answer

Answer: email

email is written as <email />, so it has no text content and no child elements. It is still a valid child of person.

Key Takeaways

  1. A standard XML element has an opening tag, content, and a matching closing tag.
  2. Attributes are optional name="value" metadata attached to an opening tag.
  3. An empty element can use paired empty tags or a compact self-closing tag; both forms mean the same thing.
  4. Nested elements create parent-child relationships in a tree-like hierarchy with one root element.
  5. Use attributes for metadata and nested elements for primary content or structured child data.

Key Takeaways

  • Read every XML element by locating its opening tag, content, and closing tag.
  • Interpret attributes as metadata attached to opening tags, not as nested content.
  • Recognize both paired empty tags and self-closing tags as valid representations of empty elements.
  • Use nesting to understand which elements are parents, which are children, and which element is the root.
  • Distinguish descriptive metadata from primary data when deciding between attributes and nested elements.