Concepts / How to Parse XML

How to Parse XML

XML (eXtensible Markup Language) is a text-based format for encoding structured, hierarchical data using custom tags that describe the meaning of the data.

  • Programming

From Web Pages to Program Documents

The web was first associated mainly with displaying documents in browsers. As it became easier for programs to retrieve and parse documents over HTTP, developers began creating documents intended not for human eyes in a browser, but for other programs to read and process. XML emerged as a text-based format for encoding structured, hierarchical data with custom tags that explain what the data means.

XML, or eXtensible Markup Language, is a text-based format for encoding structured, hierarchical data using custom tags that describe the meaning of the data.

To parse XML means to process an XML document so a program can extract and use the information represented by its elements, attributes, and text.

Following the XML Tree

An XML document is organized as a tree of nested elements. Each element has a start tag, content, and an end tag. An element may contain other elements, so the document can represent relationships between pieces of data. The outermost element is the root. Inside it are child elements, which may themselves contain more child elements.

containscontainscontainscontainscontainslibraryroot elementbookfirst childtitlebook informationbooksecond childauthorbook informationyearbook information
What contains what in an XML document, and how do nested tags represent parent-child relationships?

In the library example, library is the root element. It contains multiple book elements. Each book contains child elements such as title, author, and year. A program that parses the document can follow this hierarchy to locate the book information it needs.

What do you think happens?

If a program needs the author of a book, where should it look in the library document?

  • Among the children of the library element
  • Among the child elements of the relevant book element
  • Only in the root element's name
Reveal answer

Answer: Among the child elements of the relevant book element

The source structure places title, author, and year inside each book element. The nesting tells the program which author belongs to which book.

Reading Element Anatomy

Every XML element is built from parts that work together to define a piece of data. The start tag identifies where the element begins, the content holds text or nested elements, and the end tag identifies where the element ends. XML can also use attributes to provide additional information about an element. In the library example, each book has an id attribute, while title, author, year, and ISBN are represented by child elements.

may includedescribesends at<book>start tagidattributebook datatext or nested elements</book>end tag
How do the opening tag, closing tag, attribute, and text fit together to form a program-consumed document?
PartPurposeLibrary example
Root elementTop-level element containing the document's hierarchylibrary
Start tagMarks the beginning of an element<book>
ContentText or nested elements inside an elementtitle, author, year, and ISBN
End tagMarks the end of an element</book>
AttributeProvides additional information about an elementThe id attribute on book

The parts used to describe structured data in the source's library document

These parts give XML its self-describing quality. The tags explain what each piece of data represents, so a receiving program can understand the document's organization without needing to know the sending system's internal structure.

Tracing an XML Exchange

createssent throughdeliversparsesSystem Acreates XMLXML documentstructured dataHTTPcarries documentSystem Breceives documentBook informationextracted and used
How does XML data move from one program or system to another, and how is it interpreted by the receiving program?
  1. One system creates an XML document containing structured data.
  2. The system sends the XML document to another system over HTTP.
  3. The receiving system parses the XML document.
  4. The receiving system extracts and uses the information, such as displaying it, storing it in a database, or sending it to another system.

The exchange works because XML is standardized: a program that understands XML can read a well-formed XML document even when the document uses custom tags created by another system. The tags carry meaning, while the nested structure shows how the information is organized.

Parsing a library catalog

A library database sends a catalog containing books, titles, authors, publication years, and ISBNs to a web application.

Create: The library database creates an XML document with library as the root element, book elements inside it, and title, author, year, and ISBN elements inside each book.

Send: The library database sends the XML document to the web application over HTTP.

Parse: The web application parses the document and follows its hierarchy to extract the book information.

Use: The application can use the extracted information to display the books, store the information in a database, or send it to another system.

The XML document acts as the vehicle carrying understandable, structured book information between the two systems.

Choosing XML, HTML, or JSON

guides presentationdescribes dataHTMLdescribes displayXMLdescribes meaningBrowserdisplays documentsProgramreads and processes data
What is the difference between HTML tags that describe presentation and XML tags that describe data meaning?
often suited tooften suited toXMLcomplex documenthierarchiesJSONsimple data exchangeEnterprise systemsweb services andconfigurationLightweightcommunicationdictionaries and lists
How do XML and JSON differ in structure and in the situations where programs typically use each format?
FormatPrimary purposeGood fit described by the source
HTMLDescribe how information should be displayedDocuments displayed in browsers
XMLDescribe what data means and how it is organizedComplex, document-style data with hierarchies
JSONExchange simpler structured dataDictionaries, lists, and lightweight program-to-program communication

XML and JSON are both used to exchange data across the web, but they are not interchangeable in every situation. XML is best suited for document-style data with complex hierarchies. JSON is better suited for simple dictionaries, lists, and lightweight program-to-program communication. HTML has a different emphasis: it describes how data should be displayed rather than what the data means.

Mistakes Beginners Make

  • Treating XML as if it were mainly a presentation language

    HTML describes how data should be displayed, while XML describes what the data means and how it is organized.

    Fix: Read XML tags as labels for data and relationships, not primarily as instructions for browser presentation.

  • Ignoring nesting

    The XML tree uses nesting to show which pieces of information belong together.

    Fix: Start at the root and follow parent-child relationships until you reach the data the program needs.

  • Choosing XML for every data exchange

    The source identifies JSON as a better fit for simple dictionaries, lists, and lightweight program-to-program communication.

    Fix: Choose XML when the data is document-style and has a complex hierarchy; consider JSON for simpler exchanges.

  • Assuming the receiving program must know the sender's internal database structure

    XML is self-describing: its tags explain what each piece of data represents.

    Fix: Focus on the meaning expressed by the XML tags and the hierarchy they form.

When learning to parse XML, first identify the root element, then map the elements nested inside it. After that, identify any attributes and the text or child elements that provide the actual data. This order keeps the document's structure visible instead of reducing it to unrelated strings.

Practice the Parsing Mindset

MEDIUM

A document contains a root element named library. Inside it are several book elements. Each book has an id attribute and child elements named title, author, year, and ISBN. Explain how a program would use the document's structure to obtain the title and author for one particular book, and state whether this is a better fit for XML or for JSON according to the source's use-case distinction.

Hints
  • Begin with the root element.
  • Locate the particular book element and its id attribute.
  • Look among that book's child elements for title and author.
  • Classify the data as document-style data with a natural hierarchy.

Key Takeaways

  1. XML is a text-based format that uses custom tags to describe structured, hierarchical data.
  2. XML documents form trees: a root element contains nested child elements, and the nesting represents relationships.
  3. Parsing XML means processing the document so a program can extract and use its information.
  4. Programs can create XML, send it over HTTP, and have another program parse it.
  5. HTML focuses on display, XML focuses on data meaning, and JSON is generally better for simpler dictionaries, lists, and lightweight communication.

Key Takeaways

  • XML represents structured data with custom, meaningful tags.
  • Its nested elements form a tree that programs can follow when extracting information.
  • XML commonly moves between systems over HTTP and is parsed by the receiving program.
  • XML is suited to complex document-style hierarchies, while JSON is suited to simpler data exchange.
  • HTML and XML both use tags, but HTML describes presentation whereas XML describes data meaning.