How to Parse XML
XML (eXtensible Markup Language) is a text-based format for encoding structured, hierarchical data using custom tags that describe the meaning of the data.
From Web Pages to Program Documents
The web was first associated mainly with displaying documents in browsers. As it became easier for programs to retrieve and parse documents over HTTP, developers began creating documents intended not for human eyes in a browser, but for other programs to read and process. XML emerged as a text-based format for encoding structured, hierarchical data with custom tags that explain what the data means.
XML, or eXtensible Markup Language, is a text-based format for encoding structured, hierarchical data using custom tags that describe the meaning of the data.
To parse XML means to process an XML document so a program can extract and use the information represented by its elements, attributes, and text.
Following the XML Tree
An XML document is organized as a tree of nested elements. Each element has a start tag, content, and an end tag. An element may contain other elements, so the document can represent relationships between pieces of data. The outermost element is the root. Inside it are child elements, which may themselves contain more child elements.
In the library example, library is the root element. It contains multiple book elements. Each book contains child elements such as title, author, and year. A program that parses the document can follow this hierarchy to locate the book information it needs.
What do you think happens?
If a program needs the author of a book, where should it look in the library document?
Reveal answer
Answer: Among the child elements of the relevant book element
The source structure places title, author, and year inside each book element. The nesting tells the program which author belongs to which book.
Reading Element Anatomy
Every XML element is built from parts that work together to define a piece of data. The start tag identifies where the element begins, the content holds text or nested elements, and the end tag identifies where the element ends. XML can also use attributes to provide additional information about an element. In the library example, each book has an id attribute, while title, author, year, and ISBN are represented by child elements.
| Part | Purpose | Library example |
|---|---|---|
| Root element | Top-level element containing the document's hierarchy | library |
| Start tag | Marks the beginning of an element | <book> |
| Content | Text or nested elements inside an element | title, author, year, and ISBN |
| End tag | Marks the end of an element | </book> |
| Attribute | Provides additional information about an element | The id attribute on book |
The parts used to describe structured data in the source's library document
These parts give XML its self-describing quality. The tags explain what each piece of data represents, so a receiving program can understand the document's organization without needing to know the sending system's internal structure.
Tracing an XML Exchange
- One system creates an XML document containing structured data.
- The system sends the XML document to another system over HTTP.
- The receiving system parses the XML document.
- The receiving system extracts and uses the information, such as displaying it, storing it in a database, or sending it to another system.
The exchange works because XML is standardized: a program that understands XML can read a well-formed XML document even when the document uses custom tags created by another system. The tags carry meaning, while the nested structure shows how the information is organized.
Parsing a library catalog
A library database sends a catalog containing books, titles, authors, publication years, and ISBNs to a web application.
Create: The library database creates an XML document with library as the root element, book elements inside it, and title, author, year, and ISBN elements inside each book.
Send: The library database sends the XML document to the web application over HTTP.
Parse: The web application parses the document and follows its hierarchy to extract the book information.
Use: The application can use the extracted information to display the books, store the information in a database, or send it to another system.
The XML document acts as the vehicle carrying understandable, structured book information between the two systems.
Choosing XML, HTML, or JSON
| Format | Primary purpose | Good fit described by the source |
|---|---|---|
| HTML | Describe how information should be displayed | Documents displayed in browsers |
| XML | Describe what data means and how it is organized | Complex, document-style data with hierarchies |
| JSON | Exchange simpler structured data | Dictionaries, lists, and lightweight program-to-program communication |
XML and JSON are both used to exchange data across the web, but they are not interchangeable in every situation. XML is best suited for document-style data with complex hierarchies. JSON is better suited for simple dictionaries, lists, and lightweight program-to-program communication. HTML has a different emphasis: it describes how data should be displayed rather than what the data means.
Mistakes Beginners Make
Treating XML as if it were mainly a presentation language
HTML describes how data should be displayed, while XML describes what the data means and how it is organized.
Fix:
Read XML tags as labels for data and relationships, not primarily as instructions for browser presentation.Ignoring nesting
The XML tree uses nesting to show which pieces of information belong together.
Fix:
Start at the root and follow parent-child relationships until you reach the data the program needs.Choosing XML for every data exchange
The source identifies JSON as a better fit for simple dictionaries, lists, and lightweight program-to-program communication.
Fix:
Choose XML when the data is document-style and has a complex hierarchy; consider JSON for simpler exchanges.Assuming the receiving program must know the sender's internal database structure
XML is self-describing: its tags explain what each piece of data represents.
Fix:
Focus on the meaning expressed by the XML tags and the hierarchy they form.
When learning to parse XML, first identify the root element, then map the elements nested inside it. After that, identify any attributes and the text or child elements that provide the actual data. This order keeps the document's structure visible instead of reducing it to unrelated strings.
Practice the Parsing Mindset
A document contains a root element named library. Inside it are several book elements. Each book has an id attribute and child elements named title, author, year, and ISBN. Explain how a program would use the document's structure to obtain the title and author for one particular book, and state whether this is a better fit for XML or for JSON according to the source's use-case distinction.
Hints
- Begin with the root element.
- Locate the particular book element and its id attribute.
- Look among that book's child elements for title and author.
- Classify the data as document-style data with a natural hierarchy.
Key Takeaways
- XML is a text-based format that uses custom tags to describe structured, hierarchical data.
- XML documents form trees: a root element contains nested child elements, and the nesting represents relationships.
- Parsing XML means processing the document so a program can extract and use its information.
- Programs can create XML, send it over HTTP, and have another program parse it.
- HTML focuses on display, XML focuses on data meaning, and JSON is generally better for simpler dictionaries, lists, and lightweight communication.
Key Takeaways
- XML represents structured data with custom, meaningful tags.
- Its nested elements form a tree that programs can follow when extracting information.
- XML commonly moves between systems over HTTP and is parsed by the receiving program.
- XML is suited to complex document-style hierarchies, while JSON is suited to simpler data exchange.
- HTML and XML both use tags, but HTML describes presentation whereas XML describes data meaning.