Concepts / Introduction to JSON

Introduction to JSON

XML (eXtensible Markup Language) is a text-based format for encoding structured, hierarchical data using custom tags that describe the meaning of the data.

  • Programming

From Web Pages to Program Documents

The web was initially centered on displaying documents in browsers. HTML helped describe how those documents should be displayed. As it became easier for programs to retrieve and parse documents over HTTP, developers began creating documents intended for programs rather than directly for human readers. These program-consumed documents gave systems a way to exchange structured information across the internet.

XML, or eXtensible Markup Language, is a text-based format for encoding structured, hierarchical data using custom tags that describe the meaning of the data.

XML is not mainly about making information look attractive in a browser. Its purpose is to describe and organize information so that programs can read, interpret, and exchange it.

Reading XML as a Tree

An XML document is organized as a tree of nested elements. Each element has a start tag, content, and an end tag. An element may contain other elements, so the document can express which pieces of information belong inside other pieces of information. This hierarchy makes XML suitable for complex, document-like data.

containscontainscontainscontainscontainslibraryroot elementbookone booktitlebook titlebookanother bookauthorbook authoryearpublication year
What contains what in a hierarchical XML document?

A Library Catalog

Represent a library catalog whose books have titles, authors, publication years, and ISBNs.

Choose the root: Use library as the top-level element because the document describes a library catalog.

Add repeated items: Place book elements inside library. Multiple book elements represent multiple books.

Add book details: Place title, author, year, and ISBN elements inside each book element.

Add identifying information: A book can also have an id attribute, which supplies identifying information on the book element.

The resulting XML structure has one library root, multiple book elements, and descriptive child elements inside each book. Its nesting communicates both the data and the relationships among the data.

Parts of an XML Document

The main structural parts of the library example are the root element, nested elements, attributes, and text values. The root element is library, at the top level of the document. The book elements are nested inside it. Each book can contain child elements such as title, author, year, and ISBN. An id attribute can identify a particular book. The text inside an element supplies the value being described.

identifiesdescribesis enclosed by<book>start tagidattributebook detailscontent and child elements</book>end tag
What roles do tags, attributes, and text values play in an XML document?

Moving Data Between Systems

When two systems exchange XML, the first system creates an XML document containing the data. It sends that document over HTTP to the second system. The receiving system parses the XML, extracts the information it needs, and uses that information. For example, a library database can send book information to a web application that displays the books, stores them in a database, or sends them to another system.

createssent throughdelivers toextracts and usesLibrary databasecreates XMLXML documentbook informationHTTPcarries the documentWeb applicationparses XMLBook informationdisplayed or stored
How does structured data move from one program to another through an XML document?

The receiving system does not need to know the sending system's internal structure. Because XML is self-describing, its tags explain what each piece of data represents.

XML and HTML

FeatureXMLHTML
Primary purposeDescribe and exchange structured dataDescribe how content should be displayed
TagsCustom tags defined for the dataA predefined set of tags
Typical readerPrograms that parse and use the dataBrowsers that display documents
StrengthComplex, document-like data structuresDocuments intended for browser display

The presence of angle-bracket tags does not make XML and HTML the same. HTML uses a predefined set of tags to describe how content should be displayed. XML allows developers to create tags that describe the meaning of their data. A tag such as book communicates a data concept, while the purpose of an HTML tag is connected to presentation in a browser.

XML and JSON Choices

FormatBest suited forReason to choose it
XMLDocument-style data with complex hierarchiesCustom, self-describing tags represent complex relationships and document structures
JSONSimple dictionaries, lists, and lightweight program-to-program communicationIts use is associated with simpler, lightweight data exchange

XML and JSON are not interchangeable choices in every situation. A book catalog with a natural document hierarchy is an example of data for which XML is well suited. JSON is described as better for simple dictionaries, lists, and lightweight communication between programs. The choice depends on the shape and purpose of the data rather than on the fact that both formats can travel across the web.

Mistakes with Program Documents

  • Treating XML as another display language like HTML

    XML tags describe the meaning and organization of data, while HTML tags describe how content should be displayed.

    Fix: Ask whether the document is intended primarily for browser presentation or for program interpretation and data exchange.

  • Choosing XML or JSON without considering the data structure

    The source distinguishes XML as well suited to complex, document-style hierarchies and JSON as better suited to simple dictionaries, lists, and lightweight communication.

    Fix: Choose XML for complex document-like structures and consider JSON for simpler, lightweight exchanges.

  • Thinking the receiving program must know the sender's internal design

    XML is self-describing: its custom tags explain what the data represents.

    Fix: Focus on the structure and meaning expressed in the XML document that is sent.

  • Ignoring hierarchy

    Nested XML elements express relationships, such as books belonging to a library and titles belonging to books.

    Fix: Start with the root element, then follow its nested elements to understand what contains what.

Check Your Understanding

EASY

A library wants to send a catalog to another system. The catalog contains multiple books, and each book has a title, author, publication year, and ISBN. Decide whether XML or JSON is the better fit according to the distinctions in this article. Then explain what the root element and nested elements would represent.

Hints
  • First decide whether the catalog is simple lightweight data or document-style data with a natural hierarchy.
  • The source's library example uses library as the root and book as a nested element.
  • Each book can contain title, author, year, and ISBN child elements.

Practice Review

Explain how the library catalog travels from one system to another.

Identify the sender: The library database creates a document containing the book information.

Identify the transport: The XML document is sent over HTTP.

Identify the receiver: The other system receives and parses the XML document.

Identify the result: The receiving system extracts the book information and can display it, store it, or send it onward.

XML acts as a self-describing vehicle for structured data moving between systems.

Key Takeaways

  1. XML is a text-based format for encoding structured, hierarchical data with custom tags.
  2. XML elements form a tree: a root element contains nested elements that express relationships among data.
  3. XML describes what data means, while HTML primarily describes how content should be displayed.
  4. Programs can create XML, send it over HTTP, parse it, and use the extracted information.
  5. XML is suited to complex document-style data, while JSON is better suited to simple dictionaries, lists, and lightweight program-to-program communication.

Key Takeaways

  • XML uses custom semantic tags to encode structured, hierarchical data.
  • The tree structure of XML makes relationships between document parts explicit.
  • XML differs from HTML because XML describes data meaning while HTML describes browser presentation.
  • Programs exchange XML by creating, sending, parsing, and using XML documents.
  • XML and JSON serve different purposes: XML supports complex document-style hierarchies, while JSON supports simpler and more lightweight exchanges.