Concepts / Introduction to D3.js and Data Visualization

Introduction to D3.js and Data Visualization

Visualizing a network graph requires running spjson.py to convert your PageRank database into JSON data, then opening force.html in a browser to view the D3.js visualization.

  • Programming

From Crawl to Graph

A network visualization is the final stage of a data workflow. The process begins with PageRank information stored in spider.json. You run spjson.py to extract information about the most highly linked pages and write it as JSON. You then open force.html in a browser, where D3.js reads that JSON and displays the pages and their connections as an interactive network graph.

readwriteloadrenderspider.jsonPageRank databasespjson.pyextracts highly linkedpagesJSON datanodes and linksforce.htmlbrowser pageD3.js graphinteractive network
What happens next as PageRank database data is converted into JSON and displayed as an interactive network graph?

What the JSON Represents

The generated JSON has two main arrays: nodes and links. A node represents a web page and includes properties such as its URL and PageRank score. A link represents a connection between two pages, typically using the indices or IDs of the source and target nodes. D3.js reads these arrays to position the nodes and draw lines between connected pages.

sourcetargetPage AURL and PageRank scoreLinksource to targetPage BURL and PageRank score
How do PageRanked URLs become nodes, and how do hyperlinks become connections between those nodes?

The graph is not a picture created independently of the data. Its visible objects come from the JSON structure: nodes supply the pages, and links supply the relationships between pages.

Generating the Visualization Data

A Complete Conversion Run

You have PageRank data in spider.json and want to view the most highly linked pages as a network graph.

Read the database: The workflow starts with spider.json, which stores the PageRank database produced from the web crawl.

Execute the converter: Run spjson.py from a command line or terminal. The script extracts information about the most highly linked pages.

Write JSON: spjson.py writes the extracted information in JSON format. The output contains nodes for web pages and links for connections between pages.

Open the viewer: Open force.html in a browser. D3.js reads the generated JSON and renders the network graph.

Explore the result: Use the graph's interactions to reposition nodes and reveal the URLs represented by nodes.

The PageRank database has been transformed into a D3.js network visualization.

readcreatecreateloadloadPageRank recordsfrom spider.jsonExtract pagesmost highly linkednodespagesforce.htmlD3.js readerlinksconnections
How does running spjson.py transform database records into the nodes-and-links JSON structure required by force.html?

Reading the Force Layout

When force.html opens, the graph is live and interactive. Its automatic arrangement uses a force-directed layout. Forces push nodes apart to avoid overlap, while links pull connected nodes together. The result is a natural arrangement intended to make the network easier to read.

What do you think happens?

What is the purpose of the forces in the graph layout?

  • To push every node to the same position
  • To separate nodes while keeping connected nodes together
  • To remove links from the JSON data
Reveal answer

Answer: To separate nodes while keeping connected nodes together

The force-directed layout uses forces that push nodes apart to avoid overlap and links that pull connected nodes together.

Exploring Individual Nodes

The layout gives you a broad view of the network, but a node does not automatically tell you which real web page it represents. Double-click a node to reveal its URL. The URL is essential for identifying the actual page behind that node.

You can also click and drag a node to reposition it. Dragging is useful when nodes overlap, when you want to untangle part of the graph, or when you want to focus on a particular region. The graph's layout can therefore be explored both automatically and through direct manipulation.

double-clickclick and dragNetwork nodeURL not revealedNode URLdouble-click resultRepositioned nodeclick and drag result
What changes in the graph when a learner identifies or repositions a node?

When a node matters to your investigation, double-click it to reveal its URL rather than relying only on its position or appearance in the graph. Use dragging to separate it from nearby nodes and make its connections easier to inspect.

Refreshing After a New Crawl

The web crawl can change over time. Running spider.py again may discover new pages or update link information, causing spider.json to grow or change. The existing visualization does not automatically represent those updates. To reflect the new crawl, run spjson.py again and then press refresh in the browser.

updatesreadnew JSONrefreshNew crawlspider.pyspider.jsonupdated databasespjson.pyregenerate JSONBrowser refreshreload visualizationUpdated graphnew network data
How does newly crawled data move through the conversion and loading steps to update the network graph?
  • Refreshing the browser without running spjson.py again

    The browser refresh can only load the data available to the visualization. It does not perform the conversion from the updated database.

    Fix: Run spjson.py again first, then press refresh in the browser.

  • Assuming a node's position identifies its web page

    The graph position is produced by the force-directed layout and does not itself reveal the real web page.

    Fix: Double-click the node to reveal its URL.

  • Treating the graph as a static image

    The D3.js visualization is live and interactive.

    Fix: Click and drag nodes to reposition them, and double-click nodes to identify their URLs.

Apply the Workflow

EASY

Imagine that spider.py has discovered additional pages and updated link information. Describe the exact sequence you would use to make the new information appear in the network graph. Then explain which interaction you would use to identify a particular page and which interaction you would use to separate overlapping nodes.

Hints
  • Start with the updated spider.json database.
  • Name the script that converts the database into JSON.
  • Remember that the browser must be refreshed after the data is regenerated.
  • Use double-clicking to reveal a URL and click-and-dragging to reposition a node.
  1. The workflow is: update the crawl, regenerate the visualization data, and refresh the browser. spider.json stores the PageRank database, spjson.py extracts the most highly linked pages and writes JSON, and force.html displays that JSON through D3.js. The JSON contains nodes for pages and links for connections. The force-directed layout arranges the graph, while clicking and dragging changes a node's position and double-clicking reveals its URL.

Key Takeaways

  • spjson.py converts PageRank information in spider.json into JSON for visualization.
  • The generated JSON contains nodes for web pages and links for connections between pages.
  • force.html uses D3.js to display the data as an interactive force-directed network graph.
  • Double-click a node to reveal its URL and click and drag it to reposition it.
  • After a new crawl, run spjson.py again and refresh the browser to update the graph.