Introduction to D3.js and Data Visualization
Visualizing a network graph requires running spjson.py to convert your PageRank database into JSON data, then opening force.html in a browser to view the D3.js visualization.
From Crawl to Graph
A network visualization is the final stage of a data workflow. The process begins with PageRank information stored in spider.json. You run spjson.py to extract information about the most highly linked pages and write it as JSON. You then open force.html in a browser, where D3.js reads that JSON and displays the pages and their connections as an interactive network graph.
What the JSON Represents
The generated JSON has two main arrays: nodes and links. A node represents a web page and includes properties such as its URL and PageRank score. A link represents a connection between two pages, typically using the indices or IDs of the source and target nodes. D3.js reads these arrays to position the nodes and draw lines between connected pages.
The graph is not a picture created independently of the data. Its visible objects come from the JSON structure: nodes supply the pages, and links supply the relationships between pages.
Generating the Visualization Data
A Complete Conversion Run
You have PageRank data in spider.json and want to view the most highly linked pages as a network graph.
Read the database: The workflow starts with spider.json, which stores the PageRank database produced from the web crawl.
Execute the converter: Run spjson.py from a command line or terminal. The script extracts information about the most highly linked pages.
Write JSON: spjson.py writes the extracted information in JSON format. The output contains nodes for web pages and links for connections between pages.
Open the viewer: Open force.html in a browser. D3.js reads the generated JSON and renders the network graph.
Explore the result: Use the graph's interactions to reposition nodes and reveal the URLs represented by nodes.
The PageRank database has been transformed into a D3.js network visualization.
Reading the Force Layout
When force.html opens, the graph is live and interactive. Its automatic arrangement uses a force-directed layout. Forces push nodes apart to avoid overlap, while links pull connected nodes together. The result is a natural arrangement intended to make the network easier to read.
What do you think happens?
What is the purpose of the forces in the graph layout?
Reveal answer
Answer: To separate nodes while keeping connected nodes together
The force-directed layout uses forces that push nodes apart to avoid overlap and links that pull connected nodes together.
Exploring Individual Nodes
The layout gives you a broad view of the network, but a node does not automatically tell you which real web page it represents. Double-click a node to reveal its URL. The URL is essential for identifying the actual page behind that node.
You can also click and drag a node to reposition it. Dragging is useful when nodes overlap, when you want to untangle part of the graph, or when you want to focus on a particular region. The graph's layout can therefore be explored both automatically and through direct manipulation.
When a node matters to your investigation, double-click it to reveal its URL rather than relying only on its position or appearance in the graph. Use dragging to separate it from nearby nodes and make its connections easier to inspect.
Refreshing After a New Crawl
The web crawl can change over time. Running spider.py again may discover new pages or update link information, causing spider.json to grow or change. The existing visualization does not automatically represent those updates. To reflect the new crawl, run spjson.py again and then press refresh in the browser.
Refreshing the browser without running spjson.py again
The browser refresh can only load the data available to the visualization. It does not perform the conversion from the updated database.
Fix:
Run spjson.py again first, then press refresh in the browser.Assuming a node's position identifies its web page
The graph position is produced by the force-directed layout and does not itself reveal the real web page.
Fix:
Double-click the node to reveal its URL.Treating the graph as a static image
The D3.js visualization is live and interactive.
Fix:
Click and drag nodes to reposition them, and double-click nodes to identify their URLs.
Apply the Workflow
Imagine that spider.py has discovered additional pages and updated link information. Describe the exact sequence you would use to make the new information appear in the network graph. Then explain which interaction you would use to identify a particular page and which interaction you would use to separate overlapping nodes.
Hints
- Start with the updated spider.json database.
- Name the script that converts the database into JSON.
- Remember that the browser must be refreshed after the data is regenerated.
- Use double-clicking to reveal a URL and click-and-dragging to reposition a node.
- The workflow is: update the crawl, regenerate the visualization data, and refresh the browser. spider.json stores the PageRank database, spjson.py extracts the most highly linked pages and writes JSON, and force.html displays that JSON through D3.js. The JSON contains nodes for pages and links for connections. The force-directed layout arranges the graph, while clicking and dragging changes a node's position and double-clicking reveals its URL.
Key Takeaways
- spjson.py converts PageRank information in spider.json into JSON for visualization.
- The generated JSON contains nodes for web pages and links for connections between pages.
- force.html uses D3.js to display the data as an interactive force-directed network graph.
- Double-click a node to reveal its URL and click and drag it to reposition it.
- After a new crawl, run spjson.py again and refresh the browser to update the graph.