Understanding PageRank and Web Graph Structure
Visualizing a network graph requires running spjson.py to convert your PageRank database into JSON data, then opening force.html in a browser to view the D3.js visualization.
From Crawl Data to a Graph
A PageRank database contains information gathered from a web crawl, but the database itself is not yet a visual network. The visualization workflow has two major stages: spjson.py reads the PageRank database stored in spider.json and writes JSON data, then force.html opens that data in a browser so D3.js can render an interactive graph.
The important boundary is between stored crawl data and display data. spjson.py performs the conversion; force.html and D3.js perform the visual rendering.
Nodes, Links, and Force Layouts
The generated JSON contains two main arrays. Nodes represent web pages. Each node includes properties such as the page URL and its PageRank score. Links represent connections between pages, usually by referring to the source and target nodes through their indices or IDs. D3.js reads these arrays, draws the nodes and links, and positions them as a network graph.
The visualization focuses on the most highly linked pages and their connections. Its force-directed layout is automatic: forces push nodes apart to reduce overlap, while links pull connected nodes together. The resulting arrangement is intended to make the network easier to read; it is a layout of the relationships, not a fixed map of geographic positions.
Reading a Small Network
Suppose the generated JSON describes four web-page nodes and links connecting Page A to Page B and Page C, with both Page B and Page C connected to Page D.
Identify the nodes: Treat Page A, Page B, Page C, and Page D as web pages represented by nodes.
Identify the links: Treat each connection as a link between a source page and a target page.
Interpret the layout: Expect connected pages to be pulled toward one another while the force-directed layout pushes nodes apart enough to reduce overlap.
Find the real pages: The labels in this illustrative example are not the actual identities of pages. In force.html, double-clicking a node reveals the URL that identifies its represented page.
The graph should be read as a relationship structure: nodes stand for pages, links stand for connections, and the force-directed arrangement helps expose that structure.
Running the Visualization Workflow
- Open a command line or terminal in the environment containing the PageRank files.
- Execute spjson.py.
- Allow the script to read spider.json and extract information about the most highly linked pages.
- Let spjson.py write the extracted information in JSON format.
- Open force.html in a browser.
- Inspect the live D3.js network visualization.
Exploring Nodes in the Browser
Once force.html opens, the graph is live and interactive. Click and drag a node to reposition it. This can help untangle overlapping nodes or bring a particular region of the graph into focus. The force-directed layout responds as the node is moved, so interaction becomes a way to investigate the structure rather than merely view a static picture.
Dragging answers a layout question: where should this node sit so the graph is easier to inspect? Double-clicking answers an identity question: which real web page does this node represent? The URL is the way to identify the page represented by a node.
In force.html, choose a node that overlaps another node. Drag it to a clearer position, then double-click it. Record what the URL tells you about the page represented by that node.
Hints
- Use dragging to separate or focus on a region of the network.
- Use double-clicking when you need the node's page identity.
Refreshing After a New Crawl
The crawl data can change when spider.py discovers new pages or updates link information. Those changes affect spider.json, so the existing visualization does not automatically represent the newest crawl data. To update the graph, run spjson.py again and then press refresh in the browser.
- Run spider.py again when you want to discover new pages or update link information.
- Run spjson.py again so the changed spider.json data is converted into new JSON data.
- Return to the browser displaying force.html.
- Press refresh to load the updated visualization data.
Common Workflow Mistakes
Opening force.html without first generating the JSON data.
spjson.py is the step that reads spider.json and writes the JSON structure that D3.js can render.
Fix:
Run spjson.py before opening or refreshing force.html.Treating a node's visual position as its identity.
The force-directed layout positions nodes to make relationships readable; the URL identifies the represented page.
Fix:
Double-click the node to reveal its URL.Dragging a node but expecting the underlying web data to change.
Dragging repositions the node in the interactive graph for inspection.
Fix:
Use dragging to clarify the layout, and use the generated data and crawl workflow for information about pages and links.Running a new crawl but forgetting to regenerate and reload the visualization.
The new database contents must pass through spjson.py, and the browser must be refreshed before the graph displays the updated data.
Fix:
Run spjson.py again and press refresh in the browser.
Workflow Check
Diagnosing an Outdated Graph
A new crawl has been performed, but the browser visualization still shows the earlier network. What should you check?
Check the database stage: Confirm that the new crawl has changed or updated the PageRank database stored in spider.json.
Repeat the conversion stage: Run spjson.py again so it reads the current spider.json data and writes updated JSON containing nodes and links.
Repeat the display stage: Press refresh in the browser displaying force.html so the visualization loads the updated data.
Inspect the result: Use the refreshed graph's layout and node interactions to explore the current network, including double-clicking nodes to reveal their URLs.
The complete refresh sequence is updated crawl data, a new spjson.py conversion, and a browser refresh.
Explain the role of each item in this sequence: spider.json, spjson.py, JSON nodes and links, force.html, and the browser refresh. Then describe which interaction identifies a node's URL and which interaction changes its position.
Hints
- Separate database storage, conversion, graph data, display, and reload.
- The two main node interactions are dragging and double-clicking.
Key Takeaways
- spjson.py reads the PageRank database in spider.json and writes JSON data for the visualization.
- The JSON contains nodes for web pages and links for connections between those pages.
- force.html uses D3.js to display the data as an interactive force-directed network graph.
- Dragging repositions nodes, while double-clicking reveals the URL represented by a node.
- After a new crawl, run spjson.py again and refresh the browser to display updated network data.