Web Crawling with spider.py
Visualizing a network graph requires running spjson.py to convert your PageRank database into JSON data, then opening force.html in a browser to view the D3.js visualization.
From Crawl to Network
A web crawl produces data, but the data becomes easier to explore when it is turned into a visual network. In this workflow, spider.py discovers pages or updates link information in the PageRank database stored in spider.json. Next, spjson.py extracts information about the most highly linked pages and writes JSON data. Finally, force.html loads that JSON so D3.js can display an interactive graph.
The Conversion Step
spjson.py is the bridge between stored crawl results and the visualization. When you execute it from a command line or terminal, it reads the PageRank database in spider.json. It extracts information about the most highly linked pages and writes that information in JSON format. D3.js can then read this JSON and render the network graph.
Treat spjson.py as a regeneration step rather than as a manual editing task. The source material says that you do not need to edit the JSON file manually. Instead, let spjson.py create the data that D3.js consumes.
Nodes and Links
The generated JSON has two main arrays: nodes and links. A node represents a web page and includes properties such as its URL and PageRank score. A link represents a connection between two pages and is typically specified using the indices or IDs of the source and target nodes. D3.js uses these two parts together: it positions the nodes and draws lines for the links.
Following One Connection
Suppose the generated JSON contains a node for Page A, a node for Page B, and a link whose source is Page A and whose target is Page B. What does that part of the graph represent?
Read the nodes: The two nodes represent two web pages. Each node can include the page URL and its PageRank score.
Read the link: The link identifies Page A as the source and Page B as the target, using the nodes' indices or IDs.
Read the rendered graph: D3.js uses the node records and the link record to position the pages and draw a connection between them.
The visual connection represents the relationship encoded by the link between the two page nodes.
The Force-Directed View
After force.html opens in a browser, the D3.js visualization is live and interactive. The graph uses a force-directed layout. Forces push nodes apart to help avoid overlap, while links pull connected nodes together. The result is an automatic arrangement intended to make the network readable rather than a fixed list of pages.
Moving a node changes where it appears in the view, but the reason for moving it is interpretive: you can untangle overlapping nodes or focus on a particular region of the graph. The links continue to show the encoded connections between pages.
Reading a Node
A visible circle or node is not enough to tell you which real web page it represents. Double-click a node to reveal its URL. That interaction connects the abstract graph element to the actual page identity. After identifying it, click and drag the node if you need to separate it from nearby nodes or examine its connected region.
What do you think happens?
You see a node in force.html but do not know which page it represents. Which interaction identifies the page?
Reveal answer
Answer: Double-click the node
The visualization reveals the URL represented by a node when you double-click it. Dragging changes the node's position instead.
Refreshing New Crawl Data
The visualization reflects the data that spjson.py has generated. If you run spider.py again, the spider.json database can grow or change as new pages are discovered or link information is updated. The browser will not reflect those changes merely because the database changed. Run spjson.py again to regenerate the JSON, then press refresh in the browser to reload the visualization.
Common Workflow Mistakes
Opening force.html before generating the JSON data
The visualization depends on the JSON output that spjson.py writes for D3.js to read.
Fix:
Execute spjson.py first, then open force.html in the browser.Assuming a node's position identifies its URL
The graph position does not by itself reveal the real page identity.
Fix:
Double-click the node to reveal its URL.Refreshing the browser without regenerating the JSON
The visualization still has the previously generated JSON data.
Fix:
Run spjson.py again and then refresh the browser.Treating links as separate pages
Nodes represent web pages, while links represent connections between pages.
Fix:
Read nodes as pages and links as relationships between source and target nodes.
Practice the Pipeline
A new crawl has been completed with spider.py. Describe the exact sequence you would use to make the updated network visible in the browser. Then explain what you would do if you wanted to identify one node and separate it from overlapping nodes.
Hints
- Start with the PageRank database and the script that converts it to JSON.
- Remember that the browser needs a refresh after the JSON is regenerated.
- Use double-clicking for URL identification and click-and-drag for repositioning.
Checking an Updated Graph
You have run spider.py again and want force.html to show the updated crawl.
Update the stored data: The new crawl changes or grows the PageRank database stored in spider.json.
Regenerate the visualization data: Execute spjson.py so it reads the updated database and writes new JSON containing nodes and links for the highly linked pages.
Reload the graph: Press refresh in the browser while force.html is open so the visualization loads the regenerated data.
Inspect a page: Double-click a node to reveal its URL, or click and drag it to reposition the node for easier inspection.
The browser shows the regenerated network, and node interactions help you identify pages and examine their connections.
Key Takeaways
- spider.py discovers pages or updates link information in the PageRank database stored in spider.json.
- spjson.py reads that database and writes JSON containing nodes for web pages and links for connections between them.
- force.html loads the JSON so D3.js can display a force-directed network graph.
- Double-click a node to reveal its URL, and click and drag a node to reposition it.
- After a new crawl, run spjson.py again and refresh the browser to display the updated network.
Key Takeaways
- The workflow connects spider.py, spider.json, spjson.py, and force.html in that order.
- spjson.py converts PageRank database information about highly linked pages into JSON for D3.js.
- The JSON separates page nodes from the links that connect them.
- The force-directed graph can be explored by dragging nodes and double-clicking them to reveal URLs.
- Updated crawl results require both regenerating the JSON and refreshing the browser.