Visualizing Network Data with D3
spider.py systematically crawls web pages and stores them in a local database, recording both page content and the links between pages.
From Pages to a Network
A web page is more than a document to a crawler. It is also a location in a network. Its outgoing links point toward other pages, and those pages may point onward. spider.py collects this structure by visiting pages, recording their URLs and outgoing-link counts, and storing relationships between source pages and target pages. D3 can then turn the stored graph into an interactive visual network.
The Crawl State
The crawl begins with a starting URL and a requested number of pages. spider.py fetches a page, extracts its links, stores the page and its outgoing links, and then moves toward another unvisited link. Each page receives a unique ID and an outgoing-link count. Each relationship records a source page ID and a target page ID. Together, these records form a graph inside the database.
A Three-Page Crawl
Trace a crawl that starts with http://example.com/ and requests three pages from an empty database.
Start: spider.py receives the starting URL and the requested crawl size.
Store the first page: The starting page is fetched, recorded with a unique ID and outgoing-link count, and its links are added as candidates for later visits.
Continue to an unvisited link: The crawler selects another unvisited link, fetches that page, and records its page and outgoing relationships.
Reach the requested count: The process continues until three pages have been crawled, leaving other discovered links available for possible later work.
The database contains page records and link relationships for the collected portion of the network, while unvisited discovered links can support continued crawling.
Incremental Crawling
The crawler is additive. After a run, the database remains on disk in spider.sqlite. When spider.py is run again, it remembers pages already stored and skips them rather than fetching them again. A later run therefore adds new pages to the existing collection. On each restart, the crawler selects a random unvisited page from its queue, so successive runs can explore different branches of the web.
Independent Webs
Different crawl sessions can begin from different starting URLs and still use the same spider.sqlite database. For example, one session can begin at http://www.dr-chuck.com/ and another at http://www.wikipedia.org/. These starting points are treated as separate webs, although their pages and links are stored in the same overall database. The webs may not be directly connected, but the crawler treats all unvisited links as one unified queue.
Separate starting points do not create separate database files by themselves. Their records coexist in the same database, and the crawler can interleave visits by choosing from all unvisited links across the webs.
From PageRank to D3
Crawling creates the raw graph, but raw database records are not yet an easy visual explanation. After pages have been crawled, PageRank data identifies the most highly linked pages. Running spjson.py reads the PageRank database stored in spider.json and writes JSON containing nodes and links. Nodes represent pages and can include a URL and PageRank score; links represent connections between source and target nodes. D3.js reads this JSON structure and renders the network.
- Crawl pages with spider.py so page and link records exist.
- Use the collected data for PageRank analysis.
- Execute spjson.py from the command line or terminal.
- Let spjson.py extract information about the most highly linked pages and write the JSON data.
- Open force.html in a browser.
- Inspect the D3 network, which uses nodes for pages and links for connections.
Reading the Force Layout
When force.html opens, D3 displays pages as nodes and connections as links. A force-directed layout uses forces that push nodes apart to reduce overlap while links pull connected nodes together. The resulting arrangement is automatic, so the position of a node is a visual layout result rather than the page's URL or database ID.
Use dragging to separate overlapping nodes or focus attention on a region. Use double-clicking when you need to identify the real page represented by a node, because the URL is what tells you which web page the node corresponds to.
Before opening the visualization, explain what information you expect to find in the JSON arrays named nodes and links.
Hints
- A node represents a web page.
- A link represents a connection between two pages.
- A node can include a URL and PageRank score.
Refreshing New Crawl Data
The network is a snapshot of the data that was converted into JSON. If spider.py discovers new pages or updates link information, the existing visualization does not automatically contain those changes. Run spjson.py again to regenerate the JSON from the updated PageRank data, then press refresh in the browser. This reloads the visualization with the newer network data.
Opening force.html before generating the JSON data.
The D3 visualization consumes the JSON representation rather than the raw database records.
Fix:
Execute spjson.py before opening or refreshing force.html.Expecting a later spider.py run to re-fetch every stored page.
The crawler is designed to skip pages already in the database.
Fix:
Understand later runs as incremental collection of unvisited pages.Treating a node's screen position as its identity.
The layout positions nodes automatically according to forces and connections.
Fix:
Double-click the node to reveal its URL.Expecting updated crawl data to appear without regeneration and refresh.
The visualization is based on previously generated JSON.
Fix:
Run spjson.py again and press refresh in the browser.
A second crawl adds pages to spider.sqlite. What two actions are needed before the new pages can appear in force.html?
Hints
- The visualization reads generated JSON.
- The browser must load the updated data.
Network Visualization Workflow
- Begin with collection: spider.py visits a starting URL, extracts links, and stores pages and relationships in spider.sqlite. Continue incrementally: later runs skip pages already stored and select from unvisited links. Combine sources when needed: multiple starting URLs can contribute separate webs to the same database. Analyze the graph: PageRank identifies highly linked pages. Convert the results: spjson.py writes nodes and links as JSON. Explore the result: force.html uses D3 to display the force-directed graph, while dragging changes node positions and double-clicking reveals URLs. When the crawl changes, regenerate the JSON and refresh the browser.
Key Takeaways
- spider.py converts discovered web pages and their outgoing links into a persistent graph stored in spider.sqlite.
- Repeated crawl sessions are additive because stored pages are skipped and unvisited links remain available.
- Different starting URLs can coexist as separate webs within one database and unified crawl queue.
- spjson.py converts PageRank data into JSON nodes and links for D3.
- force.html provides a force-directed, interactive view that can be refreshed after new data is generated.