Concepts / Visualizing Network Data with D3

Visualizing Network Data with D3

spider.py systematically crawls web pages and stores them in a local database, recording both page content and the links between pages.

  • Programming

From Pages to a Network

A web page is more than a document to a crawler. It is also a location in a network. Its outgoing links point toward other pages, and those pages may point onward. spider.py collects this structure by visiting pages, recording their URLs and outgoing-link counts, and storing relationships between source pages and target pages. D3 can then turn the stored graph into an interactive visual network.

fetchoutgoing linkoutgoing linkstore page and linksstore when visitedstore when visitedStarting URLpage to fetchPage Astored pagePage Bdiscovered linkspider.sqlitepages and linksPage Cdiscovered link
How does spider.py move from a starting URL to discovered pages and record the resulting connections?

The Crawl State

The crawl begins with a starting URL and a requested number of pages. spider.py fetches a page, extracts its links, stores the page and its outgoing links, and then moves toward another unvisited link. Each page receives a unique ID and an outgoing-link count. Each relationship records a source page ID and a target page ID. Together, these records form a graph inside the database.

fetchidentify pagerecord connectionsstorestorePage requeststarting URLPage contentlinks extractedPage recordID, URL, link countspider.sqlitepersistent databaseLink recordsource ID, target ID
How do a page request and its extracted links become stored page and relationship records?

A Three-Page Crawl

Trace a crawl that starts with http://example.com/ and requests three pages from an empty database.

Start: spider.py receives the starting URL and the requested crawl size.

Store the first page: The starting page is fetched, recorded with a unique ID and outgoing-link count, and its links are added as candidates for later visits.

Continue to an unvisited link: The crawler selects another unvisited link, fetches that page, and records its page and outgoing relationships.

Reach the requested count: The process continues until three pages have been crawled, leaving other discovered links available for possible later work.

The database contains page records and link relationships for the collected portion of the network, while unvisited discovered links can support continued crawling.

Incremental Crawling

The crawler is additive. After a run, the database remains on disk in spider.sqlite. When spider.py is run again, it remembers pages already stored and skips them rather than fetching them again. A later run therefore adds new pages to the existing collection. On each restart, the crawler selects a random unvisited page from its queue, so successive runs can explore different branches of the web.

checkcheckfetch and storeExisting databasestored pagesStored pageskip re-fetchExpanded databasenew records addedUnvisited pagecandidate for crawl
What changes during a later crawl when a page is already stored, and what happens to an unvisited page?

Independent Webs

Different crawl sessions can begin from different starting URLs and still use the same spider.sqlite database. For example, one session can begin at http://www.dr-chuck.com/ and another at http://www.wikipedia.org/. These starting points are treated as separate webs, although their pages and links are stored in the same overall database. The webs may not be directly connected, but the crawler treats all unvisited links as one unified queue.

discoversdiscoversstored instored inStarting URL Afirst webWeb A pageslinked clusterspider.sqliteunified graphStarting URL Bsecond webWeb B pageslinked cluster
How can pages discovered from different starting URLs coexist while remaining separate clusters?

Separate starting points do not create separate database files by themselves. Their records coexist in the same database, and the crawler can interleave visits by choosing from all unvisited links across the webs.

From PageRank to D3

Crawling creates the raw graph, but raw database records are not yet an easy visual explanation. After pages have been crawled, PageRank data identifies the most highly linked pages. Running spjson.py reads the PageRank database stored in spider.json and writes JSON containing nodes and links. Nodes represent pages and can include a URL and PageRank score; links represent connections between source and target nodes. D3.js reads this JSON structure and renders the network.

  1. Crawl pages with spider.py so page and link records exist.
  2. Use the collected data for PageRank analysis.
  3. Execute spjson.py from the command line or terminal.
  4. Let spjson.py extract information about the most highly linked pages and write the JSON data.
  5. Open force.html in a browser.
  6. Inspect the D3 network, which uses nodes for pages and links for connections.
readwriterenderPageRank databasespider.jsonspjson.pycreates JSONJSONnodes and linksforce.htmlD3 visualization
What happens between stored PageRank results, spjson.py, and force.html?

Reading the Force Layout

When force.html opens, D3 displays pages as nodes and connections as links. A force-directed layout uses forces that push nodes apart to reduce overlap while links pull connected nodes together. The resulting arrangement is automatic, so the position of a node is a visual layout result rather than the page's URL or database ID.

selectclick and dragdouble-clickNetwork graphnodes and linksPage nodepagePage URLrevealed by double-clickPage noderepositioned
How does selecting or dragging a node reveal its URL and change the graph layout?

Use dragging to separate overlapping nodes or focus attention on a region. Use double-clicking when you need to identify the real page represented by a node, because the URL is what tells you which web page the node corresponds to.

EASY

Before opening the visualization, explain what information you expect to find in the JSON arrays named nodes and links.

Hints
  • A node represents a web page.
  • A link represents a connection between two pages.
  • A node can include a URL and PageRank score.

Refreshing New Crawl Data

The network is a snapshot of the data that was converted into JSON. If spider.py discovers new pages or updates link information, the existing visualization does not automatically contain those changes. Run spjson.py again to regenerate the JSON from the updated PageRank data, then press refresh in the browser. This reloads the visualization with the newer network data.

  • Opening force.html before generating the JSON data.

    The D3 visualization consumes the JSON representation rather than the raw database records.

    Fix: Execute spjson.py before opening or refreshing force.html.

  • Expecting a later spider.py run to re-fetch every stored page.

    The crawler is designed to skip pages already in the database.

    Fix: Understand later runs as incremental collection of unvisited pages.

  • Treating a node's screen position as its identity.

    The layout positions nodes automatically according to forces and connections.

    Fix: Double-click the node to reveal its URL.

  • Expecting updated crawl data to appear without regeneration and refresh.

    The visualization is based on previously generated JSON.

    Fix: Run spjson.py again and press refresh in the browser.

EASY

A second crawl adds pages to spider.sqlite. What two actions are needed before the new pages can appear in force.html?

Hints
  • The visualization reads generated JSON.
  • The browser must load the updated data.

Network Visualization Workflow

  1. Begin with collection: spider.py visits a starting URL, extracts links, and stores pages and relationships in spider.sqlite. Continue incrementally: later runs skip pages already stored and select from unvisited links. Combine sources when needed: multiple starting URLs can contribute separate webs to the same database. Analyze the graph: PageRank identifies highly linked pages. Convert the results: spjson.py writes nodes and links as JSON. Explore the result: force.html uses D3 to display the force-directed graph, while dragging changes node positions and double-clicking reveals URLs. When the crawl changes, regenerate the JSON and refresh the browser.

Key Takeaways

  • spider.py converts discovered web pages and their outgoing links into a persistent graph stored in spider.sqlite.
  • Repeated crawl sessions are additive because stored pages are skipped and unvisited links remain available.
  • Different starting URLs can coexist as separate webs within one database and unified crawl queue.
  • spjson.py converts PageRank data into JSON nodes and links for D3.
  • force.html provides a force-directed, interactive view that can be refreshed after new data is generated.