Concepts / HTTP Requests and Responses

HTTP Requests and Responses

Web scraping is a program-based technique for retrieving web pages and extracting specific data from their HTML source code, mimicking what a browser does but automating the data extraction process.

  • Programming

From Browsing to Automation

When you visit a website manually, you enter a URL, your browser sends a request to a server, and the server responds with HTML, CSS, and JavaScript. The browser renders that material into a readable page for you. Web scraping uses the same basic interaction, but changes who performs the reading: a program sends the request, retrieves the raw HTML, and extracts selected information automatically instead of displaying the page for a human to inspect.

Web scraping is a program-based technique for retrieving web pages and extracting specific data from their HTML source code, mimicking what a browser does while automating the data extraction process.

browser sendsserver response is renderedprogram sendsHTML is analyzedURLentered by a personURLprovided to a programHTTP requestsent by browserHTTP requestsent by programReadable pagerendered for a personTarget dataextracted from HTML
How are the actions of a manual browser and an automated scraper similar and different when retrieving a web page?

The Request and Response Cycle

The central mechanism can be followed as a sequence. First, a program sends an HTTP request for a web page. The server returns HTML in its response. The program then parses the HTML structure, identifies the information it wants, and extracts that information for storage or further processing. The important shift is that the program works with the underlying document rather than with the page as it appears to a human reader.

HTTP requestHTTP responseProgramWeb serverHTMLpage source
What happens between a program sending an HTTP request and receiving an HTML response from a web server?
  • Send an HTTP request
  • Receive HTML
  • Parse the HTML structure
  • Extract target data
  • Store or process the results

A Page Becomes a Dataset

Tracing a Focused Scrape

Suppose a program is intended to collect headlines from a news page. What happens from the moment the page is requested until the headlines are ready to use?

Request: The program requests the news page from its web server, performing the same basic retrieval action that a browser performs after a URL is entered.

Response: The server returns the page's HTML source. The program receives the underlying document rather than a human-readable rendered page.

Parsing: The program examines the structure of the HTML to locate the parts that contain headlines.

Extraction: The program selects the headline information instead of treating every part of the page as equally important.

Processing: The extracted headlines can be stored or processed as a focused collection of data.

The page has been transformed from a retrieved HTML document into a focused dataset. This is scraping because the goal is specific data extraction rather than broad discovery of many linked pages.

parseidentify and extractstore or processHTML sourcereceived documentParsed structureTarget dataselected informationProcessed resultsstored or analyzed
How does a scraper move from the received HTML document to identifying and extracting specific data?

Links Turn Pages into Paths

A scraper can stop after one page, but it can also extract hyperlinks from the HTML and follow them to other pages. This changes the task from isolated extraction to web traversal. The program retrieves one page, finds links in it, requests linked pages, extracts their links, and can repeat the process. Links therefore act as connections that let a program move from one document to another.

extract linksend requestextract linksend requestPage Aretrieved HTMLLink to Page Bextracted hyperlinkPage Bretrieved HTMLLink to Page Cextracted hyperlinkPage Cnext retrieved page
How does a scraper use links found on one page to discover and retrieve additional pages?

A focused scraper might retrieve only a handful of pages and collect one kind of information. A traversal-oriented program instead continues from page to page by following discovered links. The same retrieve, extract, and follow principle can support much broader discovery.

Search Engines at Scale

Search engines apply scraping principles on a much larger scale. Automated spiders retrieve pages, extract the hyperlinks embedded in those pages, visit linked pages, and repeat the process. This systematic traversal helps search engines discover and index publicly accessible pages across the web.

read pageinspect HTMLrequest linked pagesrepeat traversalorganize contentRetrieve pageautomated spiderCollect contentpage informationExtract linksembedded hyperlinksFollow linksdiscover pagesBuild indexorganized for search
How do search engines crawl pages, follow links, collect content, and organize it for search?
AspectWeb scrapingSearch-engine spidering
Primary purposeExtract a focused datasetDiscover pages and build an index
Typical scopeA page or a limited group of pagesLarge-scale, continuous traversal
Use of linksMay follow selected linksSystematically follows links across the web
Additional analysisCollects selected informationAlso evaluates relationships such as link frequency

Mistakes in the Mental Model

  • Treating scraping as manual browsing

    The program retrieves the raw HTML and analyzes the underlying data instead of presenting the page for a person to read.

    Fix: Think of the program as automating the retrieval and extraction steps performed around a browser visit.

  • Stopping the explanation at the HTTP request

    The scraping workflow also includes receiving HTML, parsing its structure, extracting target data, and storing or processing the results.

    Fix: Trace the complete workflow from request through extraction.

  • Using scraping and spidering as interchangeable terms

    Scraping is a general focused extraction technique, while spidering is a large-scale, continuous web traversal strategy used for indexing.

    Fix: Use scraping for focused extraction and spidering for broad link-following discovery and indexing.

  • Ignoring links after extracting page data

    Links provide connections that allow a scraper to retrieve and analyze subsequent pages.

    Fix: When traversal is required, extract links and use them as the next pages to retrieve.

A Practical Trace

What do you think happens?

A program retrieves one page, extracts three hyperlinks, and requests each linked page. Is this still web scraping, or has it become spidering?

  • It is always only web scraping
  • It is always only spidering
  • It uses scraping principles and may be part of spidering, depending on its purpose and scale
Reveal answer

Answer: It uses scraping principles and may be part of spidering, depending on its purpose and scale.

Both activities retrieve and analyze web pages. Focused extraction remains scraping, while systematic, large-scale, continuous traversal for discovery and indexing is spidering.

MEDIUM

Describe the state of the data after each step in this scenario: a program requests a page, receives its HTML, parses the structure, extracts selected information, and then finds links to two additional pages. Identify which step creates the opportunity for traversal.

Hints
  • Separate the retrieved document from the information selected from that document.
  • Links are part of the page information that can be extracted.
  • Traversal begins when extracted links are used to retrieve additional pages.
  1. Start with the request: identify which page the program asks the server to retrieve.
  2. Name the response: identify the HTML document returned by the server.
  3. Separate parsing from extraction: parsing examines structure, while extraction selects target information.
  4. Look for links: extracted hyperlinks can lead to additional requests.
  5. Classify the goal: focused collection suggests scraping, while broad continuous discovery and indexing suggests spidering.

The Complete Picture

  1. Web scraping is program-based retrieval and extraction of specific information from HTML. Its core cycle is to send an HTTP request, receive HTML, parse the structure, extract target data, and store or process the result. A scraper mimics the retrieval role of a browser but analyzes the underlying document instead of presenting it for a human reader. Extracted links can extend a one-page task into web traversal. Search-engine spiders apply these same principles continuously and at large scale to discover pages, collect content, build indexes, and analyze relationships such as link frequency.

Key Takeaways

  • Web scraping retrieves web pages programmatically and extracts selected information from their HTML source.
  • The fundamental workflow is request, response, parsing, extraction, and storage or processing.
  • A program can mimic a browser's retrieval action without rendering the page for a human reader.
  • Extracting and following links turns page retrieval into web traversal.
  • Search-engine spidering is a large-scale, continuous application of scraping principles for discovery and indexing.