Concepts / Introduction to Web Scraping

Introduction to Web Scraping

HTTP responses follow a three-part structure: status line, headers, blank line, and body.

  • Programming

The Response Before the Page

When a program connects to a web server and requests a file, the server does not send only the file data. It sends a structured HTTP response. The response begins with information about the document, continues with metadata in HTTP headers, and then contains the actual file content. A web scraper must recognize this structure before it can correctly separate metadata from the document it wants to process.

followed byends beforeseparates fromStatus lineresponse statusHeadersdocument metadataBlank linesection boundaryBodyfile content
What comes next as the response moves from its opening information to the document content?

Tracing the Four-Part Layout

The response has a fixed sequence. First comes the status line. Next come the headers, which provide metadata describing the document. After the headers comes a blank line. Finally, the body contains the actual file content. Although the topic is often described as a three-part response, the blank line is an essential delimiter inside that structure: it marks the boundary between the headers and the body.

Following a Response from Start to Finish

Suppose a server sends a response containing a status line, several headers, a blank line, and document content. What should a scraper identify at each stage?

First component: Identify the status line at the beginning of the response.

Metadata section: Read the headers that describe the document. Each header is a key-value pair.

Boundary: Find the blank line. It marks the end of the headers.

Content section: Treat the data after the blank line as the response body, which contains the actual file content.

The response is interpreted in order as status line, headers, blank line, and body.

precedesfollowed byseparatesStatus lineresponse statusHeadersdocument metadataBlank lineheader-body boundaryBodyactual file content
What information belongs to the status line, headers, boundary, and body, and how are these parts different?

Reading Header Pairs

HTTP headers are metadata about the document in the response. They are organized as key-value pairs: the header name identifies the kind of information being supplied, and the associated value supplies that information. Headers appear after the status line and before the blank line. This position matters because it tells a parser that the pairs belong to the metadata section rather than to the body.

containscontainscontainsHeadersbetween status line andblank lineContent-Typedocument typeContent-Lengthdocument lengthServerserver information
How are header names connected to their values, and where do those pairs appear in the response?
HeaderWhat it describes
Content-TypeThe type of the document in the response
Content-LengthThe length of the document in the response
ServerInformation about the server sending the response
DateDate-related information associated with the response
ConnectionConnection-related information associated with the response

Frequently encountered HTTP headers and the kind of information they provide

Finding the Body Boundary

The blank line is not empty information to discard casually. It is the delimiter that separates the header section from the response body. A parser that recognizes this boundary can treat everything before it as response metadata and everything after it as the actual file content. Ignoring the delimiter is a common source of parsing errors because metadata and document content may then be mixed together.

followed byfollowed byends atseparates fromStatus lineresponse statusHeader pairmetadataHeader pairmetadataBlank linedelimiterResponse bodyfile content
How does the blank line show where the header section ends and the response body begins?

Common Header Meanings

Several headers appear frequently when examining HTTP responses. Content-Type describes the type of the document. Content-Length describes the document's length. Server provides information about the server sending the response. Date and Connection are also frequently encountered and provide information associated with the response date and connection. These values help a scraper understand the metadata that accompanies the body rather than mistaking that metadata for the document itself.

Interpreting Three Header Names

A response contains the header names Content-Type, Content-Length, and Server. What kind of information should a scraper associate with each one?

Content-Type: Associate this pair with information about the type of the document.

Content-Length: Associate this pair with information about the document's length.

Server: Associate this pair with information about the server that sent the response.

The three headers describe different aspects of the response metadata: document type, document length, and server information.

Parsing Mistakes to Avoid

  • Treating the response as only the file content

    The server sends a structured response containing metadata before the actual file content.

    Fix: Identify the status line and headers before processing the body.

  • Ignoring the blank line

    The blank line is the delimiter between headers and the body.

    Fix: Use the blank line to separate metadata from the actual file content.

  • Treating headers as ordinary body content

    Headers describe the document; they are not the document body.

    Fix: Keep header metadata separate from the body after locating the delimiter.

Practice the Sequence

EASY

A response is presented as four consecutive regions: status information, several key-value pairs describing a document, an empty line, and file content. Name each region in order and explain why the empty line is important.

Hints
  • The first region is the status line.
  • The key-value pairs are the headers.
  • The empty line is the delimiter between headers and body.
  • The final region is the response body.

What do you think happens?

A parser has read several headers and then encounters the blank line. What should it identify next?

  • Another header section
  • The response body
  • The beginning of the status line
  • The end of the entire response structure
Reveal answer

Answer: The response body

The blank line separates the headers from the actual file content, so the data after it belongs to the body.

Key Takeaways

  1. An HTTP response begins with a status line, continues with headers, includes a blank line, and ends with the body.
  2. Headers are key-value pairs that provide metadata describing the document and response.
  3. The blank line is the delimiter that separates headers from the actual file content.
  4. Content-Type describes document type, Content-Length describes document length, and Server provides server information.
  5. Correct parsing depends on preserving the distinction between metadata and body content.

Key Takeaways

  • HTTP responses have a sequential structure: status line, headers, blank line, and body.
  • Headers are key-value pairs that describe the document and other response information.
  • The blank line is the boundary between metadata and the actual file content.
  • Content-Type, Content-Length, and Server communicate document type, document length, and server information.
  • A scraper must recognize the response structure to parse web responses correctly.