Parsing and Processing Web Responses
HTTP responses follow a three-part structure: status line, headers, blank line, and body.
From Request to Structured Response
When a program connects to a web server and requests a file, the server does not simply send the file data. It sends a structured HTTP response. The response begins with information about the response and the document, then marks the end of that information with a blank line, and finally provides the actual file content. Parsing means recognizing these parts in their order instead of treating the entire response as one undifferentiated block.
The response has a fixed reading order: status line first, headers next, a blank line as the boundary, and the body last.
Reading the Response Boundary
The most important parsing boundary is the blank line after the headers. Everything before that delimiter belongs to the response description: the status line and header fields. The content after it is the response body, which contains the actual file content. A parser that fails to recognize the blank line can mistakenly treat body content as another header or fail to locate the beginning of the file.
Headers as Document Metadata
Headers describe the document and other properties of the response. Each header is a key-value pair: the header name acts as the key, and the associated value supplies the metadata for that key. This arrangement lets a parser distinguish one kind of information from another while processing the response.
Separating Metadata from Content
Consider this generated illustrative response layout and identify which parts are metadata and which part is body content.
Find the beginning: The first line is the status line, so it begins the structured response description.
Read the header fields: The lines before the blank line are headers. Each one supplies a key and a value describing the response or document.
Locate the delimiter: The blank line marks the end of the headers.
Read the remainder: The content after the blank line is the body, meaning the actual file content.
A correct parser separates the response into its status line, header metadata, blank-line delimiter, and body content.
Common Header Roles
Some headers appear frequently while processing web responses. Content-Type, Content-Length, and Server are named examples of headers with specific meanings. They should be read as metadata fields rather than as part of the file body. Date and Connection are also frequently encountered headers.
| Header | What to recognize |
|---|---|
| Content-Type | A frequently encountered header describing a document-related property. |
| Content-Length | A frequently encountered header describing a document-related property. |
| Server | A frequently encountered header describing a server-related property. |
| Date | Another frequently encountered HTTP header. |
| Connection | Another frequently encountered HTTP header. |
Common header names and their broad role as response metadata.
Parsing Mistakes to Avoid
Treating the entire response as file content.
The response begins with structured information about the response and document before the actual file content appears.
Fix:
Separate the status line and headers from the body before processing the file content.Ignoring the blank line.
The blank line is the delimiter between headers and the response body.
Fix:
Treat the blank line as the signal that header parsing has ended and body processing should begin.Confusing a header with body content.
Headers are metadata describing the document, while the body contains the actual file content.
Fix:
Classify content by its position relative to the blank-line delimiter.Failing to recognize the key-value structure of headers.
Headers are key-value pairs that provide named pieces of metadata.
Fix:
Read each header as a key associated with a value before continuing to the next response part.
Practice the Boundary
A response has a status line, several header fields, a blank line, and then file content. Explain what a parser should identify first, what it should collect before the blank line, and what the blank line tells it to do next.
Hints
- Start with the required order of the response parts.
- Remember that headers describe the document and use key-value pairs.
- Use the blank line as the boundary between metadata and body content.
What do you think happens?
If a parser has read the status line and several headers, what should it conclude when it reaches the blank line?
Reveal answer
Answer: It should conclude that the header section has ended and that the response body begins after the blank line.
The blank line is the delimiter separating the headers from the actual file content.
Response Parsing Checklist
- Read the HTTP response in order: status line, headers, blank line, then body.
- Treat headers as key-value pairs that describe the document and response.
- Use the blank line as the delimiter between metadata and actual file content.
- Recognize Content-Type, Content-Length, Server, Date, and Connection as frequently encountered headers.
- Do not ignore the blank line: failing to use it can cause parsing errors.
Key Takeaways
- An HTTP response consists of a status line, headers, a blank-line delimiter, and a body.
- Headers are key-value pairs that provide metadata about the document or response.
- The blank line marks the exact boundary between the headers and the actual file content.
- Content-Type, Content-Length, Server, Date, and Connection are common HTTP headers to recognize.
- Ignoring the blank line or confusing headers with body content leads to parsing errors.