Concepts / HTTP Request Structure and Headers

HTTP Request Structure and Headers

Binary files must be accumulated in a bytes buffer and cannot be printed as they arrive like text files.

  • Programming

From Response to File

Retrieving an image over HTTP is not just a matter of receiving bytes and immediately saving everything that arrives. The response contains two different parts: HTTP headers that describe the response, followed by the binary image body. A reliable process must accumulate the complete response, find the boundary between those parts, discard the headers, and save only the image bytes.

sendsreturnsfollowed byfollowed byClientGET requestGET requestrequest line and headersServerHTTP responseResponse headersmetadataBlank lineCRLF CRLFImage bodybinary data
How are the request line, headers, blank line, and response body arranged within an HTTP exchange?

The central pattern has three stages: accumulate all response data in a bytes buffer, locate the double CRLF boundary, and write only the bytes after that boundary to a file opened in binary write mode.

Accumulating Socket Chunks

A socket does not necessarily deliver the complete HTTP response in one piece. The receiving process uses a loop to obtain data in chunks, typically 5120 bytes at a time, and concatenates each chunk to a growing buffer. The buffer must begin as an empty bytes object, represented by b"", because the response includes binary data rather than ordinary text.

Growing a Bytes Buffer

A server sends one HTTP response in three received chunks. Explain what the buffer represents after each chunk.

Start: The buffer is initialized as b"", an empty bytes object.

First chunk: The first received bytes are concatenated to the empty buffer, so the buffer now contains the beginning of the HTTP response.

Second chunk: The next received bytes are added with picture = picture + data. The buffer now contains the first and second chunks in order.

Third chunk: The third chunk is added in the same way. The buffer now contains the accumulated response data, including headers and the binary body received so far.

End of response: When receiving returns fewer than 1 byte, the server has closed the connection and no more response data is available.

The buffer is one bytes object containing the accumulated HTTP response rather than a collection of independently processed text fragments.

receivesreceivesreceivesconcatenateconcatenateconcatenateSocketresponse sourceChunk 1received bytesBytes bufferheaders and bodyChunk 2received bytesChunk 3received bytes
How do successive chunks of binary data arrive from the socket and combine into one complete buffer?

Finding the Header Boundary

After the socket closes, the bytes buffer contains both the HTTP response headers and the binary image body. HTTP separates these parts with a blank line represented by two consecutive CRLF sequences: \r\n\r\n. In a bytes buffer, the boundary marker is searched for with find(b'\r\n\r\n').

Locating the First Body Byte

Suppose find(b'\r\n\r\n') returns the position of the first byte in the four-byte boundary marker. Which position begins the image body?

Locate the marker: The find operation returns the index where the four-byte sequence \r\n\r\n begins.

Account for the marker: The boundary marker is four bytes long, so the marker itself occupies positions from pos through pos + 3.

Start the body slice: The binary body begins at pos + 4, immediately after the complete boundary marker.

Extract the body with a slice beginning at position pos + 4. This removes the HTTP headers and the boundary marker.

ends withfollowed byfollowed byHTTP headersmetadata bytesCRLFline endingCRLFline endingImage bodybinary bytes
Where do the CRLF header terminators occur in the byte buffer, and which bytes belong to the headers versus the image body?
ends beforefour bytes beforeHeader bytesbefore posCRLF CRLFat posImage bytesfrom pos + 4
How can the program locate the end of the HTTP headers before processing the remaining bytes as image data?

Do not hard-code a header length or assume that the body always starts at a fixed position. Locate the double CRLF marker in the actual response, because that marker identifies the boundary for that response.

Saving the Image Body

Once the boundary position has been found, the program slices the buffer from pos + 4 onward. That slice contains only the binary body. The body is then written to a file opened in binary write mode, represented by wb. Binary write mode avoids encoding transformations that could corrupt the image data.

receivesaccumulatesfind markerslice after markerwrite bytesSocket connectionHTTP GETHTTP responseheaders and bodyBytes buffercomplete responseCRLF boundarypos + 4Image bodybinary bytesOutput filebinary write mode
How does image data move from the HTTP response through the buffer and into the output file?
acceptswritesmay causeBinary modewbImage byteswritten directlySaved imagedata preservedText modeencoding transformationsCorrupted dataimage may fail
What is the difference between writing image bytes in binary mode and treating them as text?

Verifying the Response

The response headers include a Content-Type header describing the kind of data in the body. For an image, a value such as image/jpeg or image/png indicates an expected image type. Checking this header before saving is a defensive practice. If the response instead reports Content-Type: text/html, the server may have returned something other than the requested image, so the data should not be treated as an image.

Checking Content-Type Before Saving

A response has been separated into headers and body. The headers report either an image type or text/html. Decide how the Content-Type affects the next step.

Expected image type: A Content-Type such as image/jpeg or image/png supports the expectation that the body is an image.

Unexpected text type: A Content-Type of text/html indicates that the response is HTML rather than the expected image data.

Defensive decision: Use the header to decide whether the body should be saved as an image instead of blindly writing every response to an image file.

Content-Type does not replace separating the headers from the body, but it helps verify that the extracted body is the kind of data expected.

After the file is closed, the saved binary data can be opened in an image viewer to confirm that retrieval and saving worked correctly.

Mistakes That Break Binary Retrieval

  • Starting the buffer as text instead of bytes.

    Binary response data must be accumulated in a bytes buffer.

    Fix: Initialize the buffer as b"" and concatenate received byte chunks to it.

  • Saving the complete response as the image.

    The buffer contains both metadata and the binary body.

    Fix: Find the double CRLF marker and write only the slice beginning at pos + 4.

  • Using a fixed header position.

    The boundary must be located in the actual response rather than assumed.

    Fix: Use find(b'\r\n\r\n') to find the boundary dynamically.

  • Skipping the wrong number of bytes after the marker.

    The double CRLF marker is four bytes long, so those marker bytes would remain in the extracted data.

    Fix: Start the body at pos + 4.

  • Opening the output file in text mode.

    Encoding transformations can corrupt binary data.

    Fix: Open the output file in binary write mode, wb.

  • Ignoring Content-Type.

    The response may not contain the expected image.

    Fix: Check whether Content-Type indicates an image before saving it as one.

Trace the Complete Process

MEDIUM

A response arrives in several socket chunks and contains response headers followed by an image. Describe the processing sequence from the first received chunk to the final saved file. Include the buffer type, the marker used to locate the boundary, the offset used to begin the body slice, the file mode, and one check that helps confirm the body is an image.

Hints
  • Begin with an empty bytes buffer.
  • Accumulate every received chunk until the server closes the connection.
  • Search for the four-byte double CRLF marker.
  • The body begins four bytes after the marker position.
  • Use binary write mode and check Content-Type.
  1. A binary HTTP retrieval process preserves the response as bytes from beginning to end. It accumulates socket chunks into a bytes buffer, finds the header-body boundary with find(b'\r\n\r\n'), skips the four-byte marker by starting at pos + 4, and writes only the remaining binary body using binary write mode. Checking Content-Type helps confirm that the response contains the expected image rather than unrelated text.

Key Takeaways

  • Accumulate socket chunks in a bytes buffer initialized as b"".
  • HTTP response headers are followed by a blank line represented by \r\n\r\n.
  • Use find(b'\r\n\r\n') to locate the marker and begin the body slice at pos + 4.
  • Write only the extracted binary body to a file opened in binary write mode, wb.
  • Check Content-Type to help verify that the response contains the expected image data.