Concepts / Error Handling in Network Applications

Error Handling in Network Applications

Binary files must be accumulated in a bytes buffer and cannot be printed as they arrive like text files.

  • Programming

The Complete Retrieval Problem

Retrieving an image over HTTP involves more than receiving bytes and placing them directly in a file. The response first contains HTTP headers, followed by a blank line, and then the binary image body. A reliable process must accumulate the complete response, identify the exact boundary between headers and body, discard the headers, and save only the image bytes.

establishesreceivesaccumulateslocatesextracts afterwrites in wb modeSocket connectionHTTP GET requestHTTP responseBytes bufferheaders and image bodyDouble CRLFheader-body boundaryBinary bodyImage file
What is the complete sequence from the HTTP response arriving at the socket to the image being written to disk?

Accumulating Chunks as Bytes

The response arrives through the socket in chunks rather than as one already-separated image. The buffer must therefore begin as an empty bytes object, represented as b"". Each received chunk is concatenated to that growing buffer. The data remains binary throughout this process; it is not treated as printable text.

The accumulation continues until the socket reports that no more data is available. In the source procedure, recv() returns fewer than 1 byte when the server has closed the connection. At that point, the buffer contains both the HTTP headers and the binary image body.

concatenateconcatenateconcatenategrows toChunk 1bytesChunk 2bytesChunk Nbytespictureb""pictureheaders plus image bytes
How do successive chunks of binary data move from the socket and accumulate in one buffer without being treated as printable text?

Finding the Header Boundary

After receiving all data, the buffer contains two regions: the HTTP headers at the beginning and the binary body after them. HTTP separates these regions with a blank line represented by two consecutive carriage-return and line-feed sequences: \r\n\r\n.

Use find(b'\r\n\r\n') on the bytes buffer to locate the position where the four-byte boundary marker begins. The returned position identifies the start of the marker, not the start of the image. Because the marker is four bytes long, the binary body begins at position plus 4. Slicing from that point removes the headers and the separator together.

followed byfollowed byHTTP headersmetadata\r\n\r\n4 bytesImage bytesbinary body
Where do the CRLF markers occur, and how do they identify the exact boundary between the HTTP headers and the image bytes?

Extracting the Body

A complete HTTP response has been accumulated in a bytes buffer. The double CRLF begins at position pos. Which part should be saved as the image?

Locate the marker: Find the position where b'\r\n\r\n' begins in the buffer.

Skip the separator: The marker occupies four bytes, so the image begins at pos + 4.

Slice the body: Take the buffer from pos + 4 onward. This excludes the headers and the boundary marker.

The bytes from position pos + 4 onward are the binary body to write to the image file.

Saving Bytes Without Transformation

Once the binary body has been extracted, write it to a file opened in binary write mode, represented as wb. Binary mode preserves the received data. Text mode can apply encoding transformations, which can corrupt binary image data.

writeavoids transformationwritemay transformBinary bodywbText modePreserved bytesCorrupted data
What changes when image bytes are written with binary mode instead of text mode, and why does binary mode preserve the file correctly?

Defensive Checks Before Saving

The response headers contain a Content-Type header describing the kind of data in the body. For an expected image, a value such as image/jpeg or image/png supports the conclusion that the body is an image. If the server instead returns Content-Type: text/html, the response may not be the requested image, so the data should not be saved as an image file.

inspectfoundnot foundexpected image typetext/htmlcompleteResponse dataDouble CRLFContent-TypeBinary writeDo not saveSaved image
How does control flow respond when receiving data, finding the header boundary, or checking the file type produces a problem?

Treat the boundary search and Content-Type check as gates in the process. Without a located boundary, the headers cannot be separated reliably. Without an expected image Content-Type, saving the body under an image filename may hide the fact that the server returned different data.

Common Separation and Storage Mistakes

  • Initializing the buffer as text instead of bytes

    The image is binary data and must be accumulated in a bytes buffer.

    Fix: Start with an empty bytes object, b"", and concatenate received binary chunks.

  • Saving the entire response

    The beginning of the buffer contains HTTP headers, not image bytes.

    Fix: Find b'\r\n\r\n' and write only the slice beginning at the boundary position plus 4.

  • Hard-coding the header length

    The boundary must be located in the actual response rather than assumed.

    Fix: Use find(b'\r\n\r\n') for each accumulated response.

  • Skipping the wrong number of bytes

    The double CRLF marker itself is four bytes and is not part of the image body.

    Fix: Start at position plus 4.

  • Opening the output file in text mode

    Text-mode encoding transformations can corrupt binary data.

    Fix: Open the output file in wb mode.

  • Ignoring Content-Type

    The server may have returned something other than the expected image.

    Fix: Inspect Content-Type and verify that it indicates an image such as image/jpeg or image/png.

Practice the Data Trace

MEDIUM

A socket response has been accumulated in a bytes buffer. Describe the exact order of operations needed to save the image safely. Include the buffer type, the boundary marker, the slice starting position, the Content-Type check, and the file mode.

Hints
  • The buffer must contain both the headers and body before the boundary is located.
  • The boundary marker is four bytes long.
  • Only the binary body should be written to the output file.

What do you think happens?

The boundary marker begins at position pos. Which position contains the first byte of the image body?

  • pos
  • pos + 1
  • pos + 4
Reveal answer

Answer: pos + 4

The marker b'\r\n\r\n' is four bytes long, so the image body begins after all four marker bytes.

Process Checklist

  1. Establish the socket connection and send an HTTP GET request for the binary file.
  2. Initialize an empty bytes buffer and concatenate each received chunk until the server closes the connection.
  3. Locate the header-body boundary with find(b'\r\n\r\n').
  4. Extract the binary body by slicing from the boundary position plus 4.
  5. Check Content-Type to confirm that the response represents the expected image.
  6. Write only the extracted body to a file opened in binary write mode, wb.

Key Takeaways

  • Accumulate network response chunks in a bytes buffer because an image is binary data.
  • The HTTP headers end at the double CRLF marker, b'\r\n\r\n'.
  • If the marker begins at pos, the binary body begins at pos + 4.
  • Check Content-Type before treating the body as an image.
  • Write the extracted body using binary write mode, wb, to avoid corrupting the file.