Concepts / Reading and Writing Files in Binary Mode

Reading and Writing Files in Binary Mode

Binary files must be accumulated in a bytes buffer and cannot be printed as they arrive like text files.

  • Programming

From Response to Image

Saving an image retrieved over HTTP is not the same as printing text received from a server. The response contains two different regions: HTTP headers and a binary body. A reliable process accumulates the response as bytes, finds the separator between those regions, removes the headers, and writes only the image bytes to a file opened in binary write mode.

server returnsaccumulate bytesfind markerslice after markerwrite binary dataHTTP GET requestrequest sent to serverHTTP responseheaders plus binary bodybytes bufferall received bytesdouble CRLFheader-body boundarybinary bodyimage bytes onlysaved image filewritten with wb
How does image data move from the HTTP response through the buffer and into the output file?

Accumulating Socket Chunks

A socket does not necessarily provide the entire HTTP response in one arrival. The receiving process collects successive chunks and concatenates them into one growing bytes buffer. The buffer begins as an empty bytes object, represented by b"". Each received chunk is added to that buffer until recv() returns fewer than 1 byte, indicating that the server has closed the connection and no more data is available.

add chunk 1concatenatecontinue accumulationadd chunk 2empty bufferb""chunk 1received bytesbuffer after chunk 1headers and partial datachunk 2more received bytescomplete bufferheaders plus full body
How do successive chunks of bytes from a network socket accumulate into one complete image buffer?

Following the Buffer State

Suppose a response arrives in several chunks. What does the buffer represent before the connection closes?

Initialize: The buffer is an empty bytes object, b"", rather than a text string.

Receive: Each chunk returned by the socket is concatenated to the growing bytes buffer.

Continue: Accumulation continues while data is available from the server.

Finish: When recv() returns fewer than 1 byte, the connection is closed and the buffer contains the HTTP headers and binary body.

The complete response is held as bytes before the program separates the headers from the image body.

Finding the Body Boundary

After accumulation, the buffer contains both the HTTP response headers and the binary image body. HTTP places a blank line between them. In the bytes representation, that blank line is the four-byte sequence \r\n\r\n: carriage return and line feed repeated twice. The bytes method find(b'\r\n\r\n') returns the index where this marker begins.

ends beforeseparatesretain body onlyHTTP headersbefore \r\n\r\nbinary image bodyslice from pos + 4double CRLF\r\n\r\nbinary image bodyafter marker
How do the CRLF delimiters identify where the HTTP headers end and the binary image body begins?

If find() returns pos, the marker begins at pos. Because the marker is four bytes long, the binary body begins at pos + 4. Slicing the buffer from pos + 4 onward removes the headers and the separator. The header region can also be inspected before that position to verify metadata such as Content-Type.

Applying the Boundary Rule

A complete response has a double CRLF marker at position pos. Which part should be retained as the image data?

Locate: Search the bytes buffer for b'\r\n\r\n' with find().

Measure: The marker occupies four bytes.

Skip: Start the slice at pos + 4 so the separator itself is not included.

Extract: The resulting slice contains only the binary body.

Retain the buffer from pos + 4 onward and discard the headers and the double CRLF marker.

Choosing the File Mode

ModeUse in this processReason
wbWrite the extracted image bytesBinary write mode avoids encoding transformations that can corrupt the data
Text write modeNot appropriate for the extracted image bytesText handling can apply encoding transformations
write bytesmay transform datawbwrites binary bytestext write modetext handlingimage filebinary data preservedcorrupted dataencoding transformations
What is the difference between opening a file in binary write mode and text write mode when saving image bytes?

Checking the Retrieved File

send HTTP GET requestreturn headers and image byteswrite body with wbprogramserveroutput file
What happens in sequence when a program requests an image, receives the response, and completes the download?

Before saving the body as an image, inspect the Content-Type header when possible. A value such as image/jpeg or image/png indicates the kind of image data expected. If the server instead returns Content-Type: text/html, the response may indicate that something went wrong, so it should not automatically be saved as an image file.

A successful retrieval follows the same three-step pattern: send an HTTP GET request and accumulate the complete response, find the double CRLF boundary and keep the bytes after it, then write those bytes to a file opened in wb mode. After the file is closed, opening it in an image viewer provides a practical confirmation that the retrieval and save process worked.

Mistakes That Damage Downloads

  • Starting with a text string instead of a bytes buffer

    Binary data must be accumulated in a bytes buffer, and binary files cannot be handled like text files.

    Fix: Initialize the buffer as b"" and concatenate the received byte chunks.

  • Hard-coding the header length

    The header-body boundary must be located in the response rather than assumed.

    Fix: Use find(b'\r\n\r\n') to locate the boundary.

  • Forgetting to skip the separator

    The double CRLF marker is four bytes long and is not part of the binary body.

    Fix: Slice from pos + 4 onward.

  • Opening the output file in text mode

    Encoding transformations can corrupt binary data.

    Fix: Open the output file in binary write mode, wb.

  • Ignoring Content-Type

    The returned body may not be the expected image file.

    Fix: Check the header and verify that the content type matches the expected image type.

Practice the Data Flow

MEDIUM

Describe the correct order of operations for retrieving an image over HTTP and saving it to disk. Your explanation should mention the initial buffer type, the marker used to separate headers from the body, the offset applied after finding that marker, and the file mode used for writing.

Hints
  • Begin with the socket response and the growing buffer.
  • The separator is a double CRLF represented as b'\r\n\r\n'.
  • The separator occupies four bytes, so the body begins at pos + 4.
  • The output file must use binary write mode, wb.

What do you think happens?

A complete response has been accumulated in a bytes buffer, and find(b'\r\n\r\n') returns pos. Where does the binary body begin?

  • At pos
  • At pos + 1
  • At pos + 4
  • At the end of the buffer
Reveal answer

Answer: At pos + 4

The double CRLF marker is four bytes long. Starting at pos + 4 skips the marker and retains only the binary body.

Essential Sequence

  1. Accumulate every received chunk in a bytes buffer initialized as b"".
  2. Wait until the connection closes and the complete response is available.
  3. Find the header-body boundary with find(b'\r\n\r\n').
  4. Extract the binary body by slicing from the returned position plus 4.
  5. Write only the extracted body to a file opened in binary write mode, wb.
  6. Check Content-Type when possible so that an unexpected text response is not saved as an image.

Key Takeaways

  • Binary network data should be accumulated as bytes rather than treated as text.
  • The HTTP headers and binary body are separated by the four-byte marker \r\n\r\n.
  • If find() returns pos, the body begins at pos + 4.
  • Use binary write mode, wb, to avoid encoding transformations that can corrupt image data.
  • Checking Content-Type helps confirm that the response contains the expected kind of file.