Reading and Writing Files in Binary Mode
Binary files must be accumulated in a bytes buffer and cannot be printed as they arrive like text files.
From Response to Image
Saving an image retrieved over HTTP is not the same as printing text received from a server. The response contains two different regions: HTTP headers and a binary body. A reliable process accumulates the response as bytes, finds the separator between those regions, removes the headers, and writes only the image bytes to a file opened in binary write mode.
Accumulating Socket Chunks
A socket does not necessarily provide the entire HTTP response in one arrival. The receiving process collects successive chunks and concatenates them into one growing bytes buffer. The buffer begins as an empty bytes object, represented by b"". Each received chunk is added to that buffer until recv() returns fewer than 1 byte, indicating that the server has closed the connection and no more data is available.
Following the Buffer State
Suppose a response arrives in several chunks. What does the buffer represent before the connection closes?
Initialize: The buffer is an empty bytes object, b"", rather than a text string.
Receive: Each chunk returned by the socket is concatenated to the growing bytes buffer.
Continue: Accumulation continues while data is available from the server.
Finish: When recv() returns fewer than 1 byte, the connection is closed and the buffer contains the HTTP headers and binary body.
The complete response is held as bytes before the program separates the headers from the image body.
Finding the Body Boundary
After accumulation, the buffer contains both the HTTP response headers and the binary image body. HTTP places a blank line between them. In the bytes representation, that blank line is the four-byte sequence \r\n\r\n: carriage return and line feed repeated twice. The bytes method find(b'\r\n\r\n') returns the index where this marker begins.
If find() returns pos, the marker begins at pos. Because the marker is four bytes long, the binary body begins at pos + 4. Slicing the buffer from pos + 4 onward removes the headers and the separator. The header region can also be inspected before that position to verify metadata such as Content-Type.
Applying the Boundary Rule
A complete response has a double CRLF marker at position pos. Which part should be retained as the image data?
Locate: Search the bytes buffer for b'\r\n\r\n' with find().
Measure: The marker occupies four bytes.
Skip: Start the slice at pos + 4 so the separator itself is not included.
Extract: The resulting slice contains only the binary body.
Retain the buffer from pos + 4 onward and discard the headers and the double CRLF marker.
Choosing the File Mode
| Mode | Use in this process | Reason |
|---|---|---|
| wb | Write the extracted image bytes | Binary write mode avoids encoding transformations that can corrupt the data |
| Text write mode | Not appropriate for the extracted image bytes | Text handling can apply encoding transformations |
Checking the Retrieved File
Before saving the body as an image, inspect the Content-Type header when possible. A value such as image/jpeg or image/png indicates the kind of image data expected. If the server instead returns Content-Type: text/html, the response may indicate that something went wrong, so it should not automatically be saved as an image file.
A successful retrieval follows the same three-step pattern: send an HTTP GET request and accumulate the complete response, find the double CRLF boundary and keep the bytes after it, then write those bytes to a file opened in wb mode. After the file is closed, opening it in an image viewer provides a practical confirmation that the retrieval and save process worked.
Mistakes That Damage Downloads
Starting with a text string instead of a bytes buffer
Binary data must be accumulated in a bytes buffer, and binary files cannot be handled like text files.
Fix:
Initialize the buffer as b"" and concatenate the received byte chunks.Hard-coding the header length
The header-body boundary must be located in the response rather than assumed.
Fix:
Use find(b'\r\n\r\n') to locate the boundary.Forgetting to skip the separator
The double CRLF marker is four bytes long and is not part of the binary body.
Fix:
Slice from pos + 4 onward.Opening the output file in text mode
Encoding transformations can corrupt binary data.
Fix:
Open the output file in binary write mode, wb.Ignoring Content-Type
The returned body may not be the expected image file.
Fix:
Check the header and verify that the content type matches the expected image type.
Practice the Data Flow
Describe the correct order of operations for retrieving an image over HTTP and saving it to disk. Your explanation should mention the initial buffer type, the marker used to separate headers from the body, the offset applied after finding that marker, and the file mode used for writing.
Hints
- Begin with the socket response and the growing buffer.
- The separator is a double CRLF represented as b'\r\n\r\n'.
- The separator occupies four bytes, so the body begins at pos + 4.
- The output file must use binary write mode, wb.
What do you think happens?
A complete response has been accumulated in a bytes buffer, and find(b'\r\n\r\n') returns pos. Where does the binary body begin?
Reveal answer
Answer: At pos + 4
The double CRLF marker is four bytes long. Starting at pos + 4 skips the marker and retains only the binary body.
Essential Sequence
- Accumulate every received chunk in a bytes buffer initialized as b"".
- Wait until the connection closes and the complete response is available.
- Find the header-body boundary with find(b'\r\n\r\n').
- Extract the binary body by slicing from the returned position plus 4.
- Write only the extracted body to a file opened in binary write mode, wb.
- Check Content-Type when possible so that an unexpected text response is not saved as an image.
Key Takeaways
- Binary network data should be accumulated as bytes rather than treated as text.
- The HTTP headers and binary body are separated by the four-byte marker \r\n\r\n.
- If find() returns pos, the body begins at pos + 4.
- Use binary write mode, wb, to avoid encoding transformations that can corrupt image data.
- Checking Content-Type helps confirm that the response contains the expected kind of file.