Reading Files Line-by-Line with a for Loop
read() loads an entire file into a single string, including all newline characters, and exhausts the file handle resource after one call.
One Read, One Position
A file on disk can contain much more data than your program should load into memory at once. Python offers two important strategies: read the complete file into one string with read(), or iterate through the file one line at a time with a for loop. The key difference is how much data is held in memory and how the file's current position changes while it is being read.
What do you think happens?
A file contains three lines. After read() is called once without a size argument, what will a second read() call on the same file handle return?
Reveal answer
Answer: An empty string
A complete read starts at the current position and continues to the end of the file. Afterward, the file handle's position is already at the end, so another read() has no remaining data to return.
What read() Returns
Calling read() without a size argument reads from the current file position through the end of the file. Python returns the complete contents as one string, including the newline characters that separate lines. Assign that returned string to a variable immediately. The variable preserves the data, while the file handle's position continues to move.
file_handle = open("notes.txt") contents = file_handle.read() print(contents)
Calling read() twice and expecting the complete file both times.
The first complete read moves the position to the end of the file, so the second call has no remaining data.
Fix:
Assign the first read() result immediately and reuse that variable when you need the contents again.Assuming the file handle itself contains all file data.
The file handle is a connection to the file on disk. It maintains a reading position, but it does not contain the file's contents.
Fix:
Use read() to obtain a complete string, or iterate over the handle to fetch lines incrementally.
Following the File Loop
A for loop over a file handle causes Python to fetch the next line during each iteration. The open function creates the connection but does not pull the entire file into memory. The loop body runs once for each line, and the loop variable receives that line, including its newline character when the line has one.
In this pattern, line is not the entire file. On the first iteration it refers to the first line, on the next iteration it refers to the next line, and so on. After a line has been processed, it can be discarded before the next line is fetched. The loop therefore supports processing a very large file without loading all of its contents into one string.
Newline Boundaries
A line variable normally includes the newline character at the end of that line. That character is how Python represents the boundary between one line and the next, and the for loop does not remove it when placing the line in the loop variable. The complete string returned by read() also includes newline characters, but it contains all of them in one string rather than one line variable at a time.
Using repr() in this demonstration makes the newline character visible in the displayed representation. For a line that ends normally, the representation shows the newline as part of the string. The final line may not have a newline character if the file does not end with one, so the loop variable's ending characters depend on the file's actual contents.
Choosing by File Size
| Strategy | Data held during processing | Best fit |
|---|---|---|
| read() | The complete remaining file in one string | A file small enough to fit in available RAM |
| for line in file_handle | One line at a time | Large files or processing where memory efficiency matters |
Selecting a Strategy
Choose a reading strategy for a file that is 1 GB in size when the program has only a smaller amount of available RAM.
Estimate what read() would require: read() attempts to place the complete file contents into one string, so the file must fit in available memory.
Compare the file with available memory: A 1 GB file requires at least 1 GB of free memory according to the source guidance, which is not available in this situation.
Select incremental processing: Iterate over the file handle with a for loop so each line is read, processed, and then discarded before the next line arrives.
Use a for loop over the file handle rather than read().
Mistakes in File Loops
Using read() for a file that may not fit in available RAM.
read() loads the entire file into a single string, so memory use grows with the file size.
Fix:
Use a for loop over the file handle to keep only the current line in memory.Thinking that opening a file has already loaded all its contents.
The file handle is a connection to the file on disk, not a container holding all file data.
Fix:
Remember that iteration triggers the fetching of each next line.Assuming a line variable never contains a newline character.
Each iteration normally includes the line's newline character.
Fix:
Account for the newline when processing or displaying a line.Calling read() again after the complete contents have already been saved.
The file position is at the end after the first complete read, so the later call returns an empty string.
Fix:
Reuse the variable holding the first read() result.
Practice the Decision
A program needs to count the lines in a very large file. Write the file-reading pattern that keeps memory use low, and explain what the loop variable contains during one iteration.
Hints
- Use a for loop directly with the file handle.
- Increment a counter inside the loop.
- The loop variable contains one line, normally including its newline character.
A small file must be used several times after it is opened. Decide whether read() is suitable, and state what should happen immediately after the call to read().
Hints
- Consider whether the complete file can fit in available RAM.
- Assign the returned string to a variable immediately.
- Do not expect a second complete read from the same position.
Key Takeaways
- read() returns the complete remaining file contents as one string, including newline characters.
- After a complete read(), the file position is at the end, so another read() on that handle returns an empty string.
- A file handle is a connection to the file on disk and maintains a reading position; it does not contain all file data.
- A for loop over a file handle fetches one line per iteration, normally including that line's newline character.
- Use read() for files that fit in available RAM and line-by-line iteration when memory efficiency is important, especially for large files.
Key Takeaways
- read() loads a complete file into one string and moves the file position to the end.
- A second read() on the same handle returns an empty string because no data remains at the current position.
- The file handle is a connection with a position pointer, not a copy of the file's contents.
- A for loop reads one line at a time and normally keeps only the current line in memory.
- Choose read() for suitably small files and line-by-line iteration for large files or memory-sensitive processing.