Concepts / String Methods and Indexing

String Methods and Indexing

Email headers follow a predictable format; the day of the week is always the third word (index 2) on lines starting with 'From'.

  • Programming

From Headers to Useful Data

Email headers contain more information than the part you need for a particular analysis. In this task, the goal is to count how many messages arrived on each day of the week. The useful clue is that email headers follow a predictable format: lines beginning with From contain a timestamp, and the day of the week is always the third word on that line.

The solution has two connected parts. First, isolate the day from each relevant line by splitting the line into words and selecting index 2. Second, use that day as a dictionary key and update its count. Because the program processes multiple records, the dictionary represents a running summary that changes after each matching line.

next wordnext wordnext wordFromindex 0senderindex 1Monindex 2timeindex 3
How does the day of the week at the third word map to index 2 after a header line is split?

First Extraction Trace

Finding the Day

Extract the day from a relevant header line whose words include From, a sender, Mon, and a timestamp.

Filter: The line is considered because it begins with From.

Split: Splitting the line by spaces produces a list of words.

Index: The third word is at index 2 because indexing starts at 0.

Extract: The value at index 2 is the day, Mon, for this generated example.

The extracted day is Mon.

split by spacesselect index 2use for countingFrom lineword listdaydictionary key
How does a day extracted from a string become the key used in dictionary analysis?

Filtering Before Parsing

The program should not extract a day from every line in the file. It first checks whether a line starts with From. Only those lines enter the splitting and extraction process. This filtering step matters because the exercise identifies message records through the From prefix.

python
checknoyesindex 2input lineFrom prefixskip linesplit wordsday at index 2
How does the program decide which lines enter extraction and counting?

Updating the Running Count

After extracting a day, update the dictionary that stores the counts. If the day is already a key, increment its count. If it is not yet a key, add it with a count of 1. This allows one dictionary to accumulate results across all matching email header lines.

counts = {} for day in extracted_days: if day in counts: counts[day] = counts[day] + 1 else: counts[day] = 1 print(counts)

day Monday Monday Tue{}before records{Mon: 1}after Mon{Mon: 2}after Mon{Mon: 2, Tue: 1}after Tue
How does each extracted day change the running dictionary count?

Tracing Iteration State

When the final dictionary is wrong, inspect the program while it is running rather than looking only at the final result. Print the current line, the extracted day, and the dictionary after each update. This reveals whether the error occurs during filtering, extraction, or counting.

MonTueMon{}before iteration 1{Mon: 1}after iteration 1{Mon: 1, Tue: 1}after iteration 2{Mon: 2, Tue: 1}after iteration 3
What does the dictionary contain after each iteration, and where would an incorrect count first appear?
  • Using the wrong index

    Python uses zero-based indexing, so the third word is at index 2.

    Fix: Split the relevant line and read the value at index 2.

  • Processing every line

    The day extraction rule applies to the relevant From lines.

    Fix: Filter for lines beginning with From before splitting.

  • Printing the dictionary too early

    The dictionary is a running count, so an early report is incomplete.

    Fix: Print the final dictionary after the iteration finishes.

  • Skipping intermediate checks

    The first incorrect state may have been caused by filtering or extraction.

    Fix: Print the line, extracted day, and dictionary state during debugging.

Putting the Process Together

python

This complete structure follows the required order: read each line, filter for the From prefix, split the selected line, extract the day at index 2, update the dictionary, and print the result after processing. The dictionary is the connection between string manipulation and analysis: the extracted string becomes a key whose value records its frequency.

MEDIUM

Suppose the extracted days arrive in this order: Mon, Tue, Mon. Trace the dictionary after each extracted day. Then identify which line of the counting logic handles the first Mon and which line handles the second Mon.

Hints
  • A day seen for the first time must receive a count of 1.
  • A day already present as a key must have its count increased.
  • Write the dictionary state after every extracted day.

Key Takeaways

  1. Filter for lines beginning with From before attempting extraction.
  2. Split a relevant line into words and select index 2 for the third word, the day of the week.
  3. Use a dictionary as a running count: initialize a new day at 1 and increment an existing day.
  4. Debug by tracing the line, extracted day, and dictionary state after each update.
  5. Print the dictionary only after all relevant lines have been processed.

Key Takeaways

  • Email headers provide predictable structure that makes targeted extraction possible.
  • The day is the third word on a relevant From line, so its zero-based index is 2.
  • A dictionary can accumulate one count for each day as records are processed.
  • Intermediate state tracing helps locate the first incorrect filtering, extraction, or counting step.