Data Validation and Cleaning
The find() method locates the position of a substring and returns its starting index; use find(substring, start_pos) to search from a specific position onward.
From Mixed Text to Useful Data
Real-world data often arrives in mixed formats. A single line can contain an email address, a timestamp, and other information together. Parsing means breaking that larger string down to isolate the exact substring you need. In Python, a practical way to do this is to use find() to locate landmarks and string slicing to extract the characters between those landmarks.
The central workflow is: locate a starting landmark, locate an ending landmark, then slice between their positions.
Reading Positions with find()
The find() method locates a substring inside a string and returns the substring's starting index. An index identifies a character position in the larger string. For example, when find('@') returns 21, the @ character begins at position 21. The returned number is a position, not the substring itself.
The two-argument form find(substring, start_pos) begins searching at the specified position and continues onward. This is useful when the same kind of character or substring may appear more than once and you need the occurrence after a known landmark.
Turning Landmarks into a Slice
Extracting an Email Domain
Use the position of @ and the next space to isolate the domain name from a larger string.
Find the first landmark: find('@') returns 21, so the @ symbol begins at index 21.
Find the ending landmark: find(' ', atpos) searches for a space beginning at position 21 and returns 31. This identifies the first space after the @ symbol.
Set the slice start: Use atpos+1 as the start. This changes the starting position from the @ symbol at 21 to the first character after it at 22.
Set the slice end: Use sppos as the end. A slice stops before its end index, so the space at index 31 is excluded.
The slice data[atpos+1:sppos] extracts uct.ac.za.
String slicing with [start:end] extracts characters from start up to but not including end. In this example, the start is atpos+1, which skips the @ symbol, and the end is sppos, which excludes the space. The indices therefore describe exactly which characters become the extracted value.
Tracing a Parsing Failure
When parsing does not produce the expected result, trace the indices step by step. First inspect the intermediate values returned by find(). Confirm that each landmark is at the position you expect. Then inspect the slice start and end positions. A mistake in either landmark can cause the slice to include the wrong characters or omit needed characters.
Mistakes with Indices
Treating the value from find() as the extracted text
The returned value is the starting index of the @ character, not the domain or another substring.
Fix:
Use the returned index as a landmark for a later slice.Including the delimiter in the extracted value
The desired domain begins one position after the @ character.
Fix:
Use atpos+1 when the delimiter itself should be excluded.Assuming the slice includes its end position
A slice extracts up to but not including its end index.
Fix:
Remember that [start:end] stops before end.Using a -1 result as though it were a valid landmark
The expected substring was not found, so the planned boundary is unavailable.
Fix:
Check the result before using it in a slice and inspect intermediate values while debugging.Searching from the beginning when a later occurrence is needed
The search may locate a space before the relevant landmark.
Fix:
Pass the known landmark as the starting position, as in find(' ', atpos).
Guided Practice
A larger string contains an email address followed by a space and additional information. Describe the sequence of positions you would identify to extract the domain: first locate the @ symbol, then search for the next space beginning at the @ position, and finally slice from one position after @ up to the space. What should you check before performing the slice?
Hints
- The first find() result identifies the position of @.
- Pass that position as the starting position for the search for the next space.
- The slice excludes its end index.
- Check that each find() result is not -1 before using it as a boundary.
What do you think happens?
If find('@') returns 21 and the next space found from position 21 is at 31, which positions should define the slice that extracts the domain without the @ symbol or the space?
Reveal answer
Answer: 22 through 31 as an exclusive end
The @ symbol is at position 21, so the slice starts at 21+1, or 22. The end position is 31, and slicing stops before that position, excluding the space.
Parsing Workflow
- Identify a meaningful landmark in the larger string with find().
- Store or inspect the starting index returned by find().
- Use find(substring, start_pos) when the search must begin after a known position.
- Convert the landmark positions into slice boundaries.
- Remember that [start:end] includes start and excludes end.
- Check every find() result for -1 before using it in a slice.
- Print intermediate index values when debugging unexpected parsing results.
Parsing becomes manageable when you treat each index as a boundary in the original string. find() tells you where a landmark begins; a starting position narrows a later search; slicing uses the resulting boundaries to isolate the needed characters. If the result is wrong, inspect those boundaries in order rather than guessing at the final substring.
Key Takeaways
- find() returns the starting index of a located substring.
- find(substring, start_pos) searches from a specified position onward.
- A slice [start:end] includes start but excludes end.
- Use find() landmarks as slice boundaries to extract data between delimiters.
- Check for -1 and inspect intermediate indices before trusting parsing results.