Working with Multiple Strings
The find() method locates the position of a substring and returns its starting index; use find(substring, start_pos) to search from a specific position onward.
From Mixed Text to Useful Data
Real-world data often arrives as one mixed string. A single line can contain an email address, a timestamp, and other information together. Parsing means breaking that string down so that you can isolate the exact substring you need. A practical parsing workflow combines find() to locate important landmarks with slicing to extract the characters between those landmarks.
The central idea is to let find() discover positions first, then use those positions as boundaries for a slice.
Reading Positions with find()
The find() method locates a substring within a string and returns the starting index of the first matching occurrence it finds. For example, in the source parsing example, find('@') returns 21. That result means the @ character is located at position 21 in the larger string. The returned number is useful because it can be stored conceptually as a landmark and then used in a later operation.
Searching from a Known Landmark
find() can receive a starting position as its second argument. The form find(substring, start_pos) begins searching from the specified position onward. In the source example, the position of @ is stored as 21. Searching for a space from that position returns 31, identifying the first space after the @ landmark.
Two landmarks for one extraction
Use the source parsing sequence to identify the boundaries of a domain name.
Find the first landmark: find('@') returns 21, so the @ symbol is at position 21.
Continue from that landmark: find(' ', atpos) searches for a space beginning at position 21 and returns 31.
Interpret the positions: The @ at position 21 and the space at position 31 surround the domain portion that needs to be extracted.
The two find() results provide the boundaries needed for the next slicing step.
Turning Landmarks into Slices
A slice written as [start:end] extracts characters beginning at start and continuing up to, but not including, end. In the source example, the slice data[atpos+1:sppos] begins one position after the @ and ends just before the space. Adding 1 to atpos skips the @ itself. Using sppos as the ending boundary excludes the space. The extracted result is the domain name uct.ac.za.
The +1 is not arbitrary: it moves the slice start past the delimiter. The ending index remains the position of the space because the end of a slice is not included.
Tracing a Parsing Workflow
A complete parsing workflow can be traced as a sequence of dependent values. First, find() identifies the @ position. Next, that position becomes the starting point for a second find() search. Finally, both results flow into a slice. The first result is adjusted by 1 because the delimiter should not appear in the extracted text; the second result is used directly as the exclusive slice boundary.
Tracing the domain extraction
Follow the source example from the original string to the extracted domain.
Record atpos: The first search locates @ at position 21.
Record sppos: The second search starts at atpos and locates the next space at position 31.
Build the boundaries: The start boundary is atpos + 1, which is position 22. The end boundary is sppos, which is position 31.
Apply the slice: Characters from position 22 up to, but not including, position 31 form the desired substring.
The extracted substring is uct.ac.za.
Checking Search Results
Using a find() result without checking whether it is valid
The source guidance requires checking that find() returned a valid index rather than -1 before using it in a slice.
Fix:
Verify the result first, and trace the intermediate values if the parsing result is unexpected.Starting the slice at the delimiter
The delimiter becomes part of the extracted substring.
Fix:
Use the position after the delimiter when the delimiter itself should be skipped.Treating the slice end as included
String slicing extracts up to but does not include the end boundary.
Fix:
Use the delimiter's position as the end boundary when the delimiter should be excluded.Debugging only the final extracted text
The error may come from either a landmark search or a slice boundary.
Fix:
Inspect each find() result and then check how those indices flow into the slice.
Practice the Index Flow
Explain the parsing sequence in your own words: first identify the position returned by find('@'), then identify the position returned by searching for a space from that position, and finally explain why the extraction starts one position after @ and ends at the space position.
Hints
- Treat the two find() results as landmarks in the original string.
- Remember that the slice end is excluded.
- Ask which delimiter should be skipped at the beginning of the extracted substring.
What do you think happens?
The @ landmark is at position 21 and the following space is at position 31. What boundaries should be used to extract the domain without including either delimiter?
Reveal answer
Answer: Start at 22 and end at 31.
Adding 1 to the @ position skips the @. Using 31 as the end boundary stops immediately before the space because slice end positions are not included.
Key Takeaways
- find() locates a substring and returns its starting index.
- find(substring, start_pos) begins the search from a specified position onward.
- A slice [start:end] includes the start boundary and stops before the end boundary.
- Parsing combines landmark searches with slicing to extract a precise substring.
- When parsing fails, verify that find() returned a valid index and trace every intermediate value into the slice.
Key Takeaways
- Use find() to turn important characters or substrings into usable position landmarks.
- Give find() a starting position when the search should continue from a known point.
- Combine find() results with slicing to extract the characters between delimiters.
- Use one position after a starting delimiter and the delimiter's position as an exclusive ending boundary when appropriate.
- Debug parsing by checking every search result and tracing how each index becomes a slice boundary.