Concepts / Working with Multiple Strings

Working with Multiple Strings

The find() method locates the position of a substring and returns its starting index; use find(substring, start_pos) to search from a specific position onward.

  • Programming

From Mixed Text to Useful Data

Real-world data often arrives as one mixed string. A single line can contain an email address, a timestamp, and other information together. Parsing means breaking that string down so that you can isolate the exact substring you need. A practical parsing workflow combines find() to locate important landmarks with slicing to extract the characters between those landmarks.

The central idea is to let find() discover positions first, then use those positions as boundaries for a slice.

Reading Positions with find()

The find() method locates a substring within a string and returns the starting index of the first matching occurrence it finds. For example, in the source parsing example, find('@') returns 21. That result means the @ character is located at position 21 in the larger string. The returned number is useful because it can be stored conceptually as a landmark and then used in a later operation.

containsfind('@') returnslarger stringemail and surrounding text@position 2121starting index
Which character position does find() return, and how does that index identify a substring's location in the original string?

Searching from a Known Landmark

find() can receive a starting position as its second argument. The form find(substring, start_pos) begins searching from the specified position onward. In the source example, the position of @ is stored as 21. Searching for a space from that position returns 31, identifying the first space after the @ landmark.

find('@')find(' ', 21)larger stringcontains @ and spaces21search begins here31first space found afterward
How does providing a start position change where find() begins searching and which occurrence it returns?

Two landmarks for one extraction

Use the source parsing sequence to identify the boundaries of a domain name.

Find the first landmark: find('@') returns 21, so the @ symbol is at position 21.

Continue from that landmark: find(' ', atpos) searches for a space beginning at position 21 and returns 31.

Interpret the positions: The @ at position 21 and the space at position 31 surround the domain portion that needs to be extracted.

The two find() results provide the boundaries needed for the next slicing step.

Turning Landmarks into Slices

A slice written as [start:end] extracts characters beginning at start and continuing up to, but not including, end. In the source example, the slice data[atpos+1:sppos] begins one position after the @ and ends just before the space. Adding 1 to atpos skips the @ itself. Using sppos as the ending boundary excludes the space. The extracted result is the domain name uct.ac.za.

skip @slice startslice end excludedatpos21atpos + 122uct.ac.zacharacters betweenboundariessppos31
How does an index returned by find() become the start or end boundary of a slice?

The +1 is not arbitrary: it moves the slice start past the delimiter. The ending index remains the position of the space because the end of a slice is not included.

Tracing a Parsing Workflow

A complete parsing workflow can be traced as a sequence of dependent values. First, find() identifies the @ position. Next, that position becomes the starting point for a second find() search. Finally, both results flow into a slice. The first result is adjusted by 1 because the delimiter should not appear in the extracted text; the second result is used directly as the exclusive slice boundary.

locate @start search at 21supply end 31supply start 21 + 1extractsource textmixed stringfind('@')21find(' ', 21)31data[22:31]start after @uct.ac.zaextracted domain
What happens to each index as parsing moves from find() results into successive slice operations?

Tracing the domain extraction

Follow the source example from the original string to the extracted domain.

Record atpos: The first search locates @ at position 21.

Record sppos: The second search starts at atpos and locates the next space at position 31.

Build the boundaries: The start boundary is atpos + 1, which is position 22. The end boundary is sppos, which is position 31.

Apply the slice: Characters from position 22 up to, but not including, position 31 form the desired substring.

The extracted substring is uct.ac.za.

Checking Search Results

check resultyesno or unexpectedfind() resultcandidate indexvalid indexnot -1slice datause boundariesintermediate valuesinspect parsing
How should parsing logic be checked before a find() result is used in slicing?
  • Using a find() result without checking whether it is valid

    The source guidance requires checking that find() returned a valid index rather than -1 before using it in a slice.

    Fix: Verify the result first, and trace the intermediate values if the parsing result is unexpected.

  • Starting the slice at the delimiter

    The delimiter becomes part of the extracted substring.

    Fix: Use the position after the delimiter when the delimiter itself should be skipped.

  • Treating the slice end as included

    String slicing extracts up to but does not include the end boundary.

    Fix: Use the delimiter's position as the end boundary when the delimiter should be excluded.

  • Debugging only the final extracted text

    The error may come from either a landmark search or a slice boundary.

    Fix: Inspect each find() result and then check how those indices flow into the slice.

Practice the Index Flow

MEDIUM

Explain the parsing sequence in your own words: first identify the position returned by find('@'), then identify the position returned by searching for a space from that position, and finally explain why the extraction starts one position after @ and ends at the space position.

Hints
  • Treat the two find() results as landmarks in the original string.
  • Remember that the slice end is excluded.
  • Ask which delimiter should be skipped at the beginning of the extracted substring.

What do you think happens?

The @ landmark is at position 21 and the following space is at position 31. What boundaries should be used to extract the domain without including either delimiter?

  • Start at 21 and end at 31
  • Start at 22 and end at 31
  • Start at 21 and end at 32
  • Start at 22 and end at 32
Reveal answer

Answer: Start at 22 and end at 31.

Adding 1 to the @ position skips the @. Using 31 as the end boundary stops immediately before the space because slice end positions are not included.

Key Takeaways

  1. find() locates a substring and returns its starting index.
  2. find(substring, start_pos) begins the search from a specified position onward.
  3. A slice [start:end] includes the start boundary and stops before the end boundary.
  4. Parsing combines landmark searches with slicing to extract a precise substring.
  5. When parsing fails, verify that find() returned a valid index and trace every intermediate value into the slice.

Key Takeaways

  • Use find() to turn important characters or substrings into usable position landmarks.
  • Give find() a starting position when the search should continue from a known point.
  • Combine find() results with slicing to extract the characters between delimiters.
  • Use one position after a starting delimiter and the delimiter's position as an exclusive ending boundary when appropriate.
  • Debug parsing by checking every search result and tracing how each index becomes a slice boundary.