Concepts / Data Validation and Cleaning

Data Validation and Cleaning

The find() method locates the position of a substring and returns its starting index; use find(substring, start_pos) to search from a specific position onward.

  • Programming

From Mixed Text to Useful Data

Real-world data often arrives in mixed formats. A single line can contain an email address, a timestamp, and other information together. Parsing means breaking that larger string down to isolate the exact substring you need. In Python, a practical way to do this is to use find() to locate landmarks and string slicing to extract the characters between those landmarks.

The central workflow is: locate a starting landmark, locate an ending landmark, then slice between their positions.

locatesupply indicesextractMixed stringemail and other informationLandmark positionsfind() resultsSlice rangestart:endExtracted substringdata to validate or clean
How does a larger unstructured string become a smaller extracted value that can be checked or cleaned?

Reading Positions with find()

The find() method locates a substring inside a string and returns the substring's starting index. An index identifies a character position in the larger string. For example, when find('@') returns 21, the @ character begins at position 21. The returned number is a position, not the substring itself.

containsfind('@') returnsemail textlarger string@index 2121starting index
Which character position does find() return when it locates a substring inside a larger string?

The two-argument form find(substring, start_pos) begins searching at the specified position and continues onward. This is useful when the same kind of character or substring may appear more than once and you need the occurrence after a known landmark.

begin search atfind(' ', 21)positionmixed stringcontains multiple spaces21start_posspacefirst space at or after 2131returned index
How does changing the starting position alter which occurrence find() returns?

Turning Landmarks into a Slice

Extracting an Email Domain

Use the position of @ and the next space to isolate the domain name from a larger string.

Find the first landmark: find('@') returns 21, so the @ symbol begins at index 21.

Find the ending landmark: find(' ', atpos) searches for a space beginning at position 21 and returns 31. This identifies the first space after the @ symbol.

Set the slice start: Use atpos+1 as the start. This changes the starting position from the @ symbol at 21 to the first character after it at 22.

Set the slice end: Use sppos as the end. A slice stops before its end index, so the space at index 31 is excluded.

The slice data[atpos+1:sppos] extracts uct.ac.za.

skip @exclusive endslice startatpos21atpos+122sppos31uct.ac.zacharacters from 22 up to 31
How does the index returned by find() determine the start and end positions used by a slice?
containsskip with +1stop beforelarger stringcontains @, domain, andspace@index 21uct.ac.zaindices 22 through 30spaceindex 31
What portion of the original string is included or excluded when slicing between two index positions?

String slicing with [start:end] extracts characters from start up to but not including end. In this example, the start is atpos+1, which skips the @ symbol, and the end is sppos, which excludes the space. The indices therefore describe exactly which characters become the extracted value.

Tracing a Parsing Failure

When parsing does not produce the expected result, trace the indices step by step. First inspect the intermediate values returned by find(). Confirm that each landmark is at the position you expect. Then inspect the slice start and end positions. A mistake in either landmark can cause the slice to include the wrong characters or omit needed characters.

checkyesnofind() resultlandmark indexvalid indexnot -1slice operationuse start:enddebug outputinspect intermediate values
What should happen before a position returned by find() is used in later parsing logic?
returnscheck before slicingsubstring searchfind()-1target absentmissing landmarkdo not assume a validposition
What value does find() produce when the target is absent, and how should parsing logic respond?

Mistakes with Indices

  • Treating the value from find() as the extracted text

    The returned value is the starting index of the @ character, not the domain or another substring.

    Fix: Use the returned index as a landmark for a later slice.

  • Including the delimiter in the extracted value

    The desired domain begins one position after the @ character.

    Fix: Use atpos+1 when the delimiter itself should be excluded.

  • Assuming the slice includes its end position

    A slice extracts up to but not including its end index.

    Fix: Remember that [start:end] stops before end.

  • Using a -1 result as though it were a valid landmark

    The expected substring was not found, so the planned boundary is unavailable.

    Fix: Check the result before using it in a slice and inspect intermediate values while debugging.

  • Searching from the beginning when a later occurrence is needed

    The search may locate a space before the relevant landmark.

    Fix: Pass the known landmark as the starting position, as in find(' ', atpos).

Guided Practice

MEDIUM

A larger string contains an email address followed by a space and additional information. Describe the sequence of positions you would identify to extract the domain: first locate the @ symbol, then search for the next space beginning at the @ position, and finally slice from one position after @ up to the space. What should you check before performing the slice?

Hints
  • The first find() result identifies the position of @.
  • Pass that position as the starting position for the search for the next space.
  • The slice excludes its end index.
  • Check that each find() result is not -1 before using it as a boundary.

What do you think happens?

If find('@') returns 21 and the next space found from position 21 is at 31, which positions should define the slice that extracts the domain without the @ symbol or the space?

  • 21 through 31
  • 22 through 31 as an exclusive end
  • 21 through 30
  • 22 through 32
Reveal answer

Answer: 22 through 31 as an exclusive end

The @ symbol is at position 21, so the slice starts at 21+1, or 22. The end position is 31, and slicing stops before that position, excluding the space.

Parsing Workflow

  1. Identify a meaningful landmark in the larger string with find().
  2. Store or inspect the starting index returned by find().
  3. Use find(substring, start_pos) when the search must begin after a known position.
  4. Convert the landmark positions into slice boundaries.
  5. Remember that [start:end] includes start and excludes end.
  6. Check every find() result for -1 before using it in a slice.
  7. Print intermediate index values when debugging unexpected parsing results.

Parsing becomes manageable when you treat each index as a boundary in the original string. find() tells you where a landmark begins; a starting position narrows a later search; slicing uses the resulting boundaries to isolate the needed characters. If the result is wrong, inspect those boundaries in order rather than guessing at the final substring.

Key Takeaways

  • find() returns the starting index of a located substring.
  • find(substring, start_pos) searches from a specified position onward.
  • A slice [start:end] includes start but excludes end.
  • Use find() landmarks as slice boundaries to extract data between delimiters.
  • Check for -1 and inspect intermediate indices before trusting parsing results.