Concepts / Using re.sub() to Replace Matches

Using re.sub() to Replace Matches

re.findall() is a Python method that extracts all non-overlapping substrings matching a regex pattern from a string.

  • Programming

A Different Kind of Match Operation

The supplied topic title mentions re.sub(), but the concept covered here is re.findall(). The key difference is the kind of result you are trying to obtain: findall() extracts matching text from a string, while this lesson concentrates on how those extracted matches are located and returned.

re.findall() is a Python method that extracts all non-overlapping substrings matching a regular expression pattern from a string.

The two questions to ask first are: Which parts of the input match the pattern, and how many capturing groups does the pattern contain?

How findall() Moves Through Text

findall() scans from left to right. It searches for the first location where the pattern can match. After it finds a match, it records the result and resumes searching at the position immediately after that match ends. Because the next search begins after the previous match, the results are non-overlapping.

scan left to rightresume after match endscontinue scanninginput stringtext containing severalpossible matchesfirst matchrecordedposition after matchnext search begins herenext matchrecorded
Which parts of the input does findall() match, and where does it resume searching after each match?
python
Output
['42', '17', '93']

In this example, the pattern identifies each digit sequence. findall() records the first sequence, continues after it, then records the next sequence. The result is a list containing three strings rather than one combined match.

Capturing Groups Change the Result

A capturing group is a subpattern enclosed in parentheses. The number of capturing groups controls the structure of the value returned by findall(). This is independent of how many matches are found.

returnsreturnsreturnsZero groupslist of complete matchesstrings['42', '17']One grouplist of captured contentsstrings['42', '17']Multiple groupslist of captured group setstuples[('A', '42'), ('B', '17')]
How does the return value change when the regex has zero, one, or multiple capturing groups?
Capturing groupsReturned structureWhat each result contains
NoneList of stringsEach complete match
Exactly oneList of stringsThe content of that one group
Two or moreList of tuplesThe captured groups in order

The number of capturing groups determines the shape and contents of the findall() result.

import re text = "A42 B17" complete = re.findall(r"[A-Z]\d+", text) letter = re.findall(r"([A-Z])\d+", text) pair = re.findall(r"([A-Z])(\d+)", text) print(complete) print(letter) print(pair)

Several Matches from One Line

A single input line can contain several occurrences of the same pattern. findall() does not stop after the first successful match. It continues scanning from the end of each match and collects every later non-overlapping match that fits the pattern.

findresume after matchresume after matchreturn collected valuesstart scanfirst matchrecord resultsecond matchappend resultthird matchappend resultresult listall non-overlapping matches
How does one regex pattern produce a list of several matches from a single input line?

Extracting addresses from one log line

Extract each email address from the line "Contact a@example.com or b@example.com".

Choose a pattern: Use a pattern that matches the email-address form used in this example.

Scan left to right: findall() locates the first address, records it, and resumes immediately after that address.

Continue scanning: The method finds the second address and records it as another result.

Check the structure: Because the example pattern has no capturing groups, the result is a list of strings containing complete matches.

The result contains two extracted email-address strings.

Diagnosing Unexpected Results

When findall() produces a result that differs from your expectation, do not inspect only the returned list. Inspect the pattern and the match process that produced it. Three especially important checks are the pattern itself, the number of capturing groups, and the boundaries created by quantifiers.

investigatecheck structurecheck boundariesexplain outputunexpected resultwrong shape or boundariesinspect patternmatch text and groupscount groupszero, one, or multiplecheck quantifiersgreedy or lazyexplained resultmatches now make sense
Which part of the regex, including its capturing groups or match boundaries, explains the difference between the expected and actual result?
  • Expecting complete matches when the pattern contains one capturing group.

    With exactly one capturing group, findall() returns the contents of that group.

    Fix: Count the parentheses that create capturing groups and decide whether you want the full match or a captured subpattern.

  • Expecting a flat list when the pattern contains multiple capturing groups.

    Two or more capturing groups make findall() return a list of tuples.

    Fix: Read each tuple as the captured groups in their pattern order.

  • Assuming the pattern matches different boundaries than it actually does.

    Quantifiers affect how much text a match consumes. By default, quantifiers are greedy.

    Fix: Test the pattern in isolation and inspect whether its quantifiers should be greedy or lazy.

  • Assuming findall() overlaps matches.

    findall() resumes after the previous match ends.

    Fix: Trace the end of each match and begin the next search from that position.

Debug in a fixed order: first test whether the pattern matches the intended text, then count the capturing groups, then inspect quantifiers and their match boundaries. This separates a pattern problem from a return-structure problem.

Practice the Prediction

What do you think happens?

What does the result contain, and what is its structure? Assume the input is "A42 B17" and the pattern is ([A-Z])(\d+).

  • A list of complete strings
  • A list of one captured string per match
  • A list of tuples containing two captured strings per match
Reveal answer

Answer: A list of tuples containing two captured strings per match: [('A', '42'), ('B', '17')].

The pattern has two capturing groups. findall() returns a tuple for each non-overlapping match, with the groups in order.

MEDIUM

Write three findall() patterns for the text "A42 B17": one with no capturing groups, one with exactly one capturing group, and one with two capturing groups. Predict the result structure before checking your code.

Hints
  • Use a pattern that identifies an uppercase letter followed by digits.
  • For the one-group version, place parentheses around only the part you want returned.
  • For the two-group version, place separate capturing groups around the letter and the digits.

Key Takeaways

  1. findall() scans from left to right and extracts all non-overlapping matches.
  2. After each match, scanning resumes at the position where that match ends.
  3. With no capturing groups, findall() returns complete matches as strings.
  4. With one capturing group, findall() returns the contents of that group as strings.
  5. With multiple capturing groups, findall() returns tuples containing the captured groups in order.
  6. Unexpected results usually require checking the pattern, capturing-group count, quantifiers, and greedy or lazy behavior.

Key Takeaways

  • re.findall() extracts all non-overlapping substrings that match a regular expression.
  • The method scans left to right and resumes after each completed match.
  • The number of capturing groups determines whether the result contains complete strings, captured strings, or tuples.
  • Debug unexpected output by checking the pattern, capturing groups, quantifiers, and match boundaries.