Concepts / Anchors and Boundaries in Regular Expressions

Anchors and Boundaries in Regular Expressions

Square brackets [...] define a set of acceptable characters for a single position in a regex pattern.

  • Programming

From Dirty Text to Clean Matches

Suppose text contains the string source@collab.sakaiproject.org>;. The email address is present, but extra characters cling to its boundary. A pattern that only looks for the @ symbol can capture those unwanted characters. The central idea in this article is to make the beginning and ending of a match more specific: require an alphanumeric character at the start and a letter at the end.

One Position, Several Acceptable Characters

Square brackets define a character set. A character set describes the alternatives allowed at one position in a regular expression, so it matches exactly one character from the set. For example, [A-Z] accepts one uppercase letter. A range such as [a-z] represents every character from a through z. Combining ranges gives [a-zA-Z0-9], which accepts one lowercase letter, uppercase letter, or digit.

matches[A-Z]one uppercase letterGone accepted character
How does a pattern like [A-Z] match exactly one character from a range of acceptable characters?

A character set does not automatically match a whole word. Without a quantifier, [a-zA-Z0-9] matches one character only.

What the Quantifier Repeats

The quantifiers * and + apply only to the character immediately to their left. The quantifier * means zero or more occurrences, while + means one or more occurrences. In \S*, the asterisk applies to \S, so the pattern can consume zero or more non-whitespace characters. In a pattern containing a character set followed by a quantifier, the quantifier applies to that character set as the immediately preceding unit.

changes minimum from zero to one\S*zero or more non-whitespacecharacters\S+one or more non-whitespacecharacters
What changes when * or + is placed after a character or character set, and which part of the pattern does the quantifier repeat?

Why the Email Pattern Uses an Asterisk

Choose between [a-zA-Z0-9]\S*@ and [a-zA-Z0-9]\S+@ when the pattern must allow the email address a@b.c.

Count the initial character: [a-zA-Z0-9] already matches the initial a.

Check the gap before @: There are no additional characters between the initial a and the @ symbol, so the middle part must be allowed to match zero characters.

Choose the quantifier: \S* allows zero or more non-whitespace characters. \S+ would require at least one additional non-whitespace character after the initial character.

[a-zA-Z0-9]\S*@ allows the stated short email form, while [a-zA-Z0-9]\S+@ would exclude it.

Filtering the Email Boundaries

The source pattern for extracting an email address is [a-zA-Z0-9]\S*@\S*[a-zA-Z]. It has five functional parts. The first character set requires the match to begin with a letter or digit. The first \S* allows zero or more non-whitespace characters before the @ symbol. The @ matches itself literally. The second \S* allows zero or more non-whitespace characters after the @ symbol. The final [a-zA-Z] requires the match to end with a letter.

thenthenthenthen[a-zA-Z0-9]one starting letter ordigit\S*zero or more non-whitespacecharacters@literal at sign\S*zero or more non-whitespacecharacters[a-zA-Z]one ending letter
How do character sets, quantifiers, and boundary characters work together as data flows through an email-matching pattern?

When the engine reaches the > character after org, that character cannot satisfy the final [a-zA-Z]. The engine therefore backtracks to the last letter it found, the g in org, and the match stops there. The ending character set acts as a precise stopping condition rather than allowing every non-whitespace character to remain in the result.

Reading the Extraction Result

The re.findall() function returns a Python list. Each match appears as a string element in that list. If a pattern finds multiple matches in one line, the list contains multiple strings in the order returned by the search.

For the boundary-filtering pattern, a match from source@collab.sakaiproject.org>; is represented as the string source@collab.sakaiproject.org rather than including the trailing > or ;. Multiple matches are represented as separate string elements in the returned list.

Mistakes in Character-Set Patterns

  • Treating a character set as if it matches a whole word

    A character set defines acceptable characters for one position and matches exactly one character unless a quantifier changes that behavior.

    Fix: Use a quantifier when the pattern needs to consume additional characters, and remember that the quantifier applies only to the item immediately before it.

  • Using + after the initial character set in the email pattern

    The initial character set already accounts for one non-whitespace character. The + then requires at least one more non-whitespace character before @, excluding a@b.c.

    Fix: Use \S* when zero additional characters must be allowed after the initial alphanumeric character.

  • Allowing \S at both boundaries without restrictions

    \S matches non-whitespace characters, including boundary symbols such as <, >, or ;.

    Fix: Use [a-zA-Z0-9] at the start and [a-zA-Z] at the end when those are the required boundary characters.

  • Assuming the @ symbol alone produces clean email data

    Extra punctuation can cling to the email address in the source text.

    Fix: Make the acceptable starting and ending characters explicit.

Build and Check a Precise Pattern

MEDIUM

Construct the five-part pattern for an email address that starts with a letter or digit, allows zero or more non-whitespace characters before and after @, and ends with a letter. Then explain why the pattern stops before a trailing > or ;.

Hints
  • Begin with [a-zA-Z0-9].
  • Place \S* before the literal @ and another \S* after it.
  • Finish with [a-zA-Z].
  • Check the final character of source@collab.sakaiproject.org>; against the ending character set.

What do you think happens?

What value should the pattern [a-zA-Z0-9]\S*@\S*[a-zA-Z] extract from source@collab.sakaiproject.org>;

  • source@collab.sakaiproject.org>;
  • source@collab.sakaiproject.org
  • source@collab
  • No match
Reveal answer

Answer: source@collab.sakaiproject.org

The final [a-zA-Z] requires the match to end with a letter. The > and ; characters cannot satisfy it, so the match stops at the g in org.

Key Takeaways

  1. Square brackets define a set of acceptable characters for one position in a regular expression.
  2. Ranges such as [a-z] represent all characters in the stated range, and [a-zA-Z0-9] accepts one letter or digit.
  3. The quantifiers * and + apply only to the character or pattern unit immediately to their left; * allows zero or more, while + allows one or more.
  4. In [a-zA-Z0-9]\S*@\S*[a-zA-Z], the boundary character sets restrict the beginning and ending characters while \S* provides flexibility in the middle.
  5. re.findall() returns matches as strings inside a Python list.

Key Takeaways

  • Character sets describe the acceptable character at a single position.
  • Quantifiers change how many times the immediately preceding character or set can match.
  • Boundary character sets prevent unwanted punctuation from entering extracted email addresses.
  • The pattern [a-zA-Z0-9]\S*@\S*[a-zA-Z] uses flexible middle sections and restrictive edges.
  • re.findall() packages each match as a string in a Python list.