Regular Expression in Python

import re

text = "Contact us at [email protected] or [email protected]"
pattern = r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b'
emails = re.findall(pattern, text)
print(emails)  # ['[email protected]', '[email protected]']

That code extracts email addresses from text using Python regex. The pattern looks cryptic at first, but each piece serves a specific purpose in matching the structure of an email address. Regular expressions give you a tiny language for describing text patterns, and Python’s re module makes them practical to use.

Understanding Python regex syntax

Python regex patterns combine ordinary characters with special metacharacters that represent broader matching rules. The character \d matches any digit from 0 to 9. The symbol \w matches word characters (letters, digits, and underscores). The . matches any single character except newlines.

import re

# Match a three-digit number
pattern = r'\d{3}'
text = "My PIN is 4829"
match = re.search(pattern, text)
print(match.group())  # 482

The \b creates a word boundary, which helps you match complete words instead of partial strings. This becomes critical when you need to find “cat” but not “category.”

# Without word boundaries
pattern = r'cat'
text = "The cat in the category"
matches = re.findall(pattern, text)
print(matches)  # ['cat', 'cat']

# With word boundaries
pattern = r'\bcat\b'
matches = re.findall(pattern, text)
print(matches)  # ['cat']

Character classes let you define sets of acceptable characters using square brackets. The pattern [aeiou] matches any single vowel. You can specify ranges like [a-z] for lowercase letters or [0-9] for digits.

# Match valid hexadecimal digits
pattern = r'[0-9A-Fa-f]+'
text = "Color code: #FF5733"
match = re.search(pattern, text)
print(match.group())  # FF5733

Quantifiers control how many times a pattern element should repeat. The * matches zero or more occurrences. The + requires at least one occurrence. The ? makes the preceding element optional (zero or one occurrence).

# Match phone numbers with optional area codes
pattern = r'\(?\d{3}\)?[-.\s]?\d{3}[-.\s]?\d{4}'
numbers = [
    "(555) 123-4567",
    "555-123-4567",
    "5551234567"
]

for num in numbers:
    if re.match(pattern, num):
        print(f"{num} is valid")

Core Python regex functions

The re.search() function scans through a string looking for the first location where the pattern matches. It returns a match object when successful or None when no match exists.

import re

text = "Python 3.12 was released in 2024"
pattern = r'\d+\.\d+'
match = re.search(pattern, text)

if match:
    print(f"Found version: {match.group()}")  # Found version: 3.12
    print(f"Position: {match.start()}-{match.end()}")  # Position: 7-11

The re.findall() function returns all non-overlapping matches as a list of strings. This becomes your go-to when you need to extract multiple instances of a pattern from text.

# Extract all URLs from text
text = """
Visit https://example.com for docs.
Try http://test.org for testing.
"""

pattern = r'https?://[^\s]+'
urls = re.findall(pattern, text)
print(urls)  # ['https://example.com', 'http://test.org']

The re.sub() function substitutes matches with replacement text. You can use backreferences with \1, \2 to reference captured groups from the pattern.

# Format phone numbers consistently
text = "Call 555-1234 or 555.9876"
pattern = r'(\d{3})[-.](\d{4})'
formatted = re.sub(pattern, r'(\1) \2', text)
print(formatted)  # Call (555) 1234 or (555) 9876

The re.match() function checks if the pattern matches at the beginning of the string. This differs from re.search(), which looks anywhere in the string.

text = "Python is great"

# match() only checks the start
match = re.match(r'Python', text)
print(match.group() if match else "No match")  # Python

# This won't match because 'great' isn't at the start
match = re.match(r'great', text)
print(match)  # None

Compiling Python regex patterns

When you use the same pattern repeatedly, compiling it first improves performance. The re.compile() function creates a reusable pattern object with methods like search(), findall(), and sub().

import re

# Compile once, use many times
email_pattern = re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b')

texts = [
    "Email me at [email protected]",
    "Support: [email protected]",
    "No email here"
]

for text in texts:
    matches = email_pattern.findall(text)
    if matches:
        print(f"Found: {matches[0]}")

Compiled patterns also let you set flags that modify matching behavior. The re.IGNORECASE flag makes matching case-insensitive.

# Case-insensitive matching
pattern = re.compile(r'python', re.IGNORECASE)
texts = ["Python", "PYTHON", "python"]

for text in texts:
    if pattern.search(text):
        print(f"{text} matches")  # All three match

Working with groups in Python regex

Parentheses create capturing groups that extract specific parts of a match. Each group gets assigned a number starting from 1, and you can access them through the match object.

import re

# Parse log entries
log = "2024-01-15 14:30:22 ERROR Database connection failed"
pattern = r'(\d{4}-\d{2}-\d{2}) (\d{2}:\d{2}:\d{2}) (\w+) (.+)'

match = re.search(pattern, log)
if match:
    date = match.group(1)
    time = match.group(2)
    level = match.group(3)
    message = match.group(4)
    
    print(f"Date: {date}")
    print(f"Level: {level}")
    print(f"Message: {message}")

Named groups make your code more readable by letting you access groups by name instead of number. Use the syntax (?P<name>pattern) to create a named group.

# Extract structured data
pattern = r'(?P<protocol>https?)://(?P<domain>[^/]+)(?P<path>/.*)?'
url = "https://api.example.com/v1/users"

match = re.search(pattern, url)
if match:
    print(f"Protocol: {match.group('protocol')}")  # https
    print(f"Domain: {match.group('domain')}")      # api.example.com
    print(f"Path: {match.group('path')}")          # /v1/users

The re.finditer() function returns an iterator of match objects, which gives you both the matched text and position information for each match.

# Find all numbers with their positions
text = "Scores: 85, 92, 78, 95"
pattern = r'\d+'

for match in re.finditer(pattern, text):
    print(f"Found {match.group()} at position {match.start()}-{match.end()}")
# Found 85 at position 8-10
# Found 92 at position 12-14
# Found 78 at position 16-18
# Found 95 at position 20-22

Handling special characters in Python regex