Pandas String Operations

Pandas provides powerful string processing capabilities. Through the.straccessor, you can process each element in a Series just like operating on Python strings.


str Accessor Overview

When the Series data type isobjectyou can use the.straccessor to perform string operations.

Example

import pandas as pd

# Create a Series containing strings
s = pd.Series(["hello", "world", "pandas", "Python"])

print("Original Series:")
print(s)
print()

# Convert to lowercase
print("Lowercase:")
print(s.str.lower())
print()

# Convert to uppercase
print("Uppercase:")
print(s.str.upper())
print()

# Capitalize first letter
print("Capitalized:")
print(s.str.capitalize())

Common String Functions

Case Conversion

Method Description Example
str.lower() Convert to lowercase "Hello" → "hello"
str.upper() Convert to uppercase "Hello" → "HELLO"
str.title() Capitalize first letter "hello world" → "Hello World"
str.capitalize() Capitalize first letter, lowercase the rest "hELLO" → "Hello"
str.swapcase() Swap case "HeLLo" → "hEllO"

String Search and Replace

Example

import pandas as pd

s = pd.Series(["hello world", "python pandas", "data science", "machine learning"])

print("Original data:")
print(s)
print()

# Check containment
print("Contains 'python':")
print(s.str.contains("python", case=False))
print()

# Check start/end
print("Starts with 'hello':")
print(s.str.startswith("hello"))
print()

# Replace
print("Replace spaces with underscores:")
print(s.str.replace(" ", "_"))
print()

# Split
print("Split by space:")
print(s.str.split(" "))

Whitespace Removal and Padding

Example

import pandas as pd

s = pd.Series(["  hello  ", "world  ", "  pandas", " python "])

print("Original data:")
print(s)
print()

# Strip leading and trailing whitespace
print("Strip leading/trailing whitespace:")
print(s.str.strip())
print()

# Strip left whitespace
print("Strip left whitespace:")
print(s.str.lstrip())
print()

# Strip right whitespace
print("Strip right whitespace:")
print(s.str.rstrip())
print()

# Pad
print("Pad with 0 on the left to length 10:")
print(s.str.pad(10, side="left", fillchar="0"))

Slicing and Concatenation

Example

import pandas as pd

s = pd.Series(["hello", "world", "python", "pandas"])

print("Original data:")
print(s)
print()

# Slice
print("First 3 characters:")
print(s.str[:3])
print()

# Take 3 characters starting from position 1
print("Take 3 characters from position 1:")
print(s.str[1:4])
print()

# String concatenation
s1 = pd.Series(["Hello", "World"])
s2 = pd.Series(["Python", "Pandas"])
print("Concatenate with +:")
print(s1 + " " + s2)
print()

# Join with join() (using a specified separator)
print("Join with join():")
print(s1.str.cat(s2, sep="-"))

Regular Expression Support

Pandas string functions support regular expressions, making them a powerful tool for handling complex string patterns.

Example

import pandas as pd

s = pd.Series([
    "user@example.com",
    "test@domain.org",
    "invalid-email",
    "admin@site.net"
])

print("Original data:")
print(s)
print()

# Extract email
print("Extract username:")
print(s.str.extract(r"(\w+)@"))
print()

print("Extract domain:")
print(s.str.extract(r"@(\w+\.\w+)"))
print()

# Replace
print("Replace email with [email]:")
print(s.str.replace(r"\w+@\w+\.\w+", "[email]", regex=True))
print()

# Match containment
print("Contains .com or .org:")
print(s.str.contains(r"\.com|\.org"))

Extraction and Splitting

Example

import pandas as pd

s = pd.Series(["2024-01-01", "2024-02-15", "2024-03-20"])

print("Original data:")
print(s)
print()

# Extract year, month, day
extract_df = s.str.extract(r"(\d{4})-(\d{2})-(\d{2})")
extract_df.columns = ["Year", "Month", "Day"]
print("Extract year/month/day:")
print(extract_df)
print()

# Split by delimiter
s2 = pd.Series(["a,b,c", "d,e,f", "g,h,i"])
print("Split by comma:")
print(s2.str.split(","))
print()

# Split and expand into a DataFrame
print("Split and expand:")
print(s2.str.split(",", expand=True))

Statistics and Conversion

Length and Counting

Example

import pandas as pd

s = pd.Series(["hello", "world", "python", "pandas"])

print("Original data:")
print(s)
print()

# String length
print("String length:")
print(s.str.len())
print()

# Character count
s2 = pd.Series(["hello", "world", "aaa", "bbb"])
print("Occurrences of character 'l':")
print(s2.str.count("l"))

Type Conversion

Example

import pandas as pd

# String to numeric
s = pd.Series(["1", "2", "3", "abc"])
print("String to numeric:")
print(pd.to_numeric(s, errors="coerce"))
print()

# Numeric to string
s2 = pd.Series([1, 2, 3, 4])
print("Numeric to string:")
print(s2.astype(str).str.zfill(4))  # Zero-pad to 4 digits

Practical: Data Cleaning

Example

import pandas as pd
import numpy as np

# Simulate dirty data
df = pd.DataFrame({
    "Name": [" Zhang San ", "Li Si", "Wang Wu "],
    "Phone": ["138-0000-0000", "13900000000", "  010-12345678  "],
    "Email": ["zhangsan@email.com", "LISI@EMAIL.COM", "wang.wu@domain.com "]
})

print("Original data:")
print(df)
print()

# Cleaning steps
# 1. Remove whitespace
df["Name"] = df["Name"].str.strip()
df["Phone"] = df["Phone"].str.replace("-", "").str.strip()
df["Email"] = df["Email"].str.strip().str.lower()

print("After cleaning:")
print(df)
print()

# 2. Validate data
print("Email validation:")
print(df["Email"].str.contains(r"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$"))

Important Notes

1. Handling missing values

String operations skip missing values (NaN) by default. If you need to handle them, you can use.fillna()to fill them first.

2. Case sensitivity

Most functions are case-sensitive by default. Use thecase=Falseparameter for case-insensitive matching.

3. Regular expression performance

Regular expressions are powerful but relatively slow. Watch performance when processing large amounts of data.

String processing is an important part of data cleaning. Using the.straccessor properly can efficiently accomplish most text processing tasks.

Other Extensions