Pandas apply / map / applymap

apply, map, and applymap are the three major functions in Pandas for data transformation. They can perform flexible element-wise or batch operations on DataFrame or Series.


Series.map

mapIs a Series method, used to transform each element in a Series.

Basic Usage

Example

import pandas as pd

# Create Series
s = pd.Series([1, 2, 3, 4, 5])

print("Original data:")
print(s)
print()

# Use a function
print("Each element * 2:")
print(s.map(lambda x: x * 2))
print()

# Use dictionary mapping
mapping = {1: "A", 2: "B", 3: "C", 4: "D", 5: "E"}
print("Using dictionary mapping:")
print(s.map(mapping))
print()

# Use Series mapping
mapping_series = pd.Series(["A", "B", "C", "D", "E"], index=[1, 2, 3, 4, 5])
print("Using Series mapping:")
print(s.map(mapping_series))

Handling Missing Values

Example

import pandas as pd
import numpy as np

s = pd.Series([1, 2, np.nan, 4, 5])

print("Contains NaN:")
print(s)
print()

# map skips NaN by default
print("map handling (skips NaN):")
print(s.map(lambda x: x * 2 if pd.notna(x) else -1))

DataFrame.applymap

applymapIs a DataFrame method that applies a function to each element one by one (Note: Pandas 2.0+ recommends usingDataFrame.mapinstead).

Example

import pandas as pd
import numpy as np

df = pd.DataFrame({
    "A": [1, 2, 3],
    "B": [4, 5, 6],
    "C": [7, 8, 9]
})

print("Original data:")
print(df)
print()

# Multiply each element by 2
print("Each element * 2:")
print(df.applymap(lambda x: x * 2))
print()

# Keep 2 decimal places
print("Keep 2 decimal places:")
print(df.applymap(lambda x: round(x, 2)))

applymap performs element-wise operations and can be slow for large data. If you only need to operate on numeric columns, consider using vectorized operations or apply with the axis parameter.


DataFrame.apply

applyIs the most flexible method and can apply functions along an axis.

Apply by Column

Example

import pandas as pd
import numpy as np

df = pd.DataFrame({
    "A": [1, 2, 3, 4, 5],
    "B": [10, 20, 30, 40, 50],
    "C": [100, 200, 300, 400, 500]
})

print("Original data:")
print(df)
print()

# Default axis=0, apply by column
print("Column sum:")
print(df.apply(sum))
print()

print("Column max:")
print(df.apply(max))

Apply by Row

Example

import pandas as pd
import numpy as np

df = pd.DataFrame({
    "A": [1, 2, 3],
    "B": [10, 20, 30],
    "C": [100, 200, 300]
})

print("Original data:")
print(df)
print()

# axis=1, apply by row
print("Row sum:")
print(df.apply(sum, axis=1))
print()

# Max - min per row
print("Row range:")
print(df.apply(lambda x: x.max() - x.min(), axis=1))

Using aggfunc for Aggregation

Example

import pandas as pd
import numpy as np

df = pd.DataFrame({
    "A": [1, 2, 3],
    "B": [10, 20, 30]
})

# Apply multiple functions at once
print("Sum and mean simultaneously:")
print(df.apply([sum, np.mean]))
print()

# Return multiple values
result = df.apply(lambda x: pd.Series({
    "sum": x.sum(),
    "mean": x.mean(),
    "max": x.max()
}, index=["sum", "mean", "max"]))

print("Returning multiple values:")
print(result)

Series.apply

Series can also use apply, which is similar to map but more flexible.

Example

import pandas as pd
import numpy as np

s = pd.Series([1, 4, 9, 16, 25])

print("Original data:")
print(s)
print()

# Square root
print("Square root:")
print(s.apply(np.sqrt))
print()

# Conditional return
print("Conditional judgment:")
print(s.apply(lambda x: "large" if x > 10 else "small"))

Performance Comparison

Example

import pandas as pd
import numpy as np
import time

# Create large data
n = 100000
s = pd.Series(np.random.randn(n))

# Test map vs apply
func = lambda x: x * 2 + 1

start = time.time()
result1 = s.map(func)
map_time = time.time() - start

start = time.time()
result2 = s.apply(func)
apply_time = time.time() - start

# Vectorized (fastest)
start = time.time()
result3 = s * 2 + 1
vec_time = time.time() - start

print(f"map time: {map_time:.4f}s")
print(f"apply time: {apply_time:.4f}s")
print(f"vectorized time: {vec_time:.4f}s")

print("\nConclusion: prioritize vectorized operations for best performance")

Practical: Data Transformation

Example

import pandas as pd
import numpy as np

# Create a sample DataFrame
df = pd.DataFrame({
    "Name": ["Zhang San", "Li Si", "Wang Wu", "Zhao Liu"],
    "Age": [25, 30, 28, 35],
    "Salary": [12000, 15000, 11000, 18000],
    "Department": ["Technology", "Sales", "Technology", "Operations"]
})

print("Original data:")
print(df)
print()

# Use apply for row-level calculation
def calculate(row):
    """Calculate annual income and after-tax salary"""
    annual = row["Salary"] * 12
    tax = annual * 0.1 if annual > 120000 else annual * 0.05
    after_tax = annual - tax
    return pd.Series({
        "Annual salary": annual,
        "Tax": tax,
        "After tax": after_tax
    })

result = df.apply(calculate, axis=1)
df_result = pd.concat([df, result], axis=1)

print("Calculation results:")
print(df_result)

Choosing Among the Three

Method Applicable object Scenario Performance
map Series Element-wise transformation, dictionary mapping Fast
applymap DataFrame Element-wise transformation (non-numeric columns) Slow
apply Series/DataFrame Row/column aggregation, custom functions Medium

When vectorized operations (directly using operators) are available, don't use apply/map; when map is available, don't use apply.

Other Extensions