Pandas Data Selection (loc / iloc / at)
Data selection is one of the most frequently used operations in Pandas. Understandingloc、iloc、atthe differences and use cases will let you process data more efficiently.
Difference between loc and iloc
| Feature | loc | iloc |
|---|---|---|
| Indexing method | Label-based indexing | Integer-based (positional) indexing |
| Slicing | Includes the end position | Excludes the end position |
| Single value | Returns a scalar | Returns a scalar |
| Recommended scenario | When there are explicit index labels | When selecting by position |
Example
# Create example DataFrame
df = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu", "Zhao Liu", "Qian Qi"],
"Age": [25, 30, 28, 35, 22],
"City": ["Beijing", "Shanghai", "Guangzhou", "Shenzhen", "Hangzhou"]
}, index=[1, 3, 5, 7, 9]) # Note: the index is not continuous
print("DataFrame:")
print(df)
print()
# loc: use label indexing (inclusive of end)
print("df.loc[1:5] (label slice, includes 5):")
print(df.loc[1:5])
print()
# iloc: use positional indexing (exclusive of end)
print("df.iloc[0:2] (position slice, excludes 2):")
print(df.iloc[0:2])
Usage of loc
Selecting rows
Example
df = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu", "Zhao Liu"],
"Age": [25, 30, 28, 35],
"City": ["Beijing", "Shanghai", "Guangzhou", "Shenzhen"]
}, index=["a", "b", "c", "d"])
# Select a single row (returns a Series)
print("Select one row:")
print(df.loc["a"])
print()
# Select multiple rows
print("Select multiple rows:")
print(df.loc[["a", "c"]])
print()
# Slice selection (includes start and end)
print("Slice selection:")
print(df.loc["a":"c"])
Selecting columns
Example
df = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu"],
"Age": [25, 30, 28],
"City": ["Beijing", "Shanghai", "Guangzhou"]
}, index=["a", "b", "c"])
# Select a single column
print("Select a single column:")
print(df.loc[:, "Name"])
print()
# Select multiple columns
print("Select multiple columns:")
print(df.loc[:, ["Name", "City"]])
print()
# Slice selection of columns
print("Slice select columns:")
print(df.loc[:, "Name":"City"])
Selecting specific rows and columns (recommended approach)
Example
df = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu", "Zhao Liu"],
"Age": [25, 30, 28, 35],
"City": ["Beijing", "Shanghai", "Guangzhou", "Shenzhen"]
}, index=["a", "b", "c", "d"])
# Select specific rows and columns
print("Select a single value:")
print(df.loc["a", "Name"]) # Returns "Zhang San"
print(type(df.loc["a", "Name"])) # The type is str
print()
# Select multiple rows and columns
print("Select subset:")
print(df.loc[["a", "c"], ["Name", "City"]])
print()
# Conditional selection
print("Rows with Age greater than 28:")
print(df.loc[df["Age"] > 28])
Usage of iloc
Selecting by position
Example
df = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu", "Zhao Liu"],
"Age": [25, 30, 28, 35],
"City": ["Beijing", "Shanghai", "Guangzhou", "Shenzhen"]
})
# Select row 0
print("Select row 0:")
print(df.iloc[0])
print()
# Select the first 3 rows
print("Select the first 3 rows:")
print(df.iloc[:3])
print()
# Select specific rows
print("Select specific rows:")
print(df.iloc[[0, 2, 3]])
print()
# Negative indexing (starting from the end)
print("Select the last row:")
print(df.iloc[-1])
print()
print("Select the last 3 rows:")
print(df.iloc[-3:])
Selecting columns
Example
df = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu"],
"Age": [25, 30, 28],
"City": ["Beijing", "Shanghai", "Guangzhou"]
})
# Select column 0
print("Column 0:")
print(df.iloc[:, 0])
print()
# Select columns 1 and 2
print("Select multiple columns:")
print(df.iloc[:, [1, 2]])
print()
# Slice selection
print("Slice select columns:")
print(df.iloc[:, 0:2])
at and iat (get a single value)
atandiatare accessors specifically for getting/setting a single value, which is faster thanlocandiloc>更快。
Example
import time
import numpy as np
# Create a large DataFrame for performance testing
df = pd.DataFrame(np.random.randn(1000, 10), columns=[f"col_{i}" for i in range(10)])
# at gets a single value (label indexing)
print(f"at: {df.at[0, 'col_0']}")
# iat gets a single value (positional indexing)
print(f"iat: {df.iat[0, 0]}")
# Performance comparison
n = 10000
start = time.time()
for _ in range(n):
_ = df.iloc[0, 0]
print(f"iloc time: {time.time() - start:.4f}s")
start = time.time()
for _ in range(n):
_ = df.iat[0, 0]
print(f"iat time: {time.time() - start:.4f}s")
Conditional selection
Filtering data using boolean conditions is one of the most common operations.
Example
df = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu", "Zhao Liu"],
"Age": [25, 30, 28, 35],
"City": ["Beijing", "Shanghai", "Guangzhou", "Beijing"],
"Salary": [12000, 15000, 11000, 18000]
})
# Single condition
print("Employees older than 28:")
print(df[df["Age"] > 28])
print()
# Multiple conditions (using & | ~)
print("Beijing and salary greater than 12000:")
print(df[(df["City"] == "Beijing") & (df["Salary"] > 12000)])
print()
# isin filtering
print("City is Beijing or Shanghai:")
print(df[df["City"].isin(["Beijing", "Shanghai"])])
print()
# String containment
print("Name contains 'San':")
print(df[df["Name"].str.contains("San")])
Notes
1. Slicing includes the end position
locslicing includes the end position,ilocdoes not include it.
2. An error occurs if the index does not exist
Usinglocif the index does not exist, a KeyError is raised. You can useloc[index_list]in combination withreindex。
3. It is recommended to useloc
becauselocMore readable and less error-prone, unless you need to select by position.
Other extensions
at/iatIs the fastest way to access a single value.loc/ilocUsed to select multiple values. Choosing the appropriate method can improve code performance and readability.