Pandas Data Concatenation (concat / append)
Data concatenation joins multiple DataFrames or Series together by rows or by columns.pd.concatis the main concatenation function,appendis the simplified version (deprecated, concat is recommended instead).
Basic Usage of concat
pd.concat()Can concatenate multiple DataFrames or Series along an axis.
Row-wise concatenation (stacking vertically)
Example
import pandas as pd
# Create two DataFrames
df1 = pd.DataFrame({
"Name": ["Zhang San", "Li Si"],
"Age": [25, 30]
})
df2 = pd.DataFrame({
"Name": ["Wang Wu", "Zhao Liu"],
"Age": [28, 35]
})
print("DataFrame 1:")
print(df1)
print()
print("DataFrame 2:")
print(df2)
print()
# Concatenate vertically
result = pd.concat([df1, df2], ignore_index=True)
print("Concatenation result:")
print(result)
# Create two DataFrames
df1 = pd.DataFrame({
"Name": ["Zhang San", "Li Si"],
"Age": [25, 30]
})
df2 = pd.DataFrame({
"Name": ["Wang Wu", "Zhao Liu"],
"Age": [28, 35]
})
print("DataFrame 1:")
print(df1)
print()
print("DataFrame 2:")
print(df2)
print()
# Concatenate vertically
result = pd.concat([df1, df2], ignore_index=True)
print("Concatenation result:")
print(result)
Column-wise concatenation (stacking horizontally)
Example
import pandas as pd
df1 = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu"]
})
df2 = pd.DataFrame({
"Age": [25, 30, 28],
"City": ["Beijing", "Shanghai", "Guangzhou"]
})
# Concatenate horizontally
result = pd.concat([df1, df2], axis=1)
print("Horizontal concatenation:")
print(result)
df1 = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu"]
})
df2 = pd.DataFrame({
"Age": [25, 30, 28],
"City": ["Beijing", "Shanghai", "Guangzhou"]
})
# Concatenate horizontally
result = pd.concat([df1, df2], axis=1)
print("Horizontal concatenation:")
print(result)
axis=0Indicates row-wise concatenation (adding rows),axis=1Indicates column-wise concatenation (adding columns).
Handling Duplicate Indexes
ignore_index
Example
import pandas as pd
df1 = pd.DataFrame({
"Name": ["Zhang San", "Li Si"],
"Age": [25, 30]
}, index=[0, 1])
df2 = pd.DataFrame({
"Name": ["Wang Wu", "Zhao Liu"],
"Age": [28, 35]
}, index=[0, 1])
# Keep the original indexes by default
print("Keep original indexes:")
print(pd.concat([df1, df2]))
print()
# Ignore old indexes and regenerate them
print("Ignore original indexes:")
print(pd.concat([df1, df2], ignore_index=True))
df1 = pd.DataFrame({
"Name": ["Zhang San", "Li Si"],
"Age": [25, 30]
}, index=[0, 1])
df2 = pd.DataFrame({
"Name": ["Wang Wu", "Zhao Liu"],
"Age": [28, 35]
}, index=[0, 1])
# Keep the original indexes by default
print("Keep original indexes:")
print(pd.concat([df1, df2]))
print()
# Ignore old indexes and regenerate them
print("Ignore original indexes:")
print(pd.concat([df1, df2], ignore_index=True))
Verifying Duplicate Keys
Example
import pandas as pd
df1 = pd.DataFrame({
"A": [1, 2]
})
df2 = pd.DataFrame({
"A": [3, 4]
})
# Check whether there are duplicate keys
print("Verification object:")
print(pd.concat([df1, df2], verify_integrity=True))
df1 = pd.DataFrame({
"A": [1, 2]
})
df2 = pd.DataFrame({
"A": [3, 4]
})
# Check whether there are duplicate keys
print("Verification object:")
print(pd.concat([df1, df2], verify_integrity=True))
Handling Column Mismatches
The join Parameter
Example
import pandas as pd
df1 = pd.DataFrame({
"A": [1, 2, 3],
"B": ["a", "b", "c"]
})
df2 = pd.DataFrame({
"B": ["x", "y", "z"],
"C": [10, 20, 30]
})
print("df1:")
print(df1)
print()
print("df2:")
print(df2)
print()
# Outer join (default): keep all columns
print("Outer join (keep all columns):")
print(pd.concat([df1, df2], join="outer"))
print()
# Inner join: keep only common columns
print("Inner join (keep common columns):")
print(pd.concat([df1, df2], join="inner"))
df1 = pd.DataFrame({
"A": [1, 2, 3],
"B": ["a", "b", "c"]
})
df2 = pd.DataFrame({
"B": ["x", "y", "z"],
"C": [10, 20, 30]
})
print("df1:")
print(df1)
print()
print("df2:")
print(df2)
print()
# Outer join (default): keep all columns
print("Outer join (keep all columns):")
print(pd.concat([df1, df2], join="outer"))
print()
# Inner join: keep only common columns
print("Inner join (keep common columns):")
print(pd.concat([df1, df2], join="inner"))
Add Only New Columns
Example
import pandas as pd
df1 = pd.DataFrame({
"Name": ["Zhang San", "Li Si"],
"Age": [25, 30]
})
df2 = pd.DataFrame({
"City": ["Beijing", "Shanghai"]
})
# Add the columns of df2 to df1
result = pd.concat([df1, df2], axis=1)
print("Add only new columns:")
print(result)
df1 = pd.DataFrame({
"Name": ["Zhang San", "Li Si"],
"Age": [25, 30]
})
df2 = pd.DataFrame({
"City": ["Beijing", "Shanghai"]
})
# Add the columns of df2 to df1
result = pd.concat([df1, df2], axis=1)
print("Add only new columns:")
print(result)
The keys Parameter for Creating a Hierarchical Index
Example
import pandas as pd
df1 = pd.DataFrame({"A": [1, 2], "B": [3, 4]})
df2 = pd.DataFrame({"A": [5, 6], "B": [7, 8]})
df3 = pd.DataFrame({"A": [9, 10], "B": [11, 12]})
# Use the keys parameter to create a hierarchical index
result = pd.concat([df1, df2, df3], keys=["First year", "Second year", "Third year"])
print("Concatenation with hierarchical index:")
print(result)
print()
# Get data from the hierarchical index
print("Get second-year data:")
print(result.loc["Second year"])
df1 = pd.DataFrame({"A": [1, 2], "B": [3, 4]})
df2 = pd.DataFrame({"A": [5, 6], "B": [7, 8]})
df3 = pd.DataFrame({"A": [9, 10], "B": [11, 12]})
# Use the keys parameter to create a hierarchical index
result = pd.concat([df1, df2, df3], keys=["First year", "Second year", "Third year"])
print("Concatenation with hierarchical index:")
print(result)
print()
# Get data from the hierarchical index
print("Get second-year data:")
print(result.loc["Second year"])
Practical: Combining Multiple Months of Data
Example
import pandas as pd
# Simulate sales data for multiple months
jan_sales = pd.DataFrame({
"Month": ["2024-01"] * 3,
"Product": ["A", "B", "C"],
"Sales": [100, 150, 80]
})
feb_sales = pd.DataFrame({
"Month": ["2024-02"] * 3,
"Product": ["A", "B", "C"],
"Sales": [120, 140, 90]
})
mar_sales = pd.DataFrame({
"Month": ["2024-03"] * 3,
"Product": ["A", "B", "C"],
"Sales": [110, 160, 85]
})
# Combine first-quarter data
quarterly = pd.concat([jan_sales, feb_sales, mar_sales], ignore_index=True)
print("First-quarter summary:")
print(quarterly)
print()
# Aggregate by month
monthly_summary = quarterly.groupby("Month")["Sales"].sum()
print("Monthly sales summary:")
print(monthly_summary)
# Simulate sales data for multiple months
jan_sales = pd.DataFrame({
"Month": ["2024-01"] * 3,
"Product": ["A", "B", "C"],
"Sales": [100, 150, 80]
})
feb_sales = pd.DataFrame({
"Month": ["2024-02"] * 3,
"Product": ["A", "B", "C"],
"Sales": [120, 140, 90]
})
mar_sales = pd.DataFrame({
"Month": ["2024-03"] * 3,
"Product": ["A", "B", "C"],
"Sales": [110, 160, 85]
})
# Combine first-quarter data
quarterly = pd.concat([jan_sales, feb_sales, mar_sales], ignore_index=True)
print("First-quarter summary:")
print(quarterly)
print()
# Aggregate by month
monthly_summary = quarterly.groupby("Month")["Sales"].sum()
print("Monthly sales summary:")
print(monthly_summary)
The append Method (Deprecated)
DataFrame.append()Deprecated in Pandas 2.0 and not recommended. Please usepd.concat()instead.
# 不推荐(已废弃) result = df1.append(df2) # 推荐 result = pd.concat([df1, df2])
Other Extensions
concatIs the standard method for concatenating data in Pandas, with better performance and more complete functionality.