Pandas concat() Function
pd.concat()is used in the Pandas library toconcatenate (join) multiple Series or DataFrame objectsfunction. It can merge multiple objects into a new object along the row or column direction.
This is one of the most commonly used merging operations in data processing, whether it is appending data vertically or merging fields horizontally,pd.concat()it is fully capable.
Word Definition: concatis an abbreviation of "concatenate" (connect, link in series), meaning to connect multiple things end to end.
Basic Syntax and Parameters
pd.concat()is a top-level function of the Pandas library, used to concatenate multiple Series or DataFrame objects along a specified axis.
Syntax Format
pd.concat(objs, axis=0, join='outer', ignore_index=False, keys=None, levels=None, names=None, verify_integrity=False, sort=False)
Parameter Description
- Parameter:
objs- Type: Series, DataFrame, or a list of them.
- Description: One or more Series/DataFrame objects to be concatenated. They can be a single object, a list, or a dictionary.
- Parameter:
axis- Type: Integer (0 or 1).
- Description: Specifies the axis of concatenation.
0means concatenating by rows (vertical appending),1means concatenating by columns (horizontal merging). Default is0。
- Parameter:
join- Type: String ('outer' or 'inner').
- Description: How to handle when the indexes along the concatenation axis are inconsistent.
'outer'means taking the union (retaining all indexes),'inner'means taking the intersection (retaining only common indexes). Default is'outer'。
- Parameter:
ignore_index- Type: Boolean.
- Description: If it is
True, the original index is not used after concatenation, instead a new integer index is created. Default isFalse。
- Parameter:
keys- Type: Sequence.
- Description: Used to create a multi-level index after concatenation, identifying the source of each original object. Default is
None。
Function Description
- Return Value: Returns the concatenated Series or DataFrame object. The type is consistent with the input objects.
- Effect: Merges multiple objects into one object along the specified axis.
Examples
Through a series of examples from simple to complex, let's thoroughly masterpd.concat()the usage.
Example 1: Basic Usage - Vertically Concatenating Series
Example
# 1. Create two Series
s1 = pd.Series(['Alice', 'Bob', 'Charlie'], name='name')
s2 = pd.Series(['Diana', 'Eve', 'Frank'], name='name')
print(=== Original Series ===)
print("s1:", s1.tolist())
print("s2:", s2.tolist())
# 2. Use pd.concat() to concatenate vertically (default axis=0)
result = pd.concat([s1, s2])
print("\n=== Result of pd.concat([s1, s2]) vertical concatenation ===)
print(result)
# 3. Ignore the original index and create a new integer index
result_ignore = pd.concat([s1, s2], ignore_index=True)
print("\n=== Result of ignore_index=True ===)
print(result_ignore)
Expected output:
=== 原始 Series === s1: ['Alice', 'Bob', 'Charlie'] s2: ['Diana', 'Eve', 'Frank'] === pd.concat([s1, s2]) 纵向拼接结果 === 0 Alice 1 Bob 2 Charlie 0 Diana 1 Eve 2 Frank dtype: object === ignore_index=True 结果 === 0 Alice 1 Bob 2 Charlie 3 Diana 4 Eve 5 Frank dtype: object
Code Analysis:
pd.concat([s1, s2])Concatenate two Series vertically into a longer Series.- By default, the original index is retained, so there will be duplicate index values 0, 1, 2.
- Setting
ignore_index=Truecan regenerate a continuous numeric index.
Example 2: Horizontally Concatenating DataFrame
Usingaxis=1you can concatenate multiple DataFrames by column, which is very useful when merging different fields.
Example
# 1. Create two DataFrames
df1 = pd.DataFrame({
'name': ['Alice', 'Bob', 'Charlie'],
'age': [25, 30, 35]
})
df2 = pd.DataFrame({
'city': ['Beijing', 'Shanghai', 'Guangzhou'],
'score': [85, 90, 95]
})
print(=== Original DataFrame ===)
print("df1:")
print(df1)
print("\ndf2:")
print(df2)
# 2. Horizontal concatenation (axis=1)
result = pd.concat([df1, df2], axis=1)
print("\n=== Horizontal concatenation with pd.concat([df1, df2], axis=1) ===)
print(result)
Expected output:
=== 原始 DataFrame ===
df1:
name age
0 Alice 25
1 Bob 30
2 Charlie 35
df2:
city score
0 Beijing 85
1 Shanghai 90
2 Guangzhou 95
=== pd.concat([df1, df2], axis=1) 横向拼接 ===
name age city score
0 Alice 25 Beijing 85
1 Bob 30 Shanghai 90
2 Charlie 35 Guangzhou 95
Code Analysis:
- Using
axis=1concatenate the two DataFrames by column. - Horizontal concatenation requires the two DataFrames to have the same number of rows (index alignment).
- The concatenated DataFrame contains all columns of both objects.
Example 3: Using the keys Parameter to Create a Hierarchical Index
keysThe parameter can help us identify the data source, which is very useful when merging content from different datasets.
Example
# 1. Create a dictionary containing multiple Series
data = {
'group_a': pd.Series([1, 2, 3, 4]),
'group_b': pd.Series([5, 6, 7, 8]),
'group_c': pd.Series([9, 10, 11, 12])
}
print(=== Concatenation using the keys parameter ===)
# 2. Use the keys parameter to create a hierarchical index
result = pd.concat(data, keys=['Group 1', 'Group 2', 'Group 3'])
print(result)
# 3. Access data at a specific level
print("\n=== Accessing the data of 'Group 1' ===)
print(result['Group 1'])
Expected output:
=== 使用 keys 参数拼接 ===
第一组 group_a 1
group_b 2
group_c 3
group_a 4
第二组 group_a 5
group_b 6
group_c 7
group_a 8
第二组 group_b 9
...
Code Analysis:
keysThe parameter creates a multi-level index (hierarchical index) for the concatenation result, with the first level identifying the source group of the data.- You can use
result['第一组']to access data of a specific group.
Example 4: Handling DataFrames with Different Columns (join Parameter)
When the DataFrames to be concatenated have different columns, you can usejointhe parameter to control the handling method.
Example
# 1. Create DataFrames with different columns
df1 = pd.DataFrame({
'name': ['Alice', 'Bob'],
'age': [25, 30]
})
df2 = pd.DataFrame({
'name': ['Charlie', 'Diana'],
'score': [85, 90]
})
print(=== Original DataFrame ===)
print("df1:")
print(df1)
print("\ndf2:")
print(df2)
# 2. Outer join (default) - keep all columns
print("\n=== join='outer' (outer join) ===)
print(pd.concat([df1, df2], join='outer'))
# 3. Inner join - keep only common columns
print("\n=== join='inner' (inner join) ===)
print(pd.concat([df1, df2], join='inner'))
Expected output:
=== 原始 DataFrame ===
df1:
name age
0 Alice 25
1 Bob 30
df2:
name score
0 Charlie 85
1 Diana 90
=== join='outer' (外连接) ===
name age score
0 Alice 25.0 NaN
1 Bob 30.0 NaN
0 Charlie NaN 85.0
1 Diana NaN 90.0
=== join='inner' (内连接) ===
name
0 Alice
1 Bob
0 Charlie
1 Diana
Code Analysis:
join='outer'(default) keeps all columns, missing columns are filled withNaN。join='inner'keeps only the columns common to both DataFrames (here it isname)。
Other ExtensionsTip:
pd.concat()is suitable for simple concatenation operations. If you need more complex merging (such as matching based on column values), please refer topd.merge()function.
Pandas Common Functions