Pandas concat() Function

Pandas 通用函数Pandas Common Functions


pd.concat()is used in the Pandas library toconcatenate (join) multiple Series or DataFrame objectsfunction. It can merge multiple objects into a new object along the row or column direction.

This is one of the most commonly used merging operations in data processing, whether it is appending data vertically or merging fields horizontally,pd.concat()it is fully capable.

Word Definition: concatis an abbreviation of "concatenate" (connect, link in series), meaning to connect multiple things end to end.


Basic Syntax and Parameters

pd.concat()is a top-level function of the Pandas library, used to concatenate multiple Series or DataFrame objects along a specified axis.

Syntax Format

pd.concat(objs, axis=0, join='outer', ignore_index=False, keys=None, levels=None, names=None, verify_integrity=False, sort=False)

Parameter Description

  • Parameter: objs
    • Type: Series, DataFrame, or a list of them.
    • Description: One or more Series/DataFrame objects to be concatenated. They can be a single object, a list, or a dictionary.
  • Parameter: axis
    • Type: Integer (0 or 1).
    • Description: Specifies the axis of concatenation.0means concatenating by rows (vertical appending),1means concatenating by columns (horizontal merging). Default is0。
  • Parameter: join
    • Type: String ('outer' or 'inner').
    • Description: How to handle when the indexes along the concatenation axis are inconsistent.'outer'means taking the union (retaining all indexes),'inner'means taking the intersection (retaining only common indexes). Default is'outer'。
  • Parameter: ignore_index
    • Type: Boolean.
    • Description: If it isTrue, the original index is not used after concatenation, instead a new integer index is created. Default isFalse。
  • Parameter: keys
    • Type: Sequence.
    • Description: Used to create a multi-level index after concatenation, identifying the source of each original object. Default isNone。

Function Description

  • Return Value: Returns the concatenated Series or DataFrame object. The type is consistent with the input objects.
  • Effect: Merges multiple objects into one object along the specified axis.

Examples

Through a series of examples from simple to complex, let's thoroughly masterpd.concat()the usage.

Example 1: Basic Usage - Vertically Concatenating Series

Example

import pandas as pd

# 1. Create two Series
s1 = pd.Series(['Alice', 'Bob', 'Charlie'], name='name')
s2 = pd.Series(['Diana', 'Eve', 'Frank'], name='name')

print(=== Original Series ===)
print("s1:", s1.tolist())
print("s2:", s2.tolist())

# 2. Use pd.concat() to concatenate vertically (default axis=0)
result = pd.concat([s1, s2])
print("\n=== Result of pd.concat([s1, s2]) vertical concatenation ===)
print(result)

# 3. Ignore the original index and create a new integer index
result_ignore = pd.concat([s1, s2], ignore_index=True)
print("\n=== Result of ignore_index=True ===)
print(result_ignore)

Expected output:

=== 原始 Series ===
s1: ['Alice', 'Bob', 'Charlie']
s2: ['Diana', 'Eve', 'Frank']

=== pd.concat([s1, s2]) 纵向拼接结果 ===
0      Alice
1        Bob
2    Charlie
0     Diana
1        Eve
2     Frank
dtype: object

=== ignore_index=True 结果 ===
0      Alice
1        Bob
2    Charlie
3      Diana
4        Eve
5      Frank
dtype: object

Code Analysis:

  1. pd.concat([s1, s2])Concatenate two Series vertically into a longer Series.
  2. By default, the original index is retained, so there will be duplicate index values 0, 1, 2.
  3. Settingignore_index=Truecan regenerate a continuous numeric index.

Example 2: Horizontally Concatenating DataFrame

Usingaxis=1you can concatenate multiple DataFrames by column, which is very useful when merging different fields.

Example

import pandas as pd

# 1. Create two DataFrames
df1 = pd.DataFrame({
    'name': ['Alice', 'Bob', 'Charlie'],
    'age': [25, 30, 35]
})

df2 = pd.DataFrame({
    'city': ['Beijing', 'Shanghai', 'Guangzhou'],
    'score': [85, 90, 95]
})

print(=== Original DataFrame ===)
print("df1:")
print(df1)
print("\ndf2:")
print(df2)

# 2. Horizontal concatenation (axis=1)
result = pd.concat([df1, df2], axis=1)
print("\n=== Horizontal concatenation with pd.concat([df1, df2], axis=1) ===)
print(result)

Expected output:

=== 原始 DataFrame ===
df1:
      name  age
0   Alice   25
1     Bob   30
2 Charlie   35

df2:
        city  score
0   Beijing     85
1  Shanghai     90
2  Guangzhou     95

=== pd.concat([df1, df2], axis=1) 横向拼接 ===
      name  age      city  score
0   Alice   25   Beijing     85
1     Bob   30  Shanghai     90
2 Charlie   35  Guangzhou     95

Code Analysis:

  • Usingaxis=1concatenate the two DataFrames by column.
  • Horizontal concatenation requires the two DataFrames to have the same number of rows (index alignment).
  • The concatenated DataFrame contains all columns of both objects.

Example 3: Using the keys Parameter to Create a Hierarchical Index

keysThe parameter can help us identify the data source, which is very useful when merging content from different datasets.

Example

import pandas as pd

# 1. Create a dictionary containing multiple Series
data = {
    'group_a': pd.Series([1, 2, 3, 4]),
    'group_b': pd.Series([5, 6, 7, 8]),
    'group_c': pd.Series([9, 10, 11, 12])
}

print(=== Concatenation using the keys parameter ===)
# 2. Use the keys parameter to create a hierarchical index
result = pd.concat(data, keys=['Group 1', 'Group 2', 'Group 3'])
print(result)

# 3. Access data at a specific level
print("\n=== Accessing the data of 'Group 1' ===)
print(result['Group 1'])

Expected output:

=== 使用 keys 参数拼接 ===
第一组  group_a    1
        group_b    2
        group_c    3
        group_a    4
第二组  group_a    5
        group_b    6
        group_c    7
        group_a    8
第二组  group_b    9
...

Code Analysis:

  • keysThe parameter creates a multi-level index (hierarchical index) for the concatenation result, with the first level identifying the source group of the data.
  • You can useresult['第一组']to access data of a specific group.

Example 4: Handling DataFrames with Different Columns (join Parameter)

When the DataFrames to be concatenated have different columns, you can usejointhe parameter to control the handling method.

Example

import pandas as pd

# 1. Create DataFrames with different columns
df1 = pd.DataFrame({
    'name': ['Alice', 'Bob'],
    'age': [25, 30]
})

df2 = pd.DataFrame({
    'name': ['Charlie', 'Diana'],
    'score': [85, 90]
})

print(=== Original DataFrame ===)
print("df1:")
print(df1)
print("\ndf2:")
print(df2)

# 2. Outer join (default) - keep all columns
print("\n=== join='outer' (outer join) ===)
print(pd.concat([df1, df2], join='outer'))

# 3. Inner join - keep only common columns
print("\n=== join='inner' (inner join) ===)
print(pd.concat([df1, df2], join='inner'))

Expected output:

=== 原始 DataFrame ===
df1:
    name  age
0  Alice   25
1    Bob   30

df2:
      name  score
0  Charlie     85
1    Diana     90

=== join='outer' (外连接) ===
      name   age  score
0    Alice  25.0   NaN
1      Bob  30.0   NaN
0  Charlie  NaN  85.0
1   Diana  NaN  90.0

=== join='inner' (内连接) ===
      name
0    Alice
1      Bob
0  Charlie
1  Diana

Code Analysis:

  • join='outer'(default) keeps all columns, missing columns are filled withNaN。
  • join='inner'keeps only the columns common to both DataFrames (here it isname)。

Tip: pd.concat()is suitable for simple concatenation operations. If you need more complex merging (such as matching based on column values), please refer topd.merge()function.

Pandas 常用函数Pandas Common Functions

Other Extensions