Pandas Series.str.split() Function
Series.str.split()It is a function in Pandas used to split strings.
In data processing, we often need to split a string into multiple parts, for example, splitting a sentence into words, or splitting comma-separated values into lists.split()The function can split a string into multiple parts according to a specified delimiter.
Word Meaning:splitsplit means "separate, split", indicating dividing a string into multiple parts.
Basic Syntax and Parameters
str.split()It is a string accessor method of Series, so you need to first have a Series containing strings, and then use.strthe accessor to call it.
Syntax Format
Series.str.split(pat=None, n=-1, expand=False)
Parameter Description
| Parameter | Type | Required? | Description | Default Value |
|---|---|---|---|---|
| pat | str | Optional | Delimiter, can be a string or a regular expression. Defaults to whitespace. | None (whitespace) |
| n | int | Optional | Number of splits. -1 means no limit, split into as many parts as possible. | -1 |
| expand | bool | Optional | Whether to expand the result into a DataFrame. Defaults to False (returns a Series of lists). | False |
Function Description
- Return Value: By default, returns a Series containing lists, where each element in each list is a substring after splitting. When
expand=Trueexpand=True, it returns a DataFrame. - Effect: Splits each string into multiple parts according to the specified delimiter.
- Note: By default, whitespace is used as the delimiter; you can specify a custom delimiter.
Examples
Let us, through a series of examples from simple to complex, thoroughly masterstr.split()the usage.
Example 1: Basic Usage - Splitting by Whitespace
Example
# Create a Series containing sentences
s = pd.Series(['hello world', 'example python', 'pandas data analysis'])
# Split by whitespace (default behavior)
result = s.str.split()
print("Original Series:")
print(s)
print("nResult after splitting:")
print(result)
Output:
原始 Series: 0 hello world 1 example python 2 pandas data analysis dtype: object 拆分后的结果: 0 [hello, world] 1 [example, python] 2 [pandas, data, analysis]
Code explanation:
s.str.split()By default, strings are split by whitespace (spaces, tabs, etc.).- The return value is a Series containing lists, each list contains the substrings after splitting.
- 'pandas data analysis' is split into three parts.
Example 2: Splitting by a Specified Delimiter
You can usepatthe parameter to specify a custom delimiter.
Example
# Create a Series containing comma-separated values
s = pd.Series(['apple,banana,orange', 'dog,cat,bird', 'red,green,blue'])
# Split by comma
result = s.str.split(',')
print("Original Series:")
print(s)
print("nResult after splitting by comma:")
print(result)
Output:
原始 Series: 0 apple,banana,orange 1 dog,cat,bird 2 red,green,blue dtype: object 拆分后的结果: 0 [apple, banana, orange] 1 [dog, cat, bird] 2 [red, green, blue]
Code explanation:
s.str.split(',')Split each string by comma.- This is a common method for processing CSV data or comma-separated values.
Example 3: Limiting the Number of Splits
Using thenparameter, you can limit the number of splits.
Example
s = pd.Series(['a,b,c,d,e', '1,2,3,4,5'])
# Limit to only 2 parts
result_2 = s.str.split(',', n=2)
# Limit to only 1 part (i.e., split only once)
result_1 = s.str.split(',', n=1)
print("Original Series:")
print(s)
print("nLimit to 2 parts:")
print(result_2)
print("nLimit to 1 part:")
print(result_1)
Output:
原始 Series: 0 a,b,c,d,e 1 1,2,3,4,5 dtype: object 限制拆分为 2 部分: 0 [a, b, c,d,e] 1 [1, 2, 3,4,5] 限制拆分为 1 部分: 0 [a, b,c,d,e] 1 [1, 2,3,4,5]
Code explanation:
n=2Means splitting into at most 2 parts, with the last part containing all the remaining content.n=1Means splitting only at the first delimiter, resulting in 2 parts.
Example 4: Expanding to a DataFrame
Whenexpand=TrueWhen expand=True, the split result is expanded into a DataFrame.
Example
# Create a Series containing comma-separated values
s = pd.Series(['apple,banana,orange', 'dog,cat,bird', 'red,green,blue'])
# Expand to DataFrame
result = s.str.split(',', expand=True)
print("Original Series:")
print(s)
print("nExpanded to DataFrame:")
print(result)
print("nDataFrame type:", type(result))
Output:
原始 Series:
0 apple,banana,orange
1 dog,cat,bird
2 red,green,blue
dtype: object
展开为 DataFrame:
0 1 2
0 apple banana orange
1 dog cat bird
2 red green blue
DataFrame 类型: <class 'pandas.core.frame.DataFrame'>
Code explanation:
expand=TrueExpand the split result into a DataFrame.- Each column represents a position after splitting.
- This is very useful when handling structured data.
Example 5: Splitting with Regular Expressions
split()It also supports using regular expressions as delimiters.
Example
# Create a Series containing different delimiters
s = pd.Series(['hello-world', 'example_python', 'pandas#tutorial'])
# Split by regular expression (match any one of -, _, #)
result = s.str.split(r'[-_#]')
print("Original Series:")
print(s)
print("nSplit by regular expression:")
print(result)
Output:
原始 Series: 0 hello-world 1 example_python 2 pandas#tutorial dtype: object 按正则表达式拆分: 0 [hello, world] 1 [example, python] 2 [pandas, tutorial]
Code explanation:
r'[-_#]'It is a regular expression that matches any one of '-', '_', or '#'.- It can handle multiple different delimiters at the same time.
Example 6: Handling Real Data
In practical applications,split()it is often used in combination with other functions.
Example
# Simulate data extracted from logs
logs = pd.Series([
'2024-01-01 10:30:45 ERROR Connection failed',
'2024-01-01 10:31:12 INFO User logged in',
'2024-01-01 10:32:00 WARNING Memory usage high'
])
# Split the log content
log_parts = logs.str.split(r's+', expand=True)
print("Original log:")
print(logs)
print("nSplit DataFrame:")
print(log_parts)
# Rename columns
log_parts.columns = ['timestamp', 'level', 'message']
print("nResult after renaming:")
print(log_parts)
Output:
原始日志:
0 2024-01-01 10:30:45 ERROR Connection failed
1 2024-01-01 10:31:12 INFO User logged in
2 2024-01-01 10:32:00 WARNING Memory usage high
dtype: object
拆分后的 DataFrame:
0 1 2
0 2024-01-01 10:30:45 ERROR Connection failed
1 2024-01-01 10:31:12 INFO User logged in
2 2024-01-01 10:32:00 WARNING Memory usage high
重命名后的结果:
timestamp level message
0 2024-01-01 10:30:45 ERROR Connection failed
1 2024-01-01 10:31:12 INFO User logged in
2 2024-01-01 10:32:00 WARNING Memory usage high
Code explanation:
r's+'Matches one or more whitespace characters, splitting the log into multiple parts.expand=TrueExpand the split result into a DataFrame.- By renaming the columns, you can conveniently access and process each part.
Notes
str.split()By default, whitespace is used as the delimiter.- When
expand=FalseWhen expand=False, it returns a Series containing lists. - When
expand=TrueWhen expand=True, it returns a DataFrame, and the number of columns depends on the longest split result. - If a string does not contain the delimiter, the returned list has only one element (the original string).
- Regular expressions can be used as delimiters, providing a more flexible way to split.
- If the Series contains NaN values, NaN is returned.
Other Extensions
Common Pandas Functions