Pandas df.head() Function
head()is one of the most commonly used functions in Pandas DataFrame and Series, used to quickly view the beginning of a dataset. It returns the first n rows of data, allowing us to understand the basic structure and content of the data without loading the entire dataset.
This function is used very frequently in data analysis work, especially when handling large datasets, where we usually first usehead()to view the first few rows of data, confirm whether the data structure meets expectations, and then proceed with further analysis and processing.
Basic Syntax and Parameters
head()is a member function of DataFrame and Series, called through the dot operator.to invoke. It does not require any mandatory parameters, but accepts an optional parameter to specify the number of rows to return.
Syntax Format
DataFrame.head(n=5) Series.head(n=5)
Parameter Description
| Parameter | Type | Required | Description | Default Value |
|---|---|---|---|---|
| n | int | Optional | Returns the first n rows of data. If n is greater than the total number of rows, returns all data. | 5 |
Return Value Description
- Return value type: When the caller is a DataFrame, a DataFrame is returned; when the caller is a Series, a Series is returned.
- Number of rows returned: Returns at most n rows; if the data has fewer than n rows, returns all data.
Examples
Let's fully masterhead()usage through a series of examples.
Example 1: Basic Usage - View the First Few Rows of a DataFrame
First create a simple DataFrame, then usehead()to view the first few rows of data.
Example
# Create a sample DataFrame containing student grade data
data = {
'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve', 'Frank', 'Grace', 'Henry', 'Iris', 'Jack'],
'age': [18, 19, 17, 18, 20, 19, 18, 17, 19, 18],
'score': [85, 92, 78, 90, 88, 95, 82, 76, 89, 91],
'grade': ['A', 'A', 'B', 'A', 'B', 'A', 'B', 'C', 'B', 'A']
}
df = pd.DataFrame(data)
# Default returns the first 5 rows
print("Default first 5 rows:")
print(df.head())
# Specify returning the first 3 rows
print("nFirst 3 rows:")
print(df.head(3))
# Return the first 7 rows
print("nFirst 7 rows:")
print(df.head(7))
Output result:
默认前 5 行:
name age score grade
0 Alice 18 85 A
1 Bob 19 92 A
2 Charlie 17 78 B
3 David 18 90 A
4 Eve 20 88 B
前 3 行:
name age score grade
0 Alice 18 85 A
1 Bob 19 92 A
2 Charlie 17 78 B
前 7 行:
name age score grade grade
0 Alice 18 85 A
1 Bob 19 92 A
2 Charlie 17 78 B
3 David 18 90 A
4 Eve 20 88 B
5 Frank 19 95 A
6 Grace 18 82 B
Code explanation:
- Created a DataFrame containing 10 rows of student data.
df.head()Without parameters, it defaults to returning the first 5 rows of data; this is the most common usage.df.head(3)Returns the first 3 rows; the number of rows to view can be adjusted as needed.- When the specified n value is greater than the total number of data rows, all data is returned without an error.
Example 2: View the First Few Rows of a Series
head()Not only applicable to DataFrames, but also works for Series objects.
Example
# Create a Series containing a series of values
s = pd.Series([10, 20, 30, 40, 50, 60, 70, 80, 90, 100])
# View the first 3 elements of the Series
print("First 3 elements:")
print(s.head(3))
# Create a Series with an index
s2 = pd.Series([100, 200, 300, 400, 500], index=['a', 'b', 'c', 'd', 'e'])
print("nFirst 2 elements of the indexed Series:")
print(s2.head(2))
Output result:
前 3 个元素: 0 10 1 20 2 30 dtype: int64 带索引 Series 的前 2 个元素: a 100 b 200 dtype: int64
Code explanation:
- Series also supports the
head()method, returning the first n elements. - A Series with a custom index can also use
head()normally, and the original index will be preserved.
Example 3: Use in Combination with Other Functions
head()Often used in combination with other DataFrame functions for data exploration and analysis.
Example
import numpy as np
# Create a larger DataFrame
np.random.seed(42) # Set random seed to ensure reproducible results
df = pd.DataFrame({
'date': pd.date_range('2024-01-01', periods=100),
'value': np.random.randn(100).round(2),
'category': np.random.choice(['A', 'B', 'C'], 100)
})
# View data type information
print("Data type information:")
print(df.dtypes)
print()
# View the first 10 rows of data
print("First 10 rows of data:")
print(df.head(10))
# Sort first, then view the first few rows
df_sorted = df.sort_values('value', ascending=False)
print("nFirst 5 rows after sorting by value in descending order:")
print(df_sorted.head())
# View statistical information for the first few rows
print("nStatistical information for the first 5 rows:")
print(df.head().describe())
Output result:
数据类型信息:
date datetime64[ns]
value float64
category object
dtype: object
前 10 行数据:
date value category
0 2024-01-01 0.34 A
1 2024-01-02 -0.23 B
2 2024-01-03 0.54 C
3 2024-01-04 -1.58 B
4 2024-01-05 -0.29 后几行
...
按 value 降序排列后的前 5 行:
date value category
72 2024-03-13 2.87 C
55 2024-02-25 2.32 A
31 2024-02-01 2.21 B
64 2024-03-05 1.71 C
84 2024-03-25 1.55 A
前 5 行的统计信息:
age score
count 5.0 5.000
mean 18.0 87.200
std 1.0 5.403
min 17.0 78.000
50% 18.0 88.000
max 20.0 95.000
Code explanation:
head()Can be used in combination withsort_values()to sort first and then view the first few rows.head()The returned result is still a DataFrame, so other DataFrame methods can continue to be called.describe()Statistical summary can be performed onhead()the data returned.
Notes
head()Does not modify the original DataFrame or Series; it returns a new object.- When n is less than or equal to 0, an empty DataFrame or empty Series is returned.
- For large datasets, first using
head()to view the structure is a good practice. head()The returned data retains the original index values and does not renumber them.
Tip: In Jupyter Notebook or JupyterLab environments, directly entering the DataFrame variable name and executing it will by default display the
head()result, which greatly facilitates data exploration work.
Summary
head()is one of the most basic and practical data viewing functions in Pandas. It can quickly preview the beginning of a dataset, helping us understand key information such as data structure, column names, and data types.
In actual data analysis work, we usually: first usehead()to view the basic structure of the data, then usetail()to view the end of the data, and then useinfo()anddescribe()to understand the overall picture of the data. This "four-step data viewing process" is a fundamental skill that every data analyst should master.
Pandas Common Functions