Pandas Series.var() Function
Series.var()It is a function in Pandas used to calculate the variance of a Series. Variance is the square of the standard deviation. It measures the dispersion of data and is a fundamental metric in statistical analysis.
The larger the variance, the more dispersed the data; the smaller the variance, the more concentrated the data. It is widely used in fields such as statistical analysis, machine learning, and signal processing.
Basic Syntax and Parameters
var()It is a member function of the Series object, called directly via the dot operator.
Syntax Format
Series.var(axis=None, skipna=True, level=None, numeric_only=None, ddof=1, **kwargs)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| axis | int | Specify the axis. Series has only one row of data; this parameter is mainly for compatibility with DataFrame. | None |
| skipna | bool | If True, skip NaN values during calculation; if False, the result will return NaN when encountering NaN. | True |
| level | int or str | If the Series has a MultiIndex, specify the level to calculate. | None |
| numeric_only | bool | If True, only calculate for numeric data; otherwise, it will try to convert to numeric. | False |
| ddof | int | Degrees of freedom adjustment parameter. ddof=1 uses sample variance (n-1), ddof=0 uses population variance (n). | 1 |
Return Value
- Return Type:
float - DescriptionReturns the variance of the elements in the Series. By default, the sample variance is used (divided by n-1).
Examples
Through a series of examples from simple to complex, let's thoroughly masterSeries.var()its usage.
Example 1: Basic Usage - Understanding the Concept of Variance
Variance is the square of the standard deviation, measuring the degree of deviation between data and the mean.
Example
# Two sets of data
group_a = pd.Series([85, 86, 87, 88, 89])
group_b = pd.Series([70, 75, 85, 95, 100])
print("Group A data (more concentrated):")
print(group_a)
print(f"Mean: {group_a.mean():.2f}")
print(f"Variance: {group_a.var():.2f}")
print(f"Standard deviation: {group_a.std():.2f}")
print()
print("Group B data (more dispersed):")
print(group_b)
print(f"Mean: {group_b.mean():.2f}")
print(f"Variance: {group_b.var():.2f}")
print(f"Standard deviation: {group_b.std():.2f}")
print()
print("Note: Variance is the square of the standard deviation (2.50^2 = 6.25, 12.50^2 = 156.25)")
Output:
A组数据(较集中): 0 85 1 86 2 87 3 88 4 89 dtype: int64 平均值:85.00 方差:2.50 标准差:1.58 B组数据(较分散): 0 70 1 75 2 85 3 95 4 100 dtype: int64 平均值:85.00 方差:156.25 标准差:12.50 注意:方差是标准差的平方 (1.58^2 ≈ 2.50, 12.50^2 = 156.25)
Code Analysis:
- Group A has a variance of 2.50, the data is very concentrated.
- Group B has a variance of 156.25, the data is very dispersed.
- Variance = Standard deviation² (the approximate value is due to the adjustment of ddof=1).
Example 2: The Role of the ddof Parameter
ddofThis parameter controls whether to use sample variance or population variance.
Example
# Create a data set
data = pd.Series([2, 4, 4, 4, 5, 5, 7, 9])
print("Data:")
print(data)
print()
# Default ddof=1, use sample variance (divided by n-1)
sample_var = data.var(ddof=1)
print(f"Sample variance (ddof=1): {sample_var:.4f}")
# ddof=0, use population variance (divided by n)
population_var = data.var(ddof=0)
print(f"Population variance (ddof=0): {population_var:.4f}")
print()
# Verify: the square root of the sample variance is approximately equal to the sample standard deviation
import math
print(f"Square root of sample variance: {math.sqrt(sample_var):.4f}")
print(f"Sample standard deviation: {data.std():.4f}")
Output:
数据: 0 2 1 4 2 4 3 4 4 5 5 5 6 7 7 9 dtype: int64 样本方差(ddof=1):5.1429 总体方差(ddof=0):4.5000 样本方差的平方根:2.2678 样本标准差:2.2678
Example 3: Handling Data Containing Missing Values
Example
import numpy as np
# Create a Series containing missing values
data_with_nan = pd.Series([10, 20, np.nan, 30, 40, np.nan, 50])
print("Data containing missing values:")
print(data_with_nan)
print()
# Default skipna=True
var_skipna = data_with_nan.var()
print(f"Variance when skipna=True (default): {var_skipna:.4f}")
# Set skipna=False
var_no_skipna = data_with_nan.var(skipna=False)
print(f"Variance when skipna=False: {var_no_skipna}")
Output:
包含缺失值的数据: 0 10.0 1 20.0 2 NaN 3 30.0 4 40.0 5 NaN 6 50.0 dtype: float64 skipna=True(默认)时的方差:250.0000 skipna=False 时的方差:nan
Example 4: Practical Application Comparison of Variance and Standard Deviation
In practical applications, variance and standard deviation each have their own advantages.
Example
# Simulate monthly returns (%) of two investment portfolios
portfolio_a = pd.Series([2.5, 3.0, -1.0, 1.5, 2.0, 2.8, -0.5, 1.2, 1.8, 2.2])
portfolio_b = pd.Series([5.0, -3.5, 8.0, -2.0, 6.5, -4.0, 7.0, -1.5, 4.5, -2.0])
print("Portfolio A (monthly return %):")
print(portfolio_a)
print(f"Variance: {portfolio_a.var():.4f}")
print(f"Standard deviation (volatility): {portfolio_a.std():.2f}%")
print()
print("Portfolio B (monthly return %):")
print(portfolio_b)
print(f"Variance: {portfolio_b.var():.4f}")
print(f"Standard deviation (volatility): {portfolio_b.std():.2f}%")
print()
print("Analysis:")
print("- Portfolio A has smaller variance and standard deviation, and performs more stably")
print("- Portfolio B has larger variance and standard deviation, and carries higher risk")
print("- From a risk perspective, Portfolio A is more suitable for conservative investors")
Output:
投资组合 A(月度收益 %): 0 2.5 1 3.0 2 1.0 3 1.5 4 2.0 5 2.8 6 0.5 7 1.2 8 1.8 9 2.2 dtype: int64 方差:1.29 标准差(波动率):1.14% 投资组合 B(月度收益 %): 0 5.0 1 -3.5 2 8.0 3 -2.0 4 6.5 5 -4.0 6 7.0 7 -1.5 8 4.5 9 -2.0 dtype: int64 方差:19.56 标准差(波动率):4.42% 分析:组合 B 的风险(波动率)约是组合 A 的 4 倍。
Notes
- Variance is the square of the standard deviation. The relationship between the two is: std = sqrt(var).
- The sample variance (ddof=1) is used by default, which is suitable for statistical analysis.
- The unit of variance is the square of the original data unit, which is not easy to understand intuitively; the standard deviation retains the original unit and is easier to interpret.
- In scenarios where variance is needed for mathematical calculations (such as covariance analysis), variance should be used.
Summary
Series.var()It is a basic function in statistical analysis. Its main features include:
- The sample variance (ddof=1) is used by default.
- Variance is the square of the standard deviation, which is more convenient in mathematical operations.
- The standard deviation retains the original unit and is easier to understand intuitively.
- In the financial field, both variance and standard deviation are used to measure risk (volatility).
In practical applications, if you need to intuitively understand the dispersion of data, using the standard deviation is more appropriate; if you need to perform mathematical operations (such as calculating covariance), using variance is more convenient.
Other Extensions
Common Pandas Functions