Matplotlib hist() Function


Matplotlib 参考文档Matplotlib Reference Documentation

hist()Used to plot a histogram, dividing data into several intervals (bins) and counting the frequency of data occurrences in each interval.

Histograms are basic tools for exploring data distribution (central tendency, dispersion, skewness, etc.).

Function Definition

pyplot Interface

matplotlib.pyplot.hist(x, bins=None, range=None, density=False,
    weights=None, cumulative=False, bottom=None, histtype='bar',
    align='mid', orientation='vertical', rwidth=None, log=False,
    color=None, label=None, stacked=False, **kwargs)

Axes Interface

Axes.hist(x, bins=None, range=None, density=False, weights=None,
    cumulative=False, bottom=None, histtype='bar', align='mid',
    orientation='vertical', rwidth=None, log=False, color=None,
    label=None, stacked=False, **kwargs)

Parameter Description

ParameterTypeDescription
xarray or sequence of arraysInput data; multiple datasets can be passed in.
binsint or sequenceNumber of bins or sequence of bin edges. By default, the 'auto' method is used to automatically select.
rangetupleData range (lower, upper); data outside the range is ignored.
densityboolIf True, plot a probability density histogram (total area equals 1) instead of counts.
weightsarray-likeWeight of each data point.
cumulativebool or -1If True, plot a cumulative histogram; -1 means cumulative from largest to smallest.
histtypestrHistogram type: 'bar' (default), 'barstacked' (stacked), 'step' (step line), 'stepfilled' (filled step).
alignstrBin alignment: 'mid' (centered, default), 'left' (left-aligned), 'right' (right-aligned).
orientationstr'vertical' (vertical, default) or 'horizontal' (horizontal)
rwidthfloatRelative width of bars; 1.0 means no gap.
logboolIf True, the y-axis uses logarithmic scale.
colorcolor or listBar color
labelstr or listLegend label
stackedboolWhether to display multiple datasets as stacked.

The return value of hist() is a tuple of three items:(n, bins, patches)n is the count/density of each bin, bins are the bin edges, and patches are the bar objects drawn.


Usage Examples

Example 1: Basic Histogram

Example

import matplotlib.pyplot as plt
import numpy as np

# Generate normally distributed random data
np.random.seed(42)
data = np.random.randn(1000)

fig, ax = plt.subplots(layout='constrained')

# Draw histogram
n, bins, patches = ax.hist(data, bins=30,
                           color='steelblue', edgecolor='white')

# Highlight the bin containing the maximum value
max_idx = np.argmax(n)
patches[max_idx].set_facecolor('#e74c3c')

ax.set_title('Histogram of Normal Distribution')
ax.set_xlabel('Value')
ax.set_ylabel('Frequency')
ax.axvline(x=0, color='red', linestyle='--', alpha=0.7)
ax.grid(axis='y', alpha=0.3)
plt.show()

Example 2: Comparison Histogram of Multiple Data Sets

Example

import matplotlib.pyplot as plt
import numpy as np

np.random.seed(42)

# Three sets of data with different distributions
data1 = np.random.normal(0, 1, 1000)     # Standard normal
data2 = np.random.normal(2, 1.5, 800)    # mean=2, std=1.5
data3 = np.random.normal(-1, 0.5, 600)   # mean=-1, std=0.5

fig, ax = plt.subplots(figsize=(8, 5), layout='constrained')

# Compare multiple datasets using transparency and different colors
ax.hist(data1, bins=30, alpha=0.6, color='steelblue',
        label='N(0, 1)', edgecolor='white')
ax.hist(data2, bins=30, alpha=0.6, color='coral',
        label='N(2, 1.5)', edgecolor='white')
ax.hist(data3, bins=30, alpha=0.6, color='mediumseagreen',
        label='N(-1, 0.5)', edgecolor='white')

ax.set_title('Comparing Multiple Distributions')
ax.set_xlabel('Value')
ax.set_ylabel('Frequency')
ax.legend()
ax.grid(axis='y', alpha=0.3)
plt.show()

Example 3: Density Histogram + Fitted Curve

Example

import matplotlib.pyplot as plt
import numpy as np

np.random.seed(42)
data = np.random.randn(1000)

fig, ax = plt.subplots(figsize=(8, 5), layout='constrained')

# density=True plots probability density (instead of counts)
ax.hist(data, bins=30, density=True, alpha=0.7,
        color='steelblue', edgecolor='white', label='Data Histogram')

# Overlay theoretical normal distribution curve
from scipy import stats
x = np.linspace(-4, 4, 200)
pdf = stats.norm.pdf(x, loc=0, scale=1)
ax.plot(x, pdf, 'r-', linewidth=2, label='Theoretical N(0,1) PDF')

ax.set_title('Density Histogram with PDF Curve')
ax.set_xlabel('Value')
ax.set_ylabel('Probability Density')
ax.legend()
ax.grid(axis='y', alpha=0.3)
plt.show()

Example 4: Cumulative Histogram

Example

import matplotlib.pyplot as plt
import numpy as np

np.random.seed(42)
data = np.random.randn(500)

fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 4),
                                layout='constrained')

# Left plot: regular histogram
ax1.hist(data, bins=30, color='steelblue', edgecolor='white')
ax1.set_title('Standard Histogram')
ax1.set_xlabel('Value')
ax1.set_ylabel('Frequency')

# Right plot: cumulative histogram
ax2.hist(data, bins=30, cumulative=True, color='coral',
         edgecolor='white')
ax2.set_title('Cumulative Histogram')
ax2.set_xlabel('Value')
ax2.set_ylabel('Cumulative Frequency')

plt.show()

Example 5: Step Histogram

Example

import matplotlib.pyplot as plt
import numpy as np

np.random.seed(42)
data = np.random.exponential(scale=2, size=500)

fig, ax = plt.subplots(layout='constrained')

# histtype='stepfilled' plots a filled step histogram
ax.hist(data, bins=30, histtype='stepfilled',
        color='#3498db', alpha=0.6, edgecolor='black',
        linewidth=1, label='Stepfilled')

# Overlay step line
ax.hist(data, bins=30, histtype='step',
        color='black', linewidth=2, label='Step outline')

ax.set_title('Step Histogram (histtype)')
ax.set_xlabel('Value')
ax.set_ylabel('Frequency')
ax.legend()
ax.grid(axis='y', alpha=0.3)
plt.show()
print("example: step histogram displayed")

FAQ

How to choose the bins parameter?

Too few will lose data characteristics, too many will introduce noise.

Rules of thumb:sqrt(n)(square root rule),log2(n)+1(Sturges' rule), or simply use the default 'auto' automatic selection.

What does density=True mean?

It does not return counts; instead, it makes the total area of the histogram equal to 1, forming an estimate of the probability density. Suitable for comparison with probability density functions (PDF).


Matplotlib 参考文档Matplotlib Reference Documentation

Other Extensions