Matplotlib hist() Function
Matplotlib Reference Documentation
hist()Used to plot a histogram, dividing data into several intervals (bins) and counting the frequency of data occurrences in each interval.
Histograms are basic tools for exploring data distribution (central tendency, dispersion, skewness, etc.).
Function Definition
pyplot Interface
matplotlib.pyplot.hist(x, bins=None, range=None, density=False,
weights=None, cumulative=False, bottom=None, histtype='bar',
align='mid', orientation='vertical', rwidth=None, log=False,
color=None, label=None, stacked=False, **kwargs)
Axes Interface
Axes.hist(x, bins=None, range=None, density=False, weights=None,
cumulative=False, bottom=None, histtype='bar', align='mid',
orientation='vertical', rwidth=None, log=False, color=None,
label=None, stacked=False, **kwargs)
Parameter Description
| Parameter | Type | Description |
|---|---|---|
| x | array or sequence of arrays | Input data; multiple datasets can be passed in. |
| bins | int or sequence | Number of bins or sequence of bin edges. By default, the 'auto' method is used to automatically select. |
| range | tuple | Data range (lower, upper); data outside the range is ignored. |
| density | bool | If True, plot a probability density histogram (total area equals 1) instead of counts. |
| weights | array-like | Weight of each data point. |
| cumulative | bool or -1 | If True, plot a cumulative histogram; -1 means cumulative from largest to smallest. |
| histtype | str | Histogram type: 'bar' (default), 'barstacked' (stacked), 'step' (step line), 'stepfilled' (filled step). |
| align | str | Bin alignment: 'mid' (centered, default), 'left' (left-aligned), 'right' (right-aligned). |
| orientation | str | 'vertical' (vertical, default) or 'horizontal' (horizontal) |
| rwidth | float | Relative width of bars; 1.0 means no gap. |
| log | bool | If True, the y-axis uses logarithmic scale. |
| color | color or list | Bar color |
| label | str or list | Legend label |
| stacked | bool | Whether to display multiple datasets as stacked. |
The return value of hist() is a tuple of three items:
(n, bins, patches)n is the count/density of each bin, bins are the bin edges, and patches are the bar objects drawn.
Usage Examples
Example 1: Basic Histogram
Example
import numpy as np
# Generate normally distributed random data
np.random.seed(42)
data = np.random.randn(1000)
fig, ax = plt.subplots(layout='constrained')
# Draw histogram
n, bins, patches = ax.hist(data, bins=30,
color='steelblue', edgecolor='white')
# Highlight the bin containing the maximum value
max_idx = np.argmax(n)
patches[max_idx].set_facecolor('#e74c3c')
ax.set_title('Histogram of Normal Distribution')
ax.set_xlabel('Value')
ax.set_ylabel('Frequency')
ax.axvline(x=0, color='red', linestyle='--', alpha=0.7)
ax.grid(axis='y', alpha=0.3)
plt.show()
Example 2: Comparison Histogram of Multiple Data Sets
Example
import numpy as np
np.random.seed(42)
# Three sets of data with different distributions
data1 = np.random.normal(0, 1, 1000) # Standard normal
data2 = np.random.normal(2, 1.5, 800) # mean=2, std=1.5
data3 = np.random.normal(-1, 0.5, 600) # mean=-1, std=0.5
fig, ax = plt.subplots(figsize=(8, 5), layout='constrained')
# Compare multiple datasets using transparency and different colors
ax.hist(data1, bins=30, alpha=0.6, color='steelblue',
label='N(0, 1)', edgecolor='white')
ax.hist(data2, bins=30, alpha=0.6, color='coral',
label='N(2, 1.5)', edgecolor='white')
ax.hist(data3, bins=30, alpha=0.6, color='mediumseagreen',
label='N(-1, 0.5)', edgecolor='white')
ax.set_title('Comparing Multiple Distributions')
ax.set_xlabel('Value')
ax.set_ylabel('Frequency')
ax.legend()
ax.grid(axis='y', alpha=0.3)
plt.show()
Example 3: Density Histogram + Fitted Curve
Example
import numpy as np
np.random.seed(42)
data = np.random.randn(1000)
fig, ax = plt.subplots(figsize=(8, 5), layout='constrained')
# density=True plots probability density (instead of counts)
ax.hist(data, bins=30, density=True, alpha=0.7,
color='steelblue', edgecolor='white', label='Data Histogram')
# Overlay theoretical normal distribution curve
from scipy import stats
x = np.linspace(-4, 4, 200)
pdf = stats.norm.pdf(x, loc=0, scale=1)
ax.plot(x, pdf, 'r-', linewidth=2, label='Theoretical N(0,1) PDF')
ax.set_title('Density Histogram with PDF Curve')
ax.set_xlabel('Value')
ax.set_ylabel('Probability Density')
ax.legend()
ax.grid(axis='y', alpha=0.3)
plt.show()
Example 4: Cumulative Histogram
Example
import numpy as np
np.random.seed(42)
data = np.random.randn(500)
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 4),
layout='constrained')
# Left plot: regular histogram
ax1.hist(data, bins=30, color='steelblue', edgecolor='white')
ax1.set_title('Standard Histogram')
ax1.set_xlabel('Value')
ax1.set_ylabel('Frequency')
# Right plot: cumulative histogram
ax2.hist(data, bins=30, cumulative=True, color='coral',
edgecolor='white')
ax2.set_title('Cumulative Histogram')
ax2.set_xlabel('Value')
ax2.set_ylabel('Cumulative Frequency')
plt.show()
Example 5: Step Histogram
Example
import numpy as np
np.random.seed(42)
data = np.random.exponential(scale=2, size=500)
fig, ax = plt.subplots(layout='constrained')
# histtype='stepfilled' plots a filled step histogram
ax.hist(data, bins=30, histtype='stepfilled',
color='#3498db', alpha=0.6, edgecolor='black',
linewidth=1, label='Stepfilled')
# Overlay step line
ax.hist(data, bins=30, histtype='step',
color='black', linewidth=2, label='Step outline')
ax.set_title('Step Histogram (histtype)')
ax.set_xlabel('Value')
ax.set_ylabel('Frequency')
ax.legend()
ax.grid(axis='y', alpha=0.3)
plt.show()
print("example: step histogram displayed")
FAQ
How to choose the bins parameter?
Too few will lose data characteristics, too many will introduce noise.
Rules of thumb:sqrt(n)(square root rule),log2(n)+1(Sturges' rule), or simply use the default 'auto' automatic selection.
What does density=True mean?
It does not return counts; instead, it makes the total area of the histogram equal to 1, forming an estimate of the probability density. Suitable for comparison with probability density functions (PDF).
Other Extensions