From a single number, to a sequence of numbers, to a table—when dimensions keep increasing, we get a tensor. It is the most universal container for describing data in the AI world.
1. Intuitive Introduction: What exactly is a tensor?
Let’s not look at formulas first; look at a photo. A small 3×3 pixel color image, each pixel has three brightness values: R, G, B. To fully describe this image, you need coordinates in three directions:Which row, which column, and which color channel。
This is a “third-order tensor” — it is not a mysterious new concept, but simply the idea of a “table” extended one more layer along a third direction. You can understand it as “multiple matrices stacked together.”
2. Mathematical Definition
2.1 Unified representation from scalar to tensor
They are essentially the same thing: an object whose values are indexed by several coordinates, only the number of indices differs.
| Name | Rank | How many coordinates are needed to locate an element | Notation |
|---|---|---|---|
| Scalar | 0 | 0 | $s$ |
| Vector | 1 | 1: the $i$-th | $v_i$ |
| Matrix | 2 | 2: the $i$-th row and $j$-th column | $M_{ij}$ |
| Tensor | $n$ | $n$ coordinates | $T_{i_1 i_2 \dots i_n}$ |
2.2 Shape
The “shape” of a tensor describes how many elements it has in each direction. For the 3×3 RGB image in this lecture:
$$ \text{shape}(T) = (3,\ 3,\ 3) \quad \text{corresponds to } (\text{row}, \text{column}, \text{channel}) $$
The total number of elements in a tensor is the product of all the numbers in its shape:
$$ \text{total number of elements} = \prod_{k=1}^{n} d_k $$
Here $d_k$ is the size of the $k$-th dimension. In the example above, the total number of elements is $3 \times 3 \times 3 = 27$.
2.3 Addition of two tensors
For tensors with the same shape, addition is performed by adding the elements at corresponding positions. This holds consistently from vectors and matrices to higher-order tensors:
$$ (A + B)_{i_1 i_2 \dots i_n} = A_{i_1 i_2 \dots i_n} + B_{i_1 i_2 \dots i_n} $$
3. Python in Practice
Use NumPy to run through the 3×3 RGB image example above and get an intuitive feel for the two concepts of “rank” and “shape.”
3.1 Create a third-order tensor
import numpy as np
# 模拟一张 3x3 像素的彩色小图,最后一维是 RGB 三个通道
image_tensor = np.random.randint(0, 256, size=(3, 3, 3))
print("阶数 (ndim):", image_tensor.ndim)
print("形状 (shape):", image_tensor.shape)
print("元素总数 (size):", image_tensor.size)
3.2 Access values by coordinates
# 取第 0 行、第 1 列、G 通道(索引 1)的像素值
pixel_g = image_tensor[0, 1, 1]
print("该位置 G 通道值:", pixel_g)
# 取第 0 行、第 1 列的完整 RGB 三个值(一个 1 阶张量,即向量)
pixel_rgb = image_tensor[0, 1, :]
print("该位置 RGB 向量:", pixel_rgb)
[0, 1, :]The result obtained from a third-order tensor is a first-order tensor (vector). This shows that “fixing a few coordinates in a higher-order tensor yields a lower-order tensor” — this rule will be used repeatedly later when discussing batch data processing in neural networks.4. 3D Visualization: “Seeing” the tensor
Below, we use Plotly to draw this 3×3×3 tensor as 27 points in space: the three coordinate axes correspond to “row, column, channel” respectively, and the color intensity of each point represents the magnitude of the value at that position. Drag the mouse to rotate the view.
Observe carefully: keep the “channel” layer fixed, and you will see a 3×3 plane—that is exactly a second-order tensor (matrix). This is also why we say “a matrix is a special case of a tensor.”
5. Connection with AI
The reason tensors are the core data structure in deep learning frameworks (PyTorch, TensorFlow) is that almost all AI data can be naturally represented as tensors:
| Data type | Typical shape | Meaning of each dimension |
|---|---|---|
| Word vectors of a piece of text | (sequence length, word vector dimension) | Which word, and which dimension of the word vector |
| A color image | (height, width, channels) | row, column, RGB channels |
| A batch of color images | (batch size, height, width, channels) | Which image, plus the three dimensions of the image itself |
| A video clip | (number of frames, height, width, channels) | Which frame, plus the three dimensions of a single-frame image |
It can be seen that “batch size”, the parameter most commonly adjusted during training, is essentially adding an extra dimension to the tensor. Once you understand tensor rank and shape, when you seeshape mismatchan error, you can immediately locate which dimension does not match.
More content:https://www.example.com/ai-math/ai-math-tutorial.html