NumPy Array Attributes

In this chapter, we will learn some basic attributes of NumPy arrays.

The number of dimensions of a NumPy array is called its rank. The rank is the number of axes, that is, the number of dimensions of the array.

A one-dimensional array has rank 1, a two-dimensional array has rank 2, and so on.

In NumPy, each dimension of an array is called an axis, which is also a dimension.

For example, a two-dimensional array can be regarded as a "one-dimensional array composed of one-dimensional arrays": each element of the outer array is itself a one-dimensional array.

Here, axis 0 corresponds to the direction of the outermost array, axis 1 corresponds to the direction of the inner array, and so on.

The number of axes—the rank—is the number of dimensions of the array.

Many NumPy functions support the axis parameter.

For a two-dimensional array, axis=0 means operating along axis 0, that is, operating on each column; axis=1 means operating along axis 1, that is, operating on each row.

Below, taking a two-dimensional array with shape (2, 3) as an example, we visually show the directions of the two axes and the meanings of the attributes:

axis = 1: horizontal arrow axis = 1 (axis 1) Horizontal direction, along each row axis = 0: vertical arrow axis = 0 (axis 0) Vertical direction, along each column Column index Column 0 Column 1 Column 2 Row index Row 0 Row 1 2×3 grid 1 2 3 4 5 6 Dashed-line association between rows and the shape card shape description card shape = (2, 3) 2 → length of axis 0 (number of rows) 3 → length of axis 1 (number of columns) Bottom attribute summary ndim = 2 (number of axes, i.e., rank) · size = 2 × 3 = 6 (total number of elements)

axis indicates "which direction to compress along", not "which row/column": summing along axis=0 compresses multiple rows into one row, so the result is the sum of each column.

Example

import numpy as np

a = np.array([[1, 2, 3], [4, 5, 6]])

# Sum along axis 0: compress across rows, sum each column
print(a.sum(axis=0))

# Sum along axis 1: compress across columns, sum each row
print(a.sum(axis=1))

Output result:

[5 7 9]
[ 6 15]

The interactive demonstration below can more intuitively show the summing direction of the two axes. Click the buttons to switch:

↓↓↓
1
2
3
4
5
6
5
7
9
→
6
→
15
a.sum(axis=0) → [5 7 9], compress the elements of each column along axis 0 into one number

The more important ndarray object attributes in NumPy arrays include:

AttributeDescription
ndarray.ndimThe rank of the array, i.e., the number of dimensions or the number of axes of the array.
ndarray.shapeThe dimensions of the array, representing the size of the array on each axis. For a two-dimensional array (matrix), it represents the number of rows and columns.
ndarray.sizeThe total number of elements in the array, equal tondarray.shapethe product of the sizes of each axis in shape.
ndarray.dtypeThe data type of the elements in the array.
ndarray.itemsizeThe size of each element in the array, in bytes.
ndarray.flagsContains information about the memory layout, such as whether it is C or Fortran contiguous storage, whether it is read-only, etc.
ndarray.realThe real part of each element in the array (if the element type is complex).
ndarray.imagThe imaginary part of each element in the array (if the element type is complex).
ndarray.dataThe buffer that actually stores the array elements, usually accessed by indexing, and this attribute is not used directly.

ndarray.ndim

ndarray.ndim is used to get the number of dimensions of an array (i.e., the number of axes), which is the rank.

Example

import numpy as np

a = np.arange(24)
print(a.ndim)             # a now has only one dimension

# Use reshape to adjust its size, turning it into 2 pages, 4 rows, and 3 columns
b = a.reshape(2, 4, 3)    # b now has three dimensions
print(b.ndim)

Output result:

1
3

ndarray.shape

ndarray.shape represents the dimensions of the array and returns a tuple. The length of this tuple is the number of dimensions, i.e., the ndim property (rank).

For example, the shape of a two-dimensional array represents its "number of rows" and "number of columns".

The order of the shape tuple is consistent with the order of the axes: the 0th element is the size of axis 0 (number of rows), and the 1st element is the size of axis 1 (number of columns).

Example

import numpy as np

a = np.array([[1, 2, 3], [4, 5, 6]])
print(a.shape)

Output result:

(2, 3)

ndarray.shape can also be used to adjust the size of the array.

Example

import numpy as np

a = np.array([[1, 2, 3], [4, 5, 6]])
a.shape = (3, 2)
print(a)

Output result:

[[1 2]
 [3 4]
 [5 6]]

NumPy also provides the reshape function to adjust the size of arrays.

Example

import numpy as np

a = np.array([[1, 2, 3], [4, 5, 6]])
b = a.reshape(3, 2)
print(b)

Output result:

[[1 2]
 [3 4]
 [5 6]]

The difference between the two methods: directly assigning a value to a.shape modifies the array a itself in place; while a.reshape(3, 2) returns a new array object (usually sharing the same data with a), and a itself remains unchanged.


ndarray.size

ndarray.size returns the total number of elements in the array, equal to the product of the sizes of each axis in shape.

Example

import numpy as np

a = np.array([[1, 2, 3], [4, 5, 6]])
print(a.shape)   # shape is (2, 3)
print(a.size)    # Total number of elements: 2 × 3 = 6

Output result:

(2, 3)
6

ndarray.dtype

ndarray.dtype returns the data type of the elements in the array.

The dtype object contains information such as the type name and the element bit width. For example, int64 represents a 64-bit integer, and float32 represents a 32-bit floating-point number.

Example

import numpy as np

# Integer lists create an int64 array by default (int32 by default on Windows)
a = np.array([1, 2, 3])
print(a.dtype)

# When creating, you can explicitly specify the data type with the dtype parameter
b = np.array([1, 2, 3], dtype=np.float32)
print(b.dtype)

Output result:

int64
float32

ndarray.real and ndarray.imag

ndarray.real and ndarray.imag return the real part and the imaginary part of each element in the array, respectively, mainly used when the element type is complex.

Example

import numpy as np

# Create a complex array, the default type is complex128
a = np.array([1+2j, 3+4j, 5+6j])
print(a.real)   # The real part of each element
print(a.imag)   # The imaginary part of each element

Output result:

[1. 3. 5.]
[2. 4. 6.]

ndarray.itemsize

ndarray.itemsize returns the size of each element in the array in bytes.

For example, for an array with element type float64, the value of the itemsize attribute is 8 (float64 occupies 64 bits, and every 8 bits is 1 byte, so 64 ÷ 8 = 8 bytes).

For another example, for an array with element type int32, the itemsize attribute value is 4 (32 ÷ 8).

The product of size and itemsize is the total number of bytes occupied by the array data, which is the value of the ndarray.nbytes attribute.

Example

import numpy as np

# The dtype of the array is int8, each element occupies 1 byte
x = np.array([1, 2, 3, 4, 5], dtype=np.int8)
print(x.itemsize)

# The dtype of the array is float64, each element occupies 8 bytes
y = np.array([1, 2, 3, 4, 5], dtype=np.float64)
print(y.itemsize)

Output result:

1
8

ndarray.flags

ndarray.flags returns the memory information of the ndarray object, including the following attributes:

AttributeDescription
C_CONTIGUOUS (C)The data is in a single, C-style contiguous memory region.
F_CONTIGUOUS (F)The data is in a single, Fortran-style contiguous memory region.
OWNDATA (O)The array owns the memory it uses, rather than borrowing it from other objects.
WRITEABLE (W)The data area can be written; setting this value to False makes the data read-only.
ALIGNED (A)The data and all elements are properly aligned to the hardware.
WRITEBACKIFCOPY (X)This array is a copy of another array; when this array is released, its content will be written back to the original array.

The legacy UPDATEIFCOPY (U) is the old name of WRITEBACKIFCOPY. It was deprecated in NumPy 1.14 and completely removed starting from NumPy 2.0, so it should no longer be used now.

C_CONTIGUOUS and F_CONTIGUOUS describe the order in which the same data is arranged in memory. For the same two-dimensional array, there are two typical storage methods:

Left: C language style (row-major) C language style (row-major) Store row 0 completely, then row 1 123 456 Storage order indices: by row 1~6 1 2 3 4 5 6 Expand sequentially by row Memory: 1 2 3 4 5 6 123 456 1 2 3 4 5 6 012 345 Memory addresses (left to right) C_CONTIGUOUS : True Right: Fortran language style (column-major) Fortran language style (column-major) Store column 0 completely, then column 1 123 456 Storage order indices: by column 1~6 1 2 3 4 5 6 Expand sequentially by column Memory: 1 4 2 5 3 6 142 536 1 2 3 4 5 6 012 345 Memory addresses (left to right) F_CONTIGUOUS : True Bottom description NumPy arrays are stored in C language style (row-major) by default

Example

import numpy as np

x = np.array([1, 2, 3, 4, 5])
print(x.flags)

The output result is:

  C_CONTIGUOUS : True
  F_CONTIGUOUS : True
  OWNDATA : True
  WRITEABLE : True
  ALIGNED : True
  WRITEBACKIFCOPY : False
Other extensions