PyTorch Introduction

PyTorch is an open-source Python machine learning library based on the Torch library, implemented at the underlying level in C++, and applied in artificial intelligence fields such as computer vision and natural language processing.

PyTorch was originally developed by Meta Platforms' AI research team and is now part of the Linux Foundation.

Many deep learning software packages are built on PyTorch, including Tesla Autopilot, Uber's Pyro, Hugging Face's Transformers, PyTorch Lightning, and Catalyst.

PyTorch has two main features:

  • Tensor computation similar to NumPy, which can be accelerated on hardware accelerators such as GPU or MPS.
  • Deep neural networks based on an automatic differentiation system.

PyTorch includes submodules such as torch.autograd, torch.nn, and torch.optim.

PyTorch includes a variety of loss functions, including MSE (mean squared error = L2 norm), cross-entropy loss, and negative log-likelihood loss (useful for classifiers), among others.

PyTorch Features

  • Dynamic Computation Graphs: PyTorch's computation graphs are dynamic, meaning they are built at runtime and can be changed at any time. This provides great flexibility for experimentation and debugging, because developers can execute code line by line and inspect intermediate results.

  • Automatic Differentiation: PyTorch's automatic differentiation system allows developers to easily compute gradients, which is crucial for training deep learning models. It automatically calculates the gradients of the loss function with respect to model parameters through the backpropagation algorithm.

  • Tensor Computation: PyTorch provides tensor operations similar to NumPy, which can be executed on CPU and GPU, thereby accelerating the computation process. Tensors are the basic data structure in PyTorch, used to store and manipulate data.

  • Rich API: PyTorch provides a large number of predefined layers, loss functions, and optimization algorithms, all of which are common components for building deep learning models.

  • Multi-language Support: Although PyTorch uses Python as its main interface, it also provides a C++ interface, allowing lower-level integration and control.

Dynamic Computation Graph

One of the most notable features of PyTorch is its dynamic computation graph mechanism.

Unlike TensorFlow's static computation graph, PyTorch builds the computation graph during execution, meaning that with each computation, the graph automatically changes according to the shape of the input data.

Advantages of the dynamic computation graph:

  • More flexible, especially suitable for scenarios requiring conditional judgment or recursion.
  • Convenient for debugging and modification, allowing direct inspection of intermediate results.
  • Closer to Python programming style and easy to get started.

Tensor and Autograd

The core data structure in PyTorch is theTensor, which is a multi-dimensional matrix that can efficiently perform computations on CPU or GPU. Tensor operations support the Autograd mechanism, enabling automatic gradient computation during backpropagation, which is crucial for gradient descent optimization algorithms in deep learning.

Tensor:

  • Supports switching between CPU and GPU.
  • Provides a NumPy-like interface and supports element-wise operations.
  • Supports automatic differentiation, making gradient computation convenient.

Autograd:

  • PyTorch's built-in automatic differentiation engine can automatically track all tensor operations and compute gradients during backpropagation.
  • Throughrequires_gradattribute, you can specify that a tensor requires gradient computation.
  • Supports efficient backpropagation, suitable for neural network training.

Model Definition and Training

PyTorch providestorch.nnmodule, allowing users to inheritnn.Moduleclass to define neural network models. Useforwardfunction to specify forward propagation; automatic backpropagation (throughautograd) and gradient computation are also handled internally by PyTorch.

Neural network module (torch.nn):

  • Provides commonly used layers (such as linear layers, convolutional layers, pooling layers, etc.).
  • Supports defining complex neural network architectures (including networks with multiple inputs and outputs).
  • Compatible with optimizers (such astorch.optim) for use together.

GPU Acceleration

PyTorch fully supports running on GPU to accelerate the training of deep learning models. Through a simple.to(device)method, users can transfer models and tensors to the GPU for computation. PyTorch supports multi-GPU training and can leverage NVIDIA CUDA technology to significantly improve computational efficiency.

GPU support:

  • Automatically selects GPU or CPU.
  • Supports acceleration via CUDA.
  • Supports multi-GPU parallel computing (DataParallelortorch.distributed)。

Ecosystem and Community Support

As an open-source project, PyTorch has a large community and ecosystem. It is not only widely used in academia, but also widely deployed in industry, especially in fields such as computer vision and natural language processing. PyTorch also provides many tools and libraries related to deep learning, such as:

  • torchvision: datasets and models for computer vision tasks.
  • torchtext: datasets and preprocessing tools for natural language processing tasks.
  • torchaudio: a toolkit for audio processing.
  • PyTorch Lightning: a high-level library that simplifies PyTorch code, focusing on rapid iteration in research and experimentation.

Comparison with Other Frameworks

Due to its flexibility, ease of use, and community support, PyTorch has become the preferred framework for many deep learning researchers and developers.

TensorFlow vs PyTorch

  • PyTorch's dynamic computation graph makes it more flexible and suitable for rapid experimentation and research, while TensorFlow's static computation graph has more room for optimization in production environments.
  • PyTorch is more convenient for debugging, while TensorFlow is more mature in deployment and supports a wider range of hardware and platforms.
  • In recent years, TensorFlow has also introduced dynamic graphs (such as TensorFlow 2.x), making the two increasingly close in functionality.
  • Other deep learning frameworks, such as Keras and Caffe, also have certain applications, but due to its flexibility, ease of use, and community support, PyTorch has become the preferred framework for many deep learning researchers and developers.
FeatureTensorFlowPyTorch
DeveloperGoogleFacebook (FAIR)
Computation graph typeStatic computation graph (defined before execution)Dynamic computation graph (executed upon definition)
FlexibilityLow (computation graph is built at compile time and not easy to modify)High (computation graph is dynamically created at execution time, easy to modify and debug)
DebuggingDifficult (requires usingtf.debuggingor external tools for debugging)Easy (can directly debug in Python)
Ease of useLow (more complex, more APIs, steeper learning curve)High (concise API, syntax closer to Python, easy to get started)
DeploymentStrong (supports a wide range of hardware, such as TensorFlow Lite and TensorFlow.js)Weaker (relatively few deployment tools and platforms, although there is TensorFlow support)
Community supportVery strong (mature and large community, extensive tutorials and documentation)Very strong (active community, especially in academia, rapidly developing ecosystem)
Model trainingSupports distributed training and supports multiple devices (such as CPU, GPU, TPU)Supports distributed training, supports multi-GPU, CPU, and TPU
API levelHigh-level API: Keras; low-level API: TensorFlow CoreHigh-level APIs: TorchVision, TorchText, etc.; low-level API: Torch
PerformanceHigh (mature in optimization, suitable for production environments)High (suitable for research and prototyping, production performance is also improving)
Automatic differentiationSupportedtf.GradientTapeImplements dynamic differentiation (more complex)SupportedautogradDynamic differentiation (more concise and intuitive)
Tuning and scalabilityStrong (supports running on multiple platforms, such as TensorFlow Serving, etc.)Weaker (although it excels in academic and experimental environments, production environment support is relatively limited)
Framework flexibilityLower (TensorFlow 2.x introduced dynamic graph features, but it is still not fully flexible)High (dynamic graph brings higher flexibility)
Supports multiple languagesSupports multiple languages (Python, C++, Java, JavaScript, etc.)Mainly supports Python (but also has a C++ API)
Compatibility and migrationTensorFlow 2.x has good compatibility with older versionsPoor compatibility with TensorFlow, migration is difficult

PyTorch vs NumPy

Features PyTorch NumPy
Goal Dedicated to deep learning General scientific computing
GPU support Natively supports CUDA Not directly supported
Automatic differentiation Built-in automatic differentiation Requires manual gradient computation
Neural networks Rich set of neural network modules Need to implement from scratch
Learning cost Relatively high Relatively low

History and Development of PyTorch

PyTorch's predecessor was Torch, a scientific computing framework based on the Lua language. With the growing popularity of Python in machine learning, the Facebook team decided to port Torch's core ideas to Python, giving birth to PyTorch.

  • 2016: Facebook released PyTorch version 0.1
  • 2017: PyTorch 0.2 introduced distributed training support
  • 2018: PyTorch 1.0 was released, adding production deployment capabilities
  • 2019: PyTorch 1.3 introduced mobile support
  • 2020: PyTorch 1.6 added automatic mixed-precision training
  • 2021: PyTorch 1.9 introduced TorchScript and C++ frontend
  • 2022: PyTorch 1.12 optimized performance and stability
  • 2023: PyTorch 2.0 was released, introducing a compilation mode that greatly improves performance
Other extensions