PyTorch is a free and open-source deep learning framework created by Meta AI in 2016 for training and inference on deep neural networks using Python and tensor computations with GPU acceleration. Meta open-sourced PyTorch in January 2017. It has been governed by the Linux Foundation via the PyTorch Foundation from September 2022. By 2026, PyTorch is utilized in around 85% of deep learning papers and 63% of production model training tasks, based on data from the Linux Foundation and industry usage. This guide describes what PyTorch is, how its main components function, provides code examples, and compares it to TensorFlow and JAX.
What Will I Learn?
What Is PyTorch, Exactly?
The PyTorch package enables construction and training of neural networks. There are two main aspects of the PyTorch framework: tensor computation with GPU support and automatic differentiation.
The neural network is an architecture based on layers of parameters. There are two steps involved in training a neural network – forward step when the input data passes through the architecture and the backward step that computes how to adjust the parameters to minimize errors (via gradients). PyTorch performs both natively.
PyTorch is available on three operating systems: Linux, Mac OS X, and Windows. It can be run on four types of hardware architectures: NVIDIA CUDA, AMD ROCm, Apple Metal Performance Shaders (MPS) and purely CPU. It is developed in Python, C++ and CUDA languages and is available under the BSD-3-Clause license.
PyTorch implements a define-by-run approach. The computation graph (that records all the computations required to map inputs to outputs) is constructed dynamically during the runtime, step by step. This is opposite to the define-and-run approach where the full computation graph is specified upfront. Originally, TensorFlow was implementing the define-and-run approach (version 1.x); TensorFlow 2.x has switched to define-by-run.
Who Built PyTorch — and Why It Exists
PyTorch is an offspring of Torch which is a machine learning library developed in the year 2002 using Lua programming language. It was developed by Ronan Collobert, Clément Farabet, and Koray Kavukcuoglu. It was then looked after by Soumith Chintala from 2012 onwards.
Development of PyTorch started in 2016 at FAIR (Facebook AI Research). Its original authors are Adam Paszke, Sam Gross, Soumith Chintala, and Gregory Chanan. PyTorch was developed in Python for research purposes and as an attempt to solve the issue of slow transition between research and production prevalent in previous deep learning frameworks. Meta launched its public beta version in January 2017 and PyTorch 1.0 was announced in October 2018.
The PyTorch governance was moved to the newly formed PyTorch Foundation, which is an independent project hosted by the Linux Foundation, in September 2022. The foundation’s initial members were Meta, AMD, AWS, Google, Microsoft, and NVIDIA. IBM joined the organization in 2023. In October 2023, Huawei became the foundation’s first Premier Member from China. Alibaba Cloud, Cambricon, and Ant Group became the organization’s members in 2026. Among others, the foundation is involved in open-source AI infrastructure projects such as vLLM, DeepSpeed, and Ray.
Soumith Chintala headed the engineering team of PyTorch at Meta for almost eight years until 2024. He took up the position of Chief Technology Officer at Thinking Machines Lab in January 2026.
How PyTorch Works, Step by Step
The architecture of PyTorch revolves around five elements: tensors, computational graphs, autograd, modules, and data loaders. These elements will be introduced below in the order of usage within the framework.
Tensors: PyTorch’s Basic Data Structure
A tensor can be referred to as a multi-dimensional numerical array. Tensors in PyTorch represent input, output, and model parameters of a neural network. There are four types of tensors depending on their number of dimensions:
| Tensor type | Dimensions | Example |
|---|---|---|
| Scalar | 0-D | A single number, such as a loss value: 3.14 |
| Vector | 1-D | A list of numbers, such as one row of pixel brightness values |
| Matrix | 2-D | A grid of numbers, such as a grayscale image |
| N-dimensional tensor | 3-D or higher | A color image (height × width × color channel), or a batch of images |
PyTorch tensors run on a CPU by default. Moving a tensor to a GPU with .to(“cuda”) or .to(“mps”) allows the same operations to execute in parallel across thousands of GPU cores, which reduces training time on large datasets.
Computation Graphs: Dynamic vs. Static
The computation graph illustrates the data flow through a set of operations from the start to the end point. The PyTorch uses dynamic computation graphs that create an operation only during its execution following the define-by-run paradigm.
It contrasts static graphs that create the whole sequence of operations once, compile them, and then run the graph multiple times without any modifications. TensorFlow 1.x used static computation graphs. Dynamic graphs implemented in PyTorch enable modification of the network architecture during the execution of the program depending on the data being processed. Thus, it makes debugging easier as operations are executed immediately and Python tools like print() can be used.
Autograd: How PyTorch Calculates Gradients
Autograd is the name of the automatic differentiation library in PyTorch. It determines the rate at which a loss function changes in relation to each of its model parameters.
In order to train a neural network, one must tweak the parameters such that the loss function reaches a minimal value. In PyTorch, autograd keeps track of all operations done to the tensor if requires_grad=True. Once .backward() is called on the output tensor, autograd will calculate gradients for all the tracked tensors in one step. This procedure is called backpropagation.
Modules, nn, and optim
The organization of layers in the neural networks created in PyTorch is achieved through modules, which can be built utilizing the torch.nn library. Modules can have sub-modules, and therefore, a complex network can be created out of a number of smaller modules.
There are three types of modules that should satisfy most requirements in building models:
- nn modules define the layers of a network, such as nn.Linear for a fully connected layer or nn.Conv2d for a convolutional layer.
- The autograd module computes gradients for any tensor with requires_grad=True.
- optim modules update model parameters using an optimization algorithm, such as stochastic gradient descent (SGD) or Adam.
Datasets and DataLoaders
PyTorch offers two classes for dealing with training data: torch.utils.data.Dataset, which stores data samples together with their labels, and
torch.utils.data.DataLoader, which loads and creates batches of training data.
Using these classes makes our training scripts smaller and more flexible because the data handling code is separated from the model code.
See PyTorch in Action: A 5-Line Example
The following code creates a tensor, performs a calculation, and calculates a gradient using autograd:
import torch
x = torch.tensor([1.0, 2.0, 3.0], requires_grad=True)
y = x.pow(2).sum() # y = x1^2 + x2^2 + x3^2
y.backward() # autograd computes dy/dx
print(x.grad) # tensor([2., 4., 6.])
The derivative of x² is 2x. Autograd calculates this automatically for each value in the tensor, without a manually written derivative formula. This is the same mechanism PyTorch uses to train models with millions of parameters.
PyTorch vs. TensorFlow vs. JAX in 2026
PyTorch, TensorFlow, and JAX are the three most-used deep learning frameworks in 2026. Each uses a different default execution model.
| Attribute | PyTorch | TensorFlow | JAX |
|---|---|---|---|
| Developer | Meta AI / PyTorch Foundation | Google DeepMind | |
| First released | 2016 | 2015 | 2018 |
| Default graph type | Dynamic (define-by-run) | Dynamic since 2.x (static in 1.x) | Static, via just-in-time (JIT) compilation |
| Primary interface | Python | Python (Keras API) | Python (NumPy-style API) |
| Research paper share (2025–2026) | ~85% | ~15% | Growing, concentrated in large-scale research labs |
| Production deployment tools | TorchServe, TorchScript, ONNX, ExecuTorch | TensorFlow Serving, TensorFlow Lite, TensorFlow.js | Limited native tooling; often paired with other frameworks for deployment |
| Mobile/edge support | ExecuTorch | TensorFlow Lite | Not a primary use case |
| License | BSD-3-Clause | Apache 2.0 | Apache 2.0 |
Both PyTorch and JAX use the same native Pythonic NumPy-like syntax. The Keras API from TensorFlow hides many implementation details, which means that writing code will be shorter for typical architecture designs. JAX uses XLA to compile functions before their execution, and therefore, it executes larger scale training runs faster but has a higher learning curve compared to PyTorch. TensorFlow is used more often in mobile applications via TensorFlow Lite and corporate pipelines because of the pre-PyTorch era.
What Is PyTorch Used For?
Four broad categories of AI applications are possible using PyTorch: Computer Vision, Natural Language Processing, Generative AI, and Reinforcement Learning.
Computer Vision
PyTorch algorithms are capable of classifying, detecting, and segmenting images. Both Convolutional Neural Networks (CNNs) and Vision Transformers (ViT) can be executed in PyTorch using the Torchvision library, which provides pre-trained models for functions such as image classification, object detection, and image segmentation.
Natural Language Processing and Large Language Models
PyTorch forms the backbone of Hugging Face Transformers, which is one of the most popular libraries for pre-trained language models. The architectures of the models released for functions such as text generation, text translation, and sentiment analysis are always released in PyTorch first before being converted into other formats.
Generative AI
The diffusion models, employed for generating images and videos, and large language models, employed for generating text, are trained and served on PyTorch. The dynamic nature of the graphs in the framework allows the handling of the varying input and output lengths prevalent in generative architecture.
Reinforcement Learning
The PyTorch environment enables reinforcement learning through the employment of algorithms such as Deep Q-Network (DQN), Policy Gradient algorithm, and Actor-Critic algorithms. In these algorithms, an agent learns to perform certain actions in order to maximize rewards.
Organizations That Use PyTorch
Members of the PyTorch Foundation include Meta, AWS, Google, Microsoft, NVIDIA, AMD, IBM, and Huawei. Outside of the PyTorch Foundation, the PyTorch environment is being used by Tesla, OpenAI, Microsoft, Meta AI, Google DeepMind, and academic AI groups.
The PyTorch Ecosystem
PyTorch’s core library is supplemented by several official and community-maintained tools:
- Torchvision — pre-trained models and datasets for image classification, object detection, and segmentation.
- TorchText and TorchAudio — datasets and pre-processing tools for text and audio tasks.
- ONNX (Open Neural Network Exchange) — a format for converting PyTorch models to run on other platforms and runtimes.
- torch.compile — a compiler, introduced in PyTorch 2.0, that converts standard PyTorch code into optimized kernels using an internal component called TorchDynamo, without requiring code rewrites.
- ExecuTorch — a runtime for deploying PyTorch models on mobile devices, embedded systems, and other edge hardware, without a full Python runtime.
How to Install PyTorch
PyTorch installs through pip, the standard Python package manager. The exact command depends on the target hardware.
- Confirm a supported Python version. PyTorch 2.9 and later require Python 3.10 or later.
- Choose a compute platform. Select CPU-only, or a GPU platform matching installed hardware: NVIDIA CUDA, AMD ROCm, or Apple MPS (built into macOS wheels).
- Generate the install command. The PyTorch website (pytorch.org/get-started/locally) generates the exact pip command for the selected OS, Python version, and compute platform.
- Run the install command. A CPU-only install on Linux, for example, uses: pip install torch torchvision torchaudio
- Verify the installation. Run python -c “import torch; print(torch.__version__)”. A version number confirms a successful install.
A GPU is not necessary to install or use PyTorch. GPU-less installation works fine for study, small models, and inference on small data. A GPU helps speed up the training of large data and large models since GPU cores perform multiple tensor computations at once.
PyTorch can be used on cloud computing services like AWS, Google Cloud, and Microsoft Azure, all of which have GPU support.
Is PyTorch Free to Use?
There are no restrictions whatsoever in using, modifying, and redistributing PyTorch, even for any commercial application. The BSD-3-Clause license governs the use of PyTorch, which is an open source license without any need to pay royalties or to attribute the author. There is no vendor involved in controlling access to PyTorch, nor do you have to pay anything for accessing it.
PyTorch’s Limitations
PyTorch has three documented limitations relative to alternative frameworks:
- Mobile and embedded deployment requires an additional step. PyTorch models must be exported through ExecuTorch or ONNX to run without a Python runtime, whereas TensorFlow Lite was designed for on-device deployment from the start.
- Static graph compilation is opt-in, not automatic. torch.compile must be added explicitly to gain compiled-graph performance; JAX compiles by default through its jit function.
- Dynamic graphs carry a small runtime overhead compared to fully static, pre-compiled graphs, since PyTorch’s default execution mode builds the graph during each run rather than reusing a fixed, optimized graph.
These trade-offs mean TensorFlow Lite or JAX may fit specific deployment or large-scale training scenarios better than PyTorch’s defaults, depending on the target hardware and required latency.
Frequently Asked Questions
Q1. Is PyTorch free?
Ans. Yes. PyTorch is distributed with a BSD-3-Clause license, which is a free and open-source license that allows commercial distribution without royalties.
Q2. Do I need a GPU to use PyTorch?
Ans. No. PyTorch works on a CPU by default. However, GPU is needed only when one wants to accelerate model training on huge models or datasets.
Q3. Who owns PyTorch?
Ans. PyTorch is not owned by any company. It was developed by Meta AI and released in 2017 as an open-source project. Since September 2022, PyTorch has been governed by the Linux Foundation via the independent PyTorch Foundation.
Q4. Is PyTorch better than TensorFlow?
Ans. Neither of the frameworks is better. PyTorch has a higher number of research papers and a native Pythonic syntax, whereas TensorFlow has more advanced deployment mechanisms on mobile devices using TensorFlow Lite. The choice is based on the use case.
Q5. How long does it take to learn PyTorch?
Ans. A developer with existing Python and basic linear algebra knowledge can typically write and train a simple PyTorch model within a few hours, using the tensor, autograd, and nn.Module concepts covered in this guide.
Q6. What is the difference between PyTorch and Torch?
Ans. In 2002, Torch was created using the programming language called Lua. PyTorch, which was developed in 2017, is basically an upgraded version of Torch with its key concepts rewritten with the use of the Python interface.
Summary
It consists of tensor operations powered by GPUs, automatic differentiation, and dynamic computation graphs. The PyTorch Foundation that operates under the umbrella of the Linux Foundation has been overseeing the development of this framework since 2022; its members include Meta, AWS, Google, Microsoft, NVIDIA, AMD, and IBM. All of this, together with the code sample and the instructions above, completes the necessary information on how to create a PyTorch model.