1. Product Lines
- GeForce
Aimed at gamers, offering powerful graphics processing capabilities and advanced gaming technologies.
Common ones include the NVIDIA GTX series, the high-end RTX series, and the Titan series.
- Quadro
Aimed at the professional market, such as designers, engineers, scientists, and content creators.
Common ones include the Quadro P series and the high-end Quadro RTX series.
- Tesla
Aimed at the data center and high-performance computing (HPC) market, providing powerful compute for scientific research and deep learning.
Common models include V100, A100, and so on.
- Clara
Aimed at medical imaging and the life sciences, providing AI and accelerated computing capabilities for medical image processing and life data analysis.
- Jetson
Aimed at the edge computing and robotics market, providing miniaturized, low-power AI compute modules suitable for embedded systems and robotics applications.
- Orin
Aimed at the autonomous driving and edge AI market, a power-efficient SoC (System on Chip) that integrates a CPU, GPU, and deep learning accelerator.
2. Naming Conventions
2.1 Series Names
- GeForce, the graphics card series for consumers and the gaming market. Usually used for mainstream and high-performance gaming graphics cards.
- Quadro, the professional graphics card series, aimed at fields such as graphic design, 3D rendering, and engineering applications.
- Tesla, the GPU series designed specifically for data centers, high-performance computing (HPC), and AI research.
- Titan, the high-end graphics card series, sitting between the consumer and professional markets, with both gaming and compute capability.
- RTX, graphics cards that include real-time ray tracing technology, suitable for gaming and high-performance computing.
- GTX, aimed at the mainstream and high-performance gaming market, without the ray tracing capability of the RTX series.
2.2 Architecture Codenames
- Each generation of graphics cards adopts a new architecture codename, such as Kepler, Maxwell, Pascal, Volta, Turing, Ampere, Hopper, and so on. It is usually not shown directly in the graphics card model, but the architecture can be inferred from the card’s codename or release date.
2.3 Model Numbers
- The first digit indicates the card’s generation. For example, the
10inGTX 1080indicates the 10th generation (Pascal architecture), and the30inRTX 3080indicates the 30 series (Ampere architecture). - The second digit indicates the card’s positioning or performance tier; the larger the number, the stronger the performance. For example,
RTX 3080is more powerful thanRTX 3070. - The trailing letters:
- Ti, short for “Titanium”, indicating a performance-enhanced version of that model, usually more powerful than the same-generation model without Ti.
- SUPER, indicating an upgraded version, usually with better performance and value than the base model.
- Ultra, rarely used, but sometimes used to denote a higher-performance version.
2.4 Special Models
- Founders Edition (FE), the version of the graphics card released by NVIDIA itself, usually launched early in the card’s release cycle, with a distinctive exterior design and cooling solution.
- OEM, graphics card models aimed at original equipment manufacturers (OEMs), which may have different specifications from the retail version.
2.5 Naming Examples
GeForce RTX 3090 Ti:
- GeForce, the consumer gaming graphics card series.
- RTX, the series that supports real-time ray tracing.
- 30, representing the 30-series graphics cards, based on the Ampere architecture.
- 90, the high-end model.
- Ti, the performance-enhanced version.
Quadro RTX 5000:
- Quadro, the professional graphics workstation card.
- RTX, supports real-time ray tracing.
- 5000, a mid-to-high-end professional graphics card model.
3. Hardware Compute Cores
- CUDA Core
The CUDA Core is the compute core unit on an NVIDIA GPU, used to execute general-purpose parallel computing tasks, and is the most commonly seen core type. NVIDIA usually expresses its compute capability in terms of the smallest arithmetic unit; a CUDA Core refers to a processing element that executes basic operations, and the number of CUDA Cores we refer to usually corresponds to the number of FP32 compute units.
- Tensor Core
The Tensor Core is a special compute unit introduced in NVIDIA’s Volta architecture and its successors (such as the Ampere architecture). They are dedicated to tensor computations in deep learning tasks, such as matrix multiplication and convolution operations. Tensor Cores are especially large, and are usually used in combination with deep learning frameworks (such as TensorFlow and PyTorch); they can load an entire matrix into registers for batch operations, achieving more than a tenfold efficiency improvement.
- RT Core (Ray Tracing Core)
The RT Core is NVIDIA’s dedicated hardware unit, mainly used to accelerate ray tracing computations. Normally, data-center-class GPU cores do not have RT Cores; it is mainly consumer-grade graphics cards that add RT Cores for ray tracing operations. RT Cores are mainly used in fields that require real-time rendering, such as game development, film production, and virtual reality.
4. NVIDIA General-Purpose GPU Architectures
4.1 Tesla Architecture
The Tesla architecture was released in 2006. The Tesla architecture was a brand-new CUDA architecture that supported GPU programming in the C language and could be used for general-purpose data-parallel computing. The Tesla architecture had 128 stream processors and bandwidth of up to 86GB/s, marking the point at which GPUs began to transform from dedicated graphics processors into general-purpose data-parallel processors.
Typical card models:
- Tesla C1060
- Tesla M1060
- Tesla S1070
4.2 Fermi Architecture
The Fermi architecture was released in 2008. The Fermi architecture was the first GPU architecture to adopt GPU-Direct technology; it had 32 SMs (streaming multiprocessors) and 16 PolyMorph Engine arrays, with each SM having 1 PolyMorph Engine and 64 CUDA cores. The architecture used a modular design with 4 chips and had 32 rasterization processing units and 16 texture units, paired with GDDR5 memory.
Typical card models:
- GeForce GTX 480
- GeForce GTX 470
- Quadro 6000
- Quadro 5000
- Quadro 4000
- Quadro Plex 7000
- GeForce GTX 465
4.4. Kepler Architecture
The Kepler architecture was released in 2012. It used a 28nm process and was the first GPU architecture to support supercomputing and double-precision computing. Kepler GK110 had 2880 stream processors and bandwidth of up to 288GB/s, with compute capability 3-4 times higher than the Fermi architecture. The arrival of the Kepler architecture made GPUs begin to become a focus of high-performance computing.
Typical card models:
- GeForce GTX 680
- GeForce GTX Titan
- GeForce GTX 780
- GeForce GTX 770
- GeForce GTX 760
- GeForce GTX 780 Ti
- GeForce GTX Titan Black
4.4. Maxwell Architecture
The Maxwell architecture was released in 2014 and used a 28nm process. The Maxwell architecture achieved major improvements in power efficiency and compute density: one stream processor had 128 CUDA cores, whereas Kepler had only 64. GM200 had 3072 CUDA cores and 336GB/s of bandwidth, yet its power consumption was only 225W, and its compute density was twice that of Kepler. Maxwell marked the arrival of the energy-efficient computing era for GPUs.
Typical card models:
- GeForce GTX 750 Ti
- GeForce GTX 750
- GeForce GTX 980
- GeForce GTX 970
- GeForce GTX Titan X
- NVIDIA Tegra X1
4.5. Pascal Architecture
The Pascal architecture was released in 2016. It used a 16nm FinFETPlus process, enhancing the GPU’s energy efficiency and compute density. Pascal GP100 had 3840 CUDA cores and 732GB/s of memory bandwidth, yet its power consumption was only 300W, an improvement of more than 50% over the Maxwell architecture. The Pascal architecture allowed GPUs to enter broader emerging application markets such as AI and automotive.
Typical card models
- Tesla P100
- GeForce GTX 10 Series
- Titan X (Pascal)
- Quadro GP100
- Quadro P6000
These GPUs lack low-precision hardware acceleration capability, but possess moderate single-precision compute. Because they are cheap, they are suitable for practicing the training of small models (such as Cifar10) or debugging model code.
4.6 Volta Architecture
The Volta architecture was released in 2017 and used a 12nm FinFET process. The Volta architecture added tensor cores, which could greatly accelerate the training and inference of AI and deep learning. Volta GV100 had 5120 CUDA cores and 900GB/s of bandwidth, plus 640 tensor cores, reaching an AI compute capability of 112 TFLOPS, nearly 3 times higher than the Pascal architecture. The arrival of Volta marked AI becoming a new direction for GPU development.
Typical card models:
- Tesla V100
- GeForce Titan V
- GeForce GTX 20 Series
- Quadro GV100
These GPUs carry Tensor Cores specially designed to accelerate low-precision (int8/float16) computation, but their single-precision compute shows little improvement over the previous generation. It is recommended to enable mixed-precision training in the deep learning framework to speed up model computation. Compared with single-precision training, mixed-precision training can usually provide more than 2x training speedup.
4.7. Turing Architecture
The Turing architecture was released in 2018 and used a 12nm FinFET process. The Turing architecture added Ray Tracing cores (RT Cores), which can hardware-accelerate ray tracing operations. Turing TU102 had 4608 CUDA cores, 576 tensor cores, and 72 RT cores, and supported GPU ray tracing, representing a new breakthrough in graphics technology. At the same time, the Turing architecture also brought a significant performance improvement in AI.
Typical card models:
- GeForce RTX 20 Series
- Quadro RTX 6000
- Quadro RTX 8000
- NVIDIA Turing T4
4.8. Ampere Architecture
The Ampere architecture was released in 2020. The Ampere architecture brought major improvements in compute capability, energy efficiency, and deep learning performance. Ampere-architecture GPUs use multiple streaming multiprocessors (SMs) and a wider bus width, providing more CUDA Cores and higher frequencies. It also introduced third-generation Tensor Cores, delivering stronger deep learning compute performance. Ampere-architecture GPUs also have higher memory capacity and bandwidth, making them suitable for large-scale data processing and machine learning tasks.
Typical card models:
- NVIDIA A100
- NVIDIA A800
- NVIDIA A40
- NVIDIA A16
- NVIDIA A10
- GeForce RTX 30 Series
- NVIDIA RTX A5000
- NVIDIA RTX A4000
- NVIDIA RTX A3000
- NVIDIA RTX A2000
These GPUs carry third-generation Tensor Cores. Compared with the previous generation, they support the TensorFloat32 format, which can directly accelerate single-precision training (PyTorch has it enabled by default). It is recommended to train models using the ultra-high-compute float16 half-precision mode, which achieves a more significant performance improvement than the previous generation of GPUs.
4.9. Hopper Architecture
The Hopper architecture was released in 2022. Compared with Ampere, the Hopper architecture supports fourth-generation Tensor Cores and adopts a new type of streaming processor, making each SM more capable. The Hopper architecture brings new innovations and improvements in compute capability, deep learning acceleration, and graphics features.
Typical card models:
- NVIDIA H100
- NVIDIA H200
- NVIDIA H800
- NVIDIA H20
4.10. Blackwell Architecture
The Blackwell architecture was released in 2024. Blackwell-architecture GPUs have 208 billion transistors, are manufactured on a specially customized, doubled-reticle-limit 4NP TSMC process, and connect the GPU dies into a single unified GPU through a 10 TB/s die-to-die interconnect.
Typical card models:
- NVIDIA B40
- NVIDIA B100
5. Compute Capability Table for Common GPU Cards
Within the same series, arranged from highest to lowest performance:
| Model | Memory | Single Precision (FP32) | Half Precision (FP16) | Details | Notes |
|---|---|---|---|---|---|
| RTX consumer cards | |||||
| 5090 | 32 GB | 104.8 T | 104.8 T | View | The new flagship |
| 4090 | 24GB | 82.58 T | 165.2 T | View | Apart from the drawbacks of relatively small memory and low multi-node multi-GPU parallel efficiency, the price-performance ratio is very high |
| 3090 | 24GB | 35.58 T | about 71T | View | Can be seen as a memory-expanded version of the 3080Ti. Both performance and memory size are more than sufficient, applicability is very strong, and it is the first choice for price-performance. Requires cuda11.x |
| 3080Ti | 12GB | 34.10 T | about 70T | View | A performance powerhouse; if your memory requirements are not high, it is a very suitable choice. Requires cuda11.x |
| 3060 | 12GB | 12.74 T | about 24T | View | If the 1080Ti’s memory is exactly the awkward point for you, the 3060 is a good choice, suitable for beginners. Requires cuda11.x |
| 2080Ti | 11GB | 13.45 T | 53.8 T | View | A Turing-architecture GPU; performance is decent, and among the older-generation models it is relatively well suited to mixed-precision computing. Good price-performance. |
| H-series compute cards | |||||
| H100 | 80GB | 51.22 T | 204.9 T | View | 900 GB/s bandwidth; currently cannot be bought in China |
| H800 | 80GB | 59.30 T | 237.2 TF | View | China-exclusive; NVLink bandwidth is only half of the H100’s at 450 GB/s |
| H20 | 96GB | 44 T | 148 T | View | China-exclusive; NVLink bandwidth is 900 GB/s, but the compute is relatively weak |
| A-series professional cards | |||||
| A100 | 80GB | 19.5 T | 77.97 T | View | No downsides except the price. Large memory, very well suited to half-precision compute; because of NVLink at 600 GB/s, the multi-GPU parallel speedup ratio is very high. Requires cuda11.x |
| A800 | 80GB | 19.5 T | 77.97 T | View | China-exclusive; compared with the A100, the main difference is that its NVLink speed is only 400 GB/s |
| A5000 | 24GB | 27.77 T | about 117T | View | A performance powerhouse; if you find the 3080Ti’s memory insufficient, the A5000 is a suitable choice, and its high half-precision compute makes it good for mixed precision. Requires cuda11.x |
| A40 | 48GB | 37.42 T | 149.7 T | View | Can be seen as a memory-expanded version of the 3090. Its compute is basically on par with the 3090, so choose based on memory size. Requires cuda11.x |
| A4000 | 16GB | 19.17 T | about 76T | View | Both memory and compute are fairly balanced; suitable for use during the intermediate stage. Requires cuda11.x |
| L-series compute cards | |||||
| L40S | 48GB | 91.61 T | 91.61 T | View | A new-generation data center compute card, supporting multimodal AI training |
| L20 | 48 GB | 59.35 T | 59.35 T | View | An entry-level AI inference card; memory bandwidth is limited |
| Tesla series | |||||
| T4 | 16 GB | 8.141 T | 65.13 T | View | A card dedicated to inference, supporting INT8/FP16 acceleration, with excellent energy efficiency |
| V100 | 16/32GB | 16.35 T | 125 T | View | The previous-generation professional compute flagship; high half-precision performance makes it good for mixed-precision computing |
| P40 | 24GB | 11.76 T | 11.76 T | View | A large-memory solution for older CUDA environments |
| Classic older cards | |||||
| TITAN Xp | 12GB | 12.15 T | 12.15 T | View | A relatively old Pascal-architecture GPU; suitable as an entry-level choice |
| 1080 Ti | 11GB | 11.34 T | 11.34 T | View | A card from the same era as the TITANXp, also suitable for getting started, but the 11GB of memory is occasionally awkward |
