This page looks best with JavaScript enabled

The Evolution of NVIDIA GPU Cores and Architectures

 ·  ☕ 11 min read

1. Product Lines

  • GeForce

Aimed at gamers, offering powerful graphics processing capabilities and advanced gaming technologies.

Common ones include the NVIDIA GTX series, the high-end RTX series, and the Titan series.

  • Quadro

Aimed at the professional market, such as designers, engineers, scientists, and content creators.

Common ones include the Quadro P series and the high-end Quadro RTX series.

  • Tesla

Aimed at the data center and high-performance computing (HPC) market, providing powerful compute for scientific research and deep learning.

Common models include V100, A100, and so on.

  • Clara

Aimed at medical imaging and the life sciences, providing AI and accelerated computing capabilities for medical image processing and life data analysis.

  • Jetson

Aimed at the edge computing and robotics market, providing miniaturized, low-power AI compute modules suitable for embedded systems and robotics applications.

  • Orin

Aimed at the autonomous driving and edge AI market, a power-efficient SoC (System on Chip) that integrates a CPU, GPU, and deep learning accelerator.

2. Naming Conventions

2.1 Series Names

  • GeForce, the graphics card series for consumers and the gaming market. Usually used for mainstream and high-performance gaming graphics cards.
  • Quadro, the professional graphics card series, aimed at fields such as graphic design, 3D rendering, and engineering applications.
  • Tesla, the GPU series designed specifically for data centers, high-performance computing (HPC), and AI research.
  • Titan, the high-end graphics card series, sitting between the consumer and professional markets, with both gaming and compute capability.
  • RTX, graphics cards that include real-time ray tracing technology, suitable for gaming and high-performance computing.
  • GTX, aimed at the mainstream and high-performance gaming market, without the ray tracing capability of the RTX series.

2.2 Architecture Codenames

  • Each generation of graphics cards adopts a new architecture codename, such as Kepler, Maxwell, Pascal, Volta, Turing, Ampere, Hopper, and so on. It is usually not shown directly in the graphics card model, but the architecture can be inferred from the card’s codename or release date.

2.3 Model Numbers

  • The first digit indicates the card’s generation. For example, the 10 in GTX 1080 indicates the 10th generation (Pascal architecture), and the 30 in RTX 3080 indicates the 30 series (Ampere architecture).
  • The second digit indicates the card’s positioning or performance tier; the larger the number, the stronger the performance. For example, RTX 3080 is more powerful than RTX 3070.
  • The trailing letters:
    • Ti, short for “Titanium”, indicating a performance-enhanced version of that model, usually more powerful than the same-generation model without Ti.
    • SUPER, indicating an upgraded version, usually with better performance and value than the base model.
    • Ultra, rarely used, but sometimes used to denote a higher-performance version.

2.4 Special Models

  • Founders Edition (FE), the version of the graphics card released by NVIDIA itself, usually launched early in the card’s release cycle, with a distinctive exterior design and cooling solution.
  • OEM, graphics card models aimed at original equipment manufacturers (OEMs), which may have different specifications from the retail version.

2.5 Naming Examples

  • GeForce RTX 3090 Ti:

    • GeForce, the consumer gaming graphics card series.
    • RTX, the series that supports real-time ray tracing.
    • 30, representing the 30-series graphics cards, based on the Ampere architecture.
    • 90, the high-end model.
    • Ti, the performance-enhanced version.
  • Quadro RTX 5000:

    • Quadro, the professional graphics workstation card.
    • RTX, supports real-time ray tracing.
    • 5000, a mid-to-high-end professional graphics card model.

3. Hardware Compute Cores

  • CUDA Core

The CUDA Core is the compute core unit on an NVIDIA GPU, used to execute general-purpose parallel computing tasks, and is the most commonly seen core type. NVIDIA usually expresses its compute capability in terms of the smallest arithmetic unit; a CUDA Core refers to a processing element that executes basic operations, and the number of CUDA Cores we refer to usually corresponds to the number of FP32 compute units.

  • Tensor Core

The Tensor Core is a special compute unit introduced in NVIDIA’s Volta architecture and its successors (such as the Ampere architecture). They are dedicated to tensor computations in deep learning tasks, such as matrix multiplication and convolution operations. Tensor Cores are especially large, and are usually used in combination with deep learning frameworks (such as TensorFlow and PyTorch); they can load an entire matrix into registers for batch operations, achieving more than a tenfold efficiency improvement.

  • RT Core (Ray Tracing Core)

The RT Core is NVIDIA’s dedicated hardware unit, mainly used to accelerate ray tracing computations. Normally, data-center-class GPU cores do not have RT Cores; it is mainly consumer-grade graphics cards that add RT Cores for ray tracing operations. RT Cores are mainly used in fields that require real-time rendering, such as game development, film production, and virtual reality.

4. NVIDIA General-Purpose GPU Architectures

4.1 Tesla Architecture

The Tesla architecture was released in 2006. The Tesla architecture was a brand-new CUDA architecture that supported GPU programming in the C language and could be used for general-purpose data-parallel computing. The Tesla architecture had 128 stream processors and bandwidth of up to 86GB/s, marking the point at which GPUs began to transform from dedicated graphics processors into general-purpose data-parallel processors.

Typical card models:

  • Tesla C1060
  • Tesla M1060
  • Tesla S1070

4.2 Fermi Architecture

The Fermi architecture was released in 2008. The Fermi architecture was the first GPU architecture to adopt GPU-Direct technology; it had 32 SMs (streaming multiprocessors) and 16 PolyMorph Engine arrays, with each SM having 1 PolyMorph Engine and 64 CUDA cores. The architecture used a modular design with 4 chips and had 32 rasterization processing units and 16 texture units, paired with GDDR5 memory.

Typical card models:

  • GeForce GTX 480
  • GeForce GTX 470
  • Quadro 6000
  • Quadro 5000
  • Quadro 4000
  • Quadro Plex 7000
  • GeForce GTX 465

4.4. Kepler Architecture

The Kepler architecture was released in 2012. It used a 28nm process and was the first GPU architecture to support supercomputing and double-precision computing. Kepler GK110 had 2880 stream processors and bandwidth of up to 288GB/s, with compute capability 3-4 times higher than the Fermi architecture. The arrival of the Kepler architecture made GPUs begin to become a focus of high-performance computing.

Typical card models:

  • GeForce GTX 680
  • GeForce GTX Titan
  • GeForce GTX 780
  • GeForce GTX 770
  • GeForce GTX 760
  • GeForce GTX 780 Ti
  • GeForce GTX Titan Black

4.4. Maxwell Architecture

The Maxwell architecture was released in 2014 and used a 28nm process. The Maxwell architecture achieved major improvements in power efficiency and compute density: one stream processor had 128 CUDA cores, whereas Kepler had only 64. GM200 had 3072 CUDA cores and 336GB/s of bandwidth, yet its power consumption was only 225W, and its compute density was twice that of Kepler. Maxwell marked the arrival of the energy-efficient computing era for GPUs.

Typical card models:

  • GeForce GTX 750 Ti
  • GeForce GTX 750
  • GeForce GTX 980
  • GeForce GTX 970
  • GeForce GTX Titan X
  • NVIDIA Tegra X1

4.5. Pascal Architecture

The Pascal architecture was released in 2016. It used a 16nm FinFETPlus process, enhancing the GPU’s energy efficiency and compute density. Pascal GP100 had 3840 CUDA cores and 732GB/s of memory bandwidth, yet its power consumption was only 300W, an improvement of more than 50% over the Maxwell architecture. The Pascal architecture allowed GPUs to enter broader emerging application markets such as AI and automotive.

Typical card models

  • Tesla P100
  • GeForce GTX 10 Series
  • Titan X (Pascal)
  • Quadro GP100
  • Quadro P6000

These GPUs lack low-precision hardware acceleration capability, but possess moderate single-precision compute. Because they are cheap, they are suitable for practicing the training of small models (such as Cifar10) or debugging model code.

4.6 Volta Architecture

The Volta architecture was released in 2017 and used a 12nm FinFET process. The Volta architecture added tensor cores, which could greatly accelerate the training and inference of AI and deep learning. Volta GV100 had 5120 CUDA cores and 900GB/s of bandwidth, plus 640 tensor cores, reaching an AI compute capability of 112 TFLOPS, nearly 3 times higher than the Pascal architecture. The arrival of Volta marked AI becoming a new direction for GPU development.

Typical card models:

  • Tesla V100
  • GeForce Titan V
  • GeForce GTX 20 Series
  • Quadro GV100

These GPUs carry Tensor Cores specially designed to accelerate low-precision (int8/float16) computation, but their single-precision compute shows little improvement over the previous generation. It is recommended to enable mixed-precision training in the deep learning framework to speed up model computation. Compared with single-precision training, mixed-precision training can usually provide more than 2x training speedup.

4.7. Turing Architecture

The Turing architecture was released in 2018 and used a 12nm FinFET process. The Turing architecture added Ray Tracing cores (RT Cores), which can hardware-accelerate ray tracing operations. Turing TU102 had 4608 CUDA cores, 576 tensor cores, and 72 RT cores, and supported GPU ray tracing, representing a new breakthrough in graphics technology. At the same time, the Turing architecture also brought a significant performance improvement in AI.

Typical card models:

  • GeForce RTX 20 Series
  • Quadro RTX 6000
  • Quadro RTX 8000
  • NVIDIA Turing T4

4.8. Ampere Architecture

The Ampere architecture was released in 2020. The Ampere architecture brought major improvements in compute capability, energy efficiency, and deep learning performance. Ampere-architecture GPUs use multiple streaming multiprocessors (SMs) and a wider bus width, providing more CUDA Cores and higher frequencies. It also introduced third-generation Tensor Cores, delivering stronger deep learning compute performance. Ampere-architecture GPUs also have higher memory capacity and bandwidth, making them suitable for large-scale data processing and machine learning tasks.

Typical card models:

  • NVIDIA A100
  • NVIDIA A800
  • NVIDIA A40
  • NVIDIA A16
  • NVIDIA A10
  • GeForce RTX 30 Series
  • NVIDIA RTX A5000
  • NVIDIA RTX A4000
  • NVIDIA RTX A3000
  • NVIDIA RTX A2000

These GPUs carry third-generation Tensor Cores. Compared with the previous generation, they support the TensorFloat32 format, which can directly accelerate single-precision training (PyTorch has it enabled by default). It is recommended to train models using the ultra-high-compute float16 half-precision mode, which achieves a more significant performance improvement than the previous generation of GPUs.

4.9. Hopper Architecture

The Hopper architecture was released in 2022. Compared with Ampere, the Hopper architecture supports fourth-generation Tensor Cores and adopts a new type of streaming processor, making each SM more capable. The Hopper architecture brings new innovations and improvements in compute capability, deep learning acceleration, and graphics features.

Typical card models:

  • NVIDIA H100
  • NVIDIA H200
  • NVIDIA H800
  • NVIDIA H20

4.10. Blackwell Architecture

The Blackwell architecture was released in 2024. Blackwell-architecture GPUs have 208 billion transistors, are manufactured on a specially customized, doubled-reticle-limit 4NP TSMC process, and connect the GPU dies into a single unified GPU through a 10 TB/s die-to-die interconnect.

Typical card models:

  • NVIDIA B40
  • NVIDIA B100

5. Compute Capability Table for Common GPU Cards

Within the same series, arranged from highest to lowest performance:

ModelMemorySingle Precision (FP32)Half Precision (FP16)DetailsNotes
RTX consumer cards
509032 GB104.8 T104.8 TViewThe new flagship
409024GB82.58 T165.2 TViewApart from the drawbacks of relatively small memory and low multi-node multi-GPU parallel efficiency, the price-performance ratio is very high
309024GB35.58 Tabout 71TViewCan be seen as a memory-expanded version of the 3080Ti. Both performance and memory size are more than sufficient, applicability is very strong, and it is the first choice for price-performance. Requires cuda11.x
3080Ti12GB34.10 Tabout 70TViewA performance powerhouse; if your memory requirements are not high, it is a very suitable choice. Requires cuda11.x
306012GB12.74 Tabout 24TViewIf the 1080Ti’s memory is exactly the awkward point for you, the 3060 is a good choice, suitable for beginners. Requires cuda11.x
2080Ti11GB13.45 T53.8 TViewA Turing-architecture GPU; performance is decent, and among the older-generation models it is relatively well suited to mixed-precision computing. Good price-performance.
H-series compute cards
H10080GB51.22 T204.9 TView900 GB/s bandwidth; currently cannot be bought in China
H80080GB59.30 T237.2 TFViewChina-exclusive; NVLink bandwidth is only half of the H100’s at 450 GB/s
H2096GB44 T148 TViewChina-exclusive; NVLink bandwidth is 900 GB/s, but the compute is relatively weak
A-series professional cards
A10080GB19.5 T77.97 TViewNo downsides except the price. Large memory, very well suited to half-precision compute; because of NVLink at 600 GB/s, the multi-GPU parallel speedup ratio is very high. Requires cuda11.x
A80080GB19.5 T77.97 TViewChina-exclusive; compared with the A100, the main difference is that its NVLink speed is only 400 GB/s
A500024GB27.77 Tabout 117TViewA performance powerhouse; if you find the 3080Ti’s memory insufficient, the A5000 is a suitable choice, and its high half-precision compute makes it good for mixed precision. Requires cuda11.x
A4048GB37.42 T149.7 TViewCan be seen as a memory-expanded version of the 3090. Its compute is basically on par with the 3090, so choose based on memory size. Requires cuda11.x
A400016GB19.17 Tabout 76TViewBoth memory and compute are fairly balanced; suitable for use during the intermediate stage. Requires cuda11.x
L-series compute cards
L40S48GB91.61 T91.61 TViewA new-generation data center compute card, supporting multimodal AI training
L2048 GB59.35 T59.35 TViewAn entry-level AI inference card; memory bandwidth is limited
Tesla series
T416 GB8.141 T65.13 TViewA card dedicated to inference, supporting INT8/FP16 acceleration, with excellent energy efficiency
V10016/32GB16.35 T125 TViewThe previous-generation professional compute flagship; high half-precision performance makes it good for mixed-precision computing
P4024GB11.76 T11.76 TViewA large-memory solution for older CUDA environments
Classic older cards
TITAN Xp12GB12.15 T12.15 TViewA relatively old Pascal-architecture GPU; suitable as an entry-level choice
1080 Ti11GB11.34 T11.34 TViewA card from the same era as the TITANXp, also suitable for getting started, but the 11GB of memory is occasionally awkward

微信公众号
WRITTEN BY
微信公众号