DEPLOYDatabase

AI chip · NVIDIA

NVIDIA A100 Tensor Core GPU

Ampere-generation data-center GPU (SXM4 / PCIe), 40/80 GB HBM2e. The workhorse of AI training 2020-2023 and still dominant in the deployed base. TSMC 7nm.

Market position

A100 was the GPU GPT-3, ChatGPT training, Stable Diffusion and the entire 2020-2022 model-scaling era ran on. Its 80 GB HBM2e SKU made it the first GPU where a large language model actually fit. Superseded by H100 for training; still shipping for older Kubernetes clusters and academic labs.

DEPLOY editorial. What the vendor PDF cannot tell you.

Physical-AI cross-link

Derived from claims already on record (vendor first-party TDP, throughput, price, and the reference-design cluster registry). The arithmetic is shown per row so it can be audited. No estimated inputs.

Perf per watt
0.78 FP16 TFLOPS/W
312 TFLOPS ÷ 400 W = 0.78
List $ per FP16 TFLOP
$32.05
$10,000 ÷ 312 TFLOPS = $32.05

What fits in 80 GB

Weights-only footprint for public open-weight LLMs. Pure arithmetic: params × bytes/param. Excludes KV cache and activation memory; add ~10-30% headroom for real serving. A model that does not fit at FP16 may still fit at INT8 or INT4 with quality trade-offs. Not a benchmark.

ModelParamsFP16INT8INT4
Llama 3.1 8B
dense
8 B16 GB ✓8 GB ✓4 GB ✓
Llama 3.1 70B
dense
70 B140 GB ✗70 GB ✓35 GB ✓
Llama 3.1 405B
dense
405 B810 GB ✗405 GB ✗203 GB ✗
Llama 3.3 70B
dense
70 B140 GB ✗70 GB ✓35 GB ✓
DeepSeek V3
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
DeepSeek R1
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
Qwen 2.5 7B
dense
7 B14 GB ✓7 GB ✓4 GB ✓
Qwen 2.5 72B
dense
72 B144 GB ✗72 GB ✓36 GB ✓
Mixtral 8x7B
MoE (46.7B total, 12.9B active per token)
46.7 B93 GB ✗47 GB ✓23 GB ✓
Mixtral 8x22B
MoE (141B total, 39B active per token)
141 B282 GB ✗141 GB ✗71 GB ✓
Gemma 2 27B
dense
27 B54 GB ✓27 GB ✓14 GB ✓
Command R+
dense
104 B208 GB ✗104 GB ✗52 GB ✓
Kimi K2
MoE (1T total, 32B active per token)
1000 B2000 GB ✗1000 GB ✗500 GB ✗

Math: FP16 = params × 2 bytes; INT8 = params × 1 byte; INT4 = params × 0.5 bytes. MoE models sum every expert (full weights on disk), not the per-token active subset.

Export-control status

US Bureau of Industry and Security actions that constrain this chip. Each row links a regulation record (dated, sourced, on DEPLOY) to this chip. This is one of the most-searched and least-structured topics in the whole category, and everyone else writes it as paywalled prose. Structured, dated, sourced.

  1. The US Bureau of Industry and Security's Oct 7 2022 interim final rule restricted export of advanced AI accelerators to China, keyed to a performance threshold (>=4,800 TOPS or >=600 GB/s interconnect). Directly blocked NVIDIA A100 and H100 sales to Chinese entities without a license, triggering the A800 and H800 China-specific variants NVIDIA released within weeks.

    Sources: BIS press release, Oct 7 2022 · Federal Register, Oct 13 2022 (rule text)

Common questions

Answer-first, sourced. Every claim below traces to a specific field on this page or a cited datasheet. FAQPage schema is emitted so LLM crawlers can lift these verbatim.

How much does NVIDIA A100 Tensor Core GPU cost?

List price at launch was $10,000 (2020-05-14). Street prices vary with generation age and supply. DEPLOY does not currently hold cloud rental $/hr as a first-party field.

How much memory does NVIDIA A100 Tensor Core GPU have?

80 GB of HBM2e at 2,039 GB/s. Vendor datasheet, first-party.

How much power does one NVIDIA A100 Tensor Core GPU draw?

400 W TDP (thermal design power) per chip. System-level draw is higher: see the reference-design section for rack-level kW.

Which open-weight LLMs fit on one NVIDIA A100 Tensor Core GPU?

Of the 13 public open-weight LLMs DEPLOY tracks: 3 fit at FP16, 7 at INT8, 9 at INT4 (weights only, excludes KV cache). The full table is above with per-model math. Sparsity, offloading and multi-GPU serving change the picture; this row is single-chip weights-only.

Is NVIDIA A100 Tensor Core GPU subject to US export controls?

Yes. 1 BIS action on record constrain this chip, most recently BIS Oct 7 2022 export controls on advanced AI chips to China (2022-10-07). Full list with citations above.

Key facts

Class
Compute SoC (AI accelerator)
Designer
NVIDIA
Safety-critical?
No (data-center inference / training)
Record as of
2026-08-25
Latest cited claim
2020-11-01 (across 2 inventory + benchmark rows)
Specifications (17 fields, click to expand)

Vendor datasheet figures (first-party). Dense throughput first; sparse (2:4) in parentheses where NVIDIA quotes it. The exhaustive spec sheet lives on the datasheet URL below the table: this row set covers what buyers actually compare on.

Process node
TSMC N7
Transistors
54 B
Die size
826 mm²
CUDA cores
6,912
Tensor cores
432 (3rd gen (Ampere))
TDP
400 W
Memory
80 GB HBM2e
Memory bandwidth
2,039 GB/s
PCIe
Gen 4 x16 (600 GB/s)
FP32
19.5 TFLOPS
TF32 (dense)
156 TFLOPS
FP16 (dense)
312 TFLOPS (624 sparse)
INT8 (dense)
624 TOPS (1,248 sparse)
Form factor
SXM4
Announced
2020-05-14
Released
2020-05-14
Launch price
$10,000 (list)

Source: vendor datasheet

Generation

Compare with

See every chip comparison →

Benchmarks

Published measurements per workload. Vendor datasheet numbers are first-party (verified posture); MLPerf / InferenceX / press results come in when their license terms + workload naming permit.

WorkloadValueUnitSourceAs of
memory capacity80GB HBM2evendor2020-11-01
peak tflops fp16 dense312TFLOPSvendor2020-05-01

Sources

Adoption rows appear as we verify chip-integration claims per model. Every row is a public claim the maker or a tier-1 source has stated; verification tier (verified / reported / inferred) is shown per row. See every chip on record for the full catalog or the NVIDIA record.