AI chip · NVIDIA
NVIDIA A100 Tensor Core GPU
Ampere-generation data-center GPU (SXM4 / PCIe), 40/80 GB HBM2e. The workhorse of AI training 2020-2023 and still dominant in the deployed base. TSMC 7nm.
Market position
A100 was the GPU GPT-3, ChatGPT training, Stable Diffusion and the entire 2020-2022 model-scaling era ran on. Its 80 GB HBM2e SKU made it the first GPU where a large language model actually fit. Superseded by H100 for training; still shipping for older Kubernetes clusters and academic labs.
DEPLOY editorial. What the vendor PDF cannot tell you.
Physical-AI cross-link
Derived from claims already on record (vendor first-party TDP, throughput, price, and the reference-design cluster registry). The arithmetic is shown per row so it can be audited. No estimated inputs.
- Perf per watt
- 0.78 FP16 TFLOPS/W312 TFLOPS ÷ 400 W = 0.78
- List $ per FP16 TFLOP
- $32.05$10,000 ÷ 312 TFLOPS = $32.05
What fits in 80 GB
Weights-only footprint for public open-weight LLMs. Pure arithmetic: params × bytes/param. Excludes KV cache and activation memory; add ~10-30% headroom for real serving. A model that does not fit at FP16 may still fit at INT8 or INT4 with quality trade-offs. Not a benchmark.
| Model | Params | FP16 | INT8 | INT4 |
|---|---|---|---|---|
| Llama 3.1 8B dense | 8 B | 16 GB ✓ | 8 GB ✓ | 4 GB ✓ |
| Llama 3.1 70B dense | 70 B | 140 GB ✗ | 70 GB ✓ | 35 GB ✓ |
| Llama 3.1 405B dense | 405 B | 810 GB ✗ | 405 GB ✗ | 203 GB ✗ |
| Llama 3.3 70B dense | 70 B | 140 GB ✗ | 70 GB ✓ | 35 GB ✓ |
| DeepSeek V3 MoE (671B total, 37B active per token) | 671 B | 1342 GB ✗ | 671 GB ✗ | 336 GB ✗ |
| DeepSeek R1 MoE (671B total, 37B active per token) | 671 B | 1342 GB ✗ | 671 GB ✗ | 336 GB ✗ |
| Qwen 2.5 7B dense | 7 B | 14 GB ✓ | 7 GB ✓ | 4 GB ✓ |
| Qwen 2.5 72B dense | 72 B | 144 GB ✗ | 72 GB ✓ | 36 GB ✓ |
| Mixtral 8x7B MoE (46.7B total, 12.9B active per token) | 46.7 B | 93 GB ✗ | 47 GB ✓ | 23 GB ✓ |
| Mixtral 8x22B MoE (141B total, 39B active per token) | 141 B | 282 GB ✗ | 141 GB ✗ | 71 GB ✓ |
| Gemma 2 27B dense | 27 B | 54 GB ✓ | 27 GB ✓ | 14 GB ✓ |
| Command R+ dense | 104 B | 208 GB ✗ | 104 GB ✗ | 52 GB ✓ |
| Kimi K2 MoE (1T total, 32B active per token) | 1000 B | 2000 GB ✗ | 1000 GB ✗ | 500 GB ✗ |
Math: FP16 = params × 2 bytes; INT8 = params × 1 byte; INT4 = params × 0.5 bytes. MoE models sum every expert (full weights on disk), not the per-token active subset.
Export-control status
US Bureau of Industry and Security actions that constrain this chip. Each row links a regulation record (dated, sourced, on DEPLOY) to this chip. This is one of the most-searched and least-structured topics in the whole category, and everyone else writes it as paywalled prose. Structured, dated, sourced.
The US Bureau of Industry and Security's Oct 7 2022 interim final rule restricted export of advanced AI accelerators to China, keyed to a performance threshold (>=4,800 TOPS or >=600 GB/s interconnect). Directly blocked NVIDIA A100 and H100 sales to Chinese entities without a license, triggering the A800 and H800 China-specific variants NVIDIA released within weeks.
Sources: BIS press release, Oct 7 2022 · Federal Register, Oct 13 2022 (rule text)
Common questions
Answer-first, sourced. Every claim below traces to a specific field on this page or a cited datasheet. FAQPage schema is emitted so LLM crawlers can lift these verbatim.
How much does NVIDIA A100 Tensor Core GPU cost?
List price at launch was $10,000 (2020-05-14). Street prices vary with generation age and supply. DEPLOY does not currently hold cloud rental $/hr as a first-party field.
How much memory does NVIDIA A100 Tensor Core GPU have?
80 GB of HBM2e at 2,039 GB/s. Vendor datasheet, first-party.
How much power does one NVIDIA A100 Tensor Core GPU draw?
400 W TDP (thermal design power) per chip. System-level draw is higher: see the reference-design section for rack-level kW.
Which open-weight LLMs fit on one NVIDIA A100 Tensor Core GPU?
Of the 13 public open-weight LLMs DEPLOY tracks: 3 fit at FP16, 7 at INT8, 9 at INT4 (weights only, excludes KV cache). The full table is above with per-model math. Sparsity, offloading and multi-GPU serving change the picture; this row is single-chip weights-only.
Is NVIDIA A100 Tensor Core GPU subject to US export controls?
Yes. 1 BIS action on record constrain this chip, most recently BIS Oct 7 2022 export controls on advanced AI chips to China (2022-10-07). Full list with citations above.
Key facts
- Class
- Compute SoC (AI accelerator)
- Designer
- NVIDIA
- Safety-critical?
- No (data-center inference / training)
- Record as of
- 2026-08-25
- Latest cited claim
- 2020-11-01 (across 2 inventory + benchmark rows)
Specifications (17 fields, click to expand)
Vendor datasheet figures (first-party). Dense throughput first; sparse (2:4) in parentheses where NVIDIA quotes it. The exhaustive spec sheet lives on the datasheet URL below the table: this row set covers what buyers actually compare on.
- Process node
- TSMC N7
- Transistors
- 54 B
- Die size
- 826 mm²
- CUDA cores
- 6,912
- Tensor cores
- 432 (3rd gen (Ampere))
- TDP
- 400 W
- Memory
- 80 GB HBM2e
- Memory bandwidth
- 2,039 GB/s
- PCIe
- Gen 4 x16 (600 GB/s)
- FP32
- 19.5 TFLOPS
- TF32 (dense)
- 156 TFLOPS
- FP16 (dense)
- 312 TFLOPS (624 sparse)
- INT8 (dense)
- 624 TOPS (1,248 sparse)
- Form factor
- SXM4
- Announced
- 2020-05-14
- Released
- 2020-05-14
- Launch price
- $10,000 (list)
Source: vendor datasheet
Generation
Compare with
Benchmarks
Published measurements per workload. Vendor datasheet numbers are first-party (verified posture); MLPerf / InferenceX / press results come in when their license terms + workload naming permit.
Sources
- NVIDIA A100Nvidia
Adoption rows appear as we verify chip-integration claims per model. Every row is a public claim the maker or a tier-1 source has stated; verification tier (verified / reported / inferred) is shown per row. See every chip on record for the full catalog or the NVIDIA record.