AI chip · NVIDIA
NVIDIA H100 Tensor Core GPU
Hopper-architecture data-center AI GPU (SXM5 / PCIe). Dominant training + inference accelerator 2023-2024; 80 GB HBM3, ~700W TDP. Fabricated by TSMC on the 4N process.
On the DEPLOY graph: 7 campuses carry a cited inventory line for it · 2 reference-design clusters include it.
Market position
H100 was the chip the 2023-2024 AI-training build-out ran on. Meta bought 350,000. Microsoft, xAI, Anthropic, OpenAI, Oracle, CoreWeave and every neocloud filled DCs with it. Its NVLink 4 + HBM3 combo set the training-cluster template through 2024 until H200 refreshed the memory and Blackwell (B200/GB200) took the crown.
DEPLOY editorial. What the vendor PDF cannot tell you.
Physical-AI cross-link
Derived from claims already on record (vendor first-party TDP, throughput, price, and the reference-design cluster registry). The arithmetic is shown per row so it can be audited. No estimated inputs.
- Perf per watt
- 2.83 FP8 TFLOPS/W1,979 TFLOPS ÷ 700 W = 2.83
- List $ per FP8 TFLOP
- $15.16$30,000 ÷ 1,979 TFLOPS = $15.16
- Rack power (in this design)
- 10.2 kW per NVIDIA HGX H100 (8-GPU baseboard) (~1275 W per chip incl. overhead)10.2 kW ÷ 8 chips = 1275 W each (system-level, includes CPU/mem/NIC/cooling)
- Cooling class (across reference designs)
- airObserved across 2 reference-design clusters containing this chip
What fits in 80 GB
Weights-only footprint for public open-weight LLMs. Pure arithmetic: params × bytes/param. Excludes KV cache and activation memory; add ~10-30% headroom for real serving. A model that does not fit at FP16 may still fit at INT8 or INT4 with quality trade-offs. Not a benchmark.
| Model | Params | FP16 | INT8 | INT4 |
|---|---|---|---|---|
| Llama 3.1 8B dense | 8 B | 16 GB ✓ | 8 GB ✓ | 4 GB ✓ |
| Llama 3.1 70B dense | 70 B | 140 GB ✗ | 70 GB ✓ | 35 GB ✓ |
| Llama 3.1 405B dense | 405 B | 810 GB ✗ | 405 GB ✗ | 203 GB ✗ |
| Llama 3.3 70B dense | 70 B | 140 GB ✗ | 70 GB ✓ | 35 GB ✓ |
| DeepSeek V3 MoE (671B total, 37B active per token) | 671 B | 1342 GB ✗ | 671 GB ✗ | 336 GB ✗ |
| DeepSeek R1 MoE (671B total, 37B active per token) | 671 B | 1342 GB ✗ | 671 GB ✗ | 336 GB ✗ |
| Qwen 2.5 7B dense | 7 B | 14 GB ✓ | 7 GB ✓ | 4 GB ✓ |
| Qwen 2.5 72B dense | 72 B | 144 GB ✗ | 72 GB ✓ | 36 GB ✓ |
| Mixtral 8x7B MoE (46.7B total, 12.9B active per token) | 46.7 B | 93 GB ✗ | 47 GB ✓ | 23 GB ✓ |
| Mixtral 8x22B MoE (141B total, 39B active per token) | 141 B | 282 GB ✗ | 141 GB ✗ | 71 GB ✓ |
| Gemma 2 27B dense | 27 B | 54 GB ✓ | 27 GB ✓ | 14 GB ✓ |
| Command R+ dense | 104 B | 208 GB ✗ | 104 GB ✗ | 52 GB ✓ |
| Kimi K2 MoE (1T total, 32B active per token) | 1000 B | 2000 GB ✗ | 1000 GB ✗ | 500 GB ✗ |
Math: FP16 = params × 2 bytes; INT8 = params × 1 byte; INT4 = params × 0.5 bytes. MoE models sum every expert (full weights on disk), not the per-token active subset.
Export-control status
US Bureau of Industry and Security actions that constrain this chip. Each row links a regulation record (dated, sourced, on DEPLOY) to this chip. This is one of the most-searched and least-structured topics in the whole category, and everyone else writes it as paywalled prose. Structured, dated, sourced.
The US Bureau of Industry and Security's Oct 7 2022 interim final rule restricted export of advanced AI accelerators to China, keyed to a performance threshold (>=4,800 TOPS or >=600 GB/s interconnect). Directly blocked NVIDIA A100 and H100 sales to Chinese entities without a license, triggering the A800 and H800 China-specific variants NVIDIA released within weeks.
Sources: BIS press release, Oct 7 2022 · Federal Register, Oct 13 2022 (rule text)
The Oct 17 2023 update closed the workaround NVIDIA used with A800 and H800: BIS replaced the interconnect-bandwidth threshold with a performance-density metric (total processing performance x performance density) and added a Notified Advanced Computing regime for chips that fell just below the ceiling. Direct effect: A800 and H800 could no longer be exported to China. Triggered NVIDIA's H20 as the next-generation China-specific variant.
Sources: BIS press release, Oct 17 2023 · Federal Register, Oct 25 2023 (rule text)
- BIS AI Diffusion Framework (Jan 13 2025)2025-01-13
The Biden administration's Jan 13 2025 Framework for Artificial Intelligence Diffusion introduced a three-tier country regime for advanced GPU exports: unrestricted allies (Tier 1), per-country compute caps (Tier 2), and blocked destinations (Tier 3, including China, Russia, Iran). Set per-entity National Validated End User caps measured in compute (TFLOPs), and required licensing for large orders even to Tier 2 allies. Rescinded May 15 2025 by the Trump administration before full effect.
Sources: Federal Register, Jan 15 2025 (framework text) · BIS AI Diffusion Framework page
Common questions
Answer-first, sourced. Every claim below traces to a specific field on this page or a cited datasheet. FAQPage schema is emitted so LLM crawlers can lift these verbatim.
How much does NVIDIA H100 Tensor Core GPU cost?
List price at launch was $30,000 (2022-10-13). Street prices vary with generation age and supply. DEPLOY does not currently hold cloud rental $/hr as a first-party field.
How much memory does NVIDIA H100 Tensor Core GPU have?
80 GB of HBM3 at 3,350 GB/s. Vendor datasheet, first-party.
How much power does one NVIDIA H100 Tensor Core GPU draw?
700 W TDP (thermal design power) per chip. System-level draw is higher: see the reference-design section for rack-level kW.
Which data centers have NVIDIA H100 Tensor Core GPU?
7 campuses on the DEPLOY registry carry a cited inventory line for this chip. Full list below with per-campus quantity hints where published.
Which open-weight LLMs fit on one NVIDIA H100 Tensor Core GPU?
Of the 13 public open-weight LLMs DEPLOY tracks: 3 fit at FP16, 7 at INT8, 9 at INT4 (weights only, excludes KV cache). The full table is above with per-model math. Sparsity, offloading and multi-GPU serving change the picture; this row is single-chip weights-only.
Is NVIDIA H100 Tensor Core GPU subject to US export controls?
Yes. 3 BIS actions on record constrain this chip, most recently BIS AI Diffusion Framework (Jan 13 2025) (2025-01-13). Full list with citations above.
Key facts
- Class
- Compute SoC (AI accelerator)
- Designer
- NVIDIA
- Safety-critical?
- No (data-center inference / training)
- Record as of
- 2026-08-25
- Latest cited claim
- 2025-07-01 (across 11 inventory + benchmark rows)
Specifications (18 fields, click to expand)
Vendor datasheet figures (first-party). Dense throughput first; sparse (2:4) in parentheses where NVIDIA quotes it. The exhaustive spec sheet lives on the datasheet URL below the table: this row set covers what buyers actually compare on.
- Process node
- TSMC 4N
- Transistors
- 80 B
- Die size
- 814 mm²
- CUDA cores
- 16,896
- Tensor cores
- 528 (4th gen (Hopper))
- TDP
- 700 W
- Memory
- 80 GB HBM3
- Memory bandwidth
- 3,350 GB/s
- PCIe
- Gen 5 x16 (900 GB/s)
- FP32
- 67 TFLOPS
- TF32 (dense)
- 495 TFLOPS
- FP16 (dense)
- 989 TFLOPS (1,979 sparse)
- FP8 (dense)
- 1,979 TFLOPS (3,958 sparse)
- INT8 (dense)
- 1,979 TOPS (3,958 sparse)
- Form factor
- SXM5
- Announced
- 2022-03-22
- Released
- 2022-10-13
- Launch price
- $30,000 (list)
Source: vendor datasheet
Generation
Used in reference designs
The named rack-scale and pod-scale designs buyers actually order that contain this chip.
- ×8NVIDIA HGX H100 (8-GPU baseboard)platform
- ×256NVIDIA DGX SuperPOD H100 (32-node reference)supercluster
Compare with
Benchmarks
Published measurements per workload. Vendor datasheet numbers are first-party (verified posture); MLPerf / InferenceX / press results come in when their license terms + workload naming permit.
Data centers stocking NVIDIA H100 Tensor Core GPU
Who has the most? →Every campus on the DEPLOY registry with a public claim of NVIDIA H100 Tensor Core GPU on site. Quantity hints are unstructured because underlying disclosures vary by order of magnitude.
- Stargate Norway (Kvandal / Narvik)as of 2025-07-01under constructioninferred
100,000 NVIDIA GPUs targeted by end of 2026 (initial phase; specific SKU may be H100/H200/Blackwell mix)
OpenAI announcement cited 100k Nvidia GPUs without SKU. Attributed conservatively to H100 as the volume-shipping Hopper baseline; upgrade as G42/Aker discloses.
- Compass Datacenters Red Oak Campus (DFW III)as of 2025-01-01partially energizedinferred
H100/H200-generation hyperscaler capacity at the Red Oak DFW III campus
Compass DFW III hosts multiple hyperscalers with H100/H200 fleets; per-tenant chip counts not published.
- ByteDance Pecem Data Center Campus (TikTok)as of 2024-11-25under constructioninferred
ByteDance TikTok/Doubao inference capacity in South America (H100/H200-class)
Chip mix not published; ByteDance's non-China DCs run predominantly NVIDIA H100/H200.
- xAI Colossus (Memphis)as of 2024-09-01partially energizedreported
~100,000 GPUs live in first cluster (Sep 2024); scaled toward 200,000 across the complex
- Crusoe Childress Campusas of 2024-06-01announcedreported
Crusoe operates H100 GPU clusters at Childress
- Meta Hyperion Data Center (Richland Parish)as of 2024-03-12under constructionreported
Nvidia complement to MTIA for training + inference at Hyperion
- Meta Prometheus Data Center (New Albany)as of 2024-03-12partially energizedreported
Nvidia complement to MTIA for training + inference at Prometheus
Sources
- NVIDIA H100 product pageNvidia
- Hopper microarchitectureWikipedia
Adoption rows appear as we verify chip-integration claims per model. Every row is a public claim the maker or a tier-1 source has stated; verification tier (verified / reported / inferred) is shown per row. See every chip on record for the full catalog or the NVIDIA record.