DEPLOYDatabase

AI chip · NVIDIA

NVIDIA H100 Tensor Core GPU

Hopper-architecture data-center AI GPU (SXM5 / PCIe). Dominant training + inference accelerator 2023-2024; 80 GB HBM3, ~700W TDP. Fabricated by TSMC on the 4N process.

On the DEPLOY graph: 7 campuses carry a cited inventory line for it · 2 reference-design clusters include it.

Market position

H100 was the chip the 2023-2024 AI-training build-out ran on. Meta bought 350,000. Microsoft, xAI, Anthropic, OpenAI, Oracle, CoreWeave and every neocloud filled DCs with it. Its NVLink 4 + HBM3 combo set the training-cluster template through 2024 until H200 refreshed the memory and Blackwell (B200/GB200) took the crown.

DEPLOY editorial. What the vendor PDF cannot tell you.

Physical-AI cross-link

Derived from claims already on record (vendor first-party TDP, throughput, price, and the reference-design cluster registry). The arithmetic is shown per row so it can be audited. No estimated inputs.

Perf per watt
2.83 FP8 TFLOPS/W
1,979 TFLOPS ÷ 700 W = 2.83
List $ per FP8 TFLOP
$15.16
$30,000 ÷ 1,979 TFLOPS = $15.16
Rack power (in this design)
10.2 kW per NVIDIA HGX H100 (8-GPU baseboard) (~1275 W per chip incl. overhead)
10.2 kW ÷ 8 chips = 1275 W each (system-level, includes CPU/mem/NIC/cooling)
Cooling class (across reference designs)
air
Observed across 2 reference-design clusters containing this chip

What fits in 80 GB

Weights-only footprint for public open-weight LLMs. Pure arithmetic: params × bytes/param. Excludes KV cache and activation memory; add ~10-30% headroom for real serving. A model that does not fit at FP16 may still fit at INT8 or INT4 with quality trade-offs. Not a benchmark.

ModelParamsFP16INT8INT4
Llama 3.1 8B
dense
8 B16 GB ✓8 GB ✓4 GB ✓
Llama 3.1 70B
dense
70 B140 GB ✗70 GB ✓35 GB ✓
Llama 3.1 405B
dense
405 B810 GB ✗405 GB ✗203 GB ✗
Llama 3.3 70B
dense
70 B140 GB ✗70 GB ✓35 GB ✓
DeepSeek V3
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
DeepSeek R1
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
Qwen 2.5 7B
dense
7 B14 GB ✓7 GB ✓4 GB ✓
Qwen 2.5 72B
dense
72 B144 GB ✗72 GB ✓36 GB ✓
Mixtral 8x7B
MoE (46.7B total, 12.9B active per token)
46.7 B93 GB ✗47 GB ✓23 GB ✓
Mixtral 8x22B
MoE (141B total, 39B active per token)
141 B282 GB ✗141 GB ✗71 GB ✓
Gemma 2 27B
dense
27 B54 GB ✓27 GB ✓14 GB ✓
Command R+
dense
104 B208 GB ✗104 GB ✗52 GB ✓
Kimi K2
MoE (1T total, 32B active per token)
1000 B2000 GB ✗1000 GB ✗500 GB ✗

Math: FP16 = params × 2 bytes; INT8 = params × 1 byte; INT4 = params × 0.5 bytes. MoE models sum every expert (full weights on disk), not the per-token active subset.

Export-control status

US Bureau of Industry and Security actions that constrain this chip. Each row links a regulation record (dated, sourced, on DEPLOY) to this chip. This is one of the most-searched and least-structured topics in the whole category, and everyone else writes it as paywalled prose. Structured, dated, sourced.

  1. The US Bureau of Industry and Security's Oct 7 2022 interim final rule restricted export of advanced AI accelerators to China, keyed to a performance threshold (>=4,800 TOPS or >=600 GB/s interconnect). Directly blocked NVIDIA A100 and H100 sales to Chinese entities without a license, triggering the A800 and H800 China-specific variants NVIDIA released within weeks.

    Sources: BIS press release, Oct 7 2022 · Federal Register, Oct 13 2022 (rule text)

  2. The Oct 17 2023 update closed the workaround NVIDIA used with A800 and H800: BIS replaced the interconnect-bandwidth threshold with a performance-density metric (total processing performance x performance density) and added a Notified Advanced Computing regime for chips that fell just below the ceiling. Direct effect: A800 and H800 could no longer be exported to China. Triggered NVIDIA's H20 as the next-generation China-specific variant.

    Sources: BIS press release, Oct 17 2023 · Federal Register, Oct 25 2023 (rule text)

  3. The Biden administration's Jan 13 2025 Framework for Artificial Intelligence Diffusion introduced a three-tier country regime for advanced GPU exports: unrestricted allies (Tier 1), per-country compute caps (Tier 2), and blocked destinations (Tier 3, including China, Russia, Iran). Set per-entity National Validated End User caps measured in compute (TFLOPs), and required licensing for large orders even to Tier 2 allies. Rescinded May 15 2025 by the Trump administration before full effect.

    Sources: Federal Register, Jan 15 2025 (framework text) · BIS AI Diffusion Framework page

Common questions

Answer-first, sourced. Every claim below traces to a specific field on this page or a cited datasheet. FAQPage schema is emitted so LLM crawlers can lift these verbatim.

How much does NVIDIA H100 Tensor Core GPU cost?

List price at launch was $30,000 (2022-10-13). Street prices vary with generation age and supply. DEPLOY does not currently hold cloud rental $/hr as a first-party field.

How much memory does NVIDIA H100 Tensor Core GPU have?

80 GB of HBM3 at 3,350 GB/s. Vendor datasheet, first-party.

How much power does one NVIDIA H100 Tensor Core GPU draw?

700 W TDP (thermal design power) per chip. System-level draw is higher: see the reference-design section for rack-level kW.

Which data centers have NVIDIA H100 Tensor Core GPU?

7 campuses on the DEPLOY registry carry a cited inventory line for this chip. Full list below with per-campus quantity hints where published.

Which open-weight LLMs fit on one NVIDIA H100 Tensor Core GPU?

Of the 13 public open-weight LLMs DEPLOY tracks: 3 fit at FP16, 7 at INT8, 9 at INT4 (weights only, excludes KV cache). The full table is above with per-model math. Sparsity, offloading and multi-GPU serving change the picture; this row is single-chip weights-only.

Is NVIDIA H100 Tensor Core GPU subject to US export controls?

Yes. 3 BIS actions on record constrain this chip, most recently BIS AI Diffusion Framework (Jan 13 2025) (2025-01-13). Full list with citations above.

Key facts

Class
Compute SoC (AI accelerator)
Designer
NVIDIA
Safety-critical?
No (data-center inference / training)
Record as of
2026-08-25
Latest cited claim
2025-07-01 (across 11 inventory + benchmark rows)
Specifications (18 fields, click to expand)

Vendor datasheet figures (first-party). Dense throughput first; sparse (2:4) in parentheses where NVIDIA quotes it. The exhaustive spec sheet lives on the datasheet URL below the table: this row set covers what buyers actually compare on.

Process node
TSMC 4N
Transistors
80 B
Die size
814 mm²
CUDA cores
16,896
Tensor cores
528 (4th gen (Hopper))
TDP
700 W
Memory
80 GB HBM3
Memory bandwidth
3,350 GB/s
PCIe
Gen 5 x16 (900 GB/s)
FP32
67 TFLOPS
TF32 (dense)
495 TFLOPS
FP16 (dense)
989 TFLOPS (1,979 sparse)
FP8 (dense)
1,979 TFLOPS (3,958 sparse)
INT8 (dense)
1,979 TOPS (3,958 sparse)
Form factor
SXM5
Announced
2022-03-22
Released
2022-10-13
Launch price
$30,000 (list)

Source: vendor datasheet

Generation

Used in reference designs

The named rack-scale and pod-scale designs buyers actually order that contain this chip.

See all reference designs →

Compare with

See every chip comparison →

Benchmarks

Published measurements per workload. Vendor datasheet numbers are first-party (verified posture); MLPerf / InferenceX / press results come in when their license terms + workload naming permit.

WorkloadValueUnitSourceAs of
memory bandwidth3350GB/svendor2023-01-01
memory capacity80GB HBM3vendor2023-01-01
peak tflops fp16 dense(SXM5 form factor)989.4TFLOPSvendor2023-01-01
peak tflops fp8 dense1978.9TFLOPSvendor2023-01-01

Data centers stocking NVIDIA H100 Tensor Core GPU

Who has the most? →

Every campus on the DEPLOY registry with a public claim of NVIDIA H100 Tensor Core GPU on site. Quantity hints are unstructured because underlying disclosures vary by order of magnitude.

  1. under constructioninferred

    100,000 NVIDIA GPUs targeted by end of 2026 (initial phase; specific SKU may be H100/H200/Blackwell mix)

    OpenAI announcement cited 100k Nvidia GPUs without SKU. Attributed conservatively to H100 as the volume-shipping Hopper baseline; upgrade as G42/Aker discloses.

    OpenAI

  2. partially energizedinferred

    H100/H200-generation hyperscaler capacity at the Red Oak DFW III campus

    Compass DFW III hosts multiple hyperscalers with H100/H200 fleets; per-tenant chip counts not published.

    Compass Datacenters

  3. under constructioninferred

    ByteDance TikTok/Doubao inference capacity in South America (H100/H200-class)

    Chip mix not published; ByteDance's non-China DCs run predominantly NVIDIA H100/H200.

    Reuters

  4. xAI Colossus (Memphis)as of 2024-09-01
    partially energizedreported

    ~100,000 GPUs live in first cluster (Sep 2024); scaled toward 200,000 across the complex

    SemiAnalysis

  5. Crusoe Childress Campusas of 2024-06-01
    announcedreported

    Crusoe operates H100 GPU clusters at Childress

    Crusoe

  6. under constructionreported

    Nvidia complement to MTIA for training + inference at Hyperion

    Meta Engineering

  7. partially energizedreported

    Nvidia complement to MTIA for training + inference at Prometheus

    Meta Engineering

Sources

Adoption rows appear as we verify chip-integration claims per model. Every row is a public claim the maker or a tier-1 source has stated; verification tier (verified / reported / inferred) is shown per row. See every chip on record for the full catalog or the NVIDIA record.