Compute Comparison
vs
All comparisons →

Hopper

NVIDIA H100 80GB

wins

4

The NVIDIA H100 is the gold standard for AI training and inference. With 80GB HBM3 memory and NVLink 4.0 support, it dominates large-model training and multi-GPU clusters.

LLM trainingLarge-scale inferenceHPC workloads

Ada Lovelace

NVIDIA L40S

wins

1

The L40S is NVIDIA's inference-optimized Ada Lovelace GPU. 48GB GDDR6 and 362 TFLOPS FP16 make it a strong choice for production inference at lower cost than H100.

InferenceVideo AIMulti-modal workloads

Performance comparison

Visual bars — winner highlighted

NVIDIA H100 80GB

Metric

NVIDIA L40S

FP16 TFLOPS

989 T
362 T

Bandwidth

3.4k GB/s
864 GB/s

VRAM

803 GB
486 GB

TDP (lower=better)

700 W
350 W

TFLOPS/watt

1.41
1.03

GB/s per watt

4.79
2.47

Full specification table

SpecNVIDIA H100 80GBNVIDIA L40S
ArchitectureHopperAda Lovelace
VRAM80 GB HBM348 GB GDDR6
VRAM typeHBM3GDDR6
Memory bandwidth3,350 GB/s864 GB/s
FP16 TFLOPS989 TFLOPS362 TFLOPS
TDP700 W350 W
TFLOPS/watt1.41 T/W1.03 T/W
GB/s per watt4.79 GB/s/W2.47 GB/s/W
NVLinkYesNo
Release year20222023

Live cloud pricing

On-demand $/hr across providers — updated in real time

Loading prices…

Power efficiency analysis

TFLOPS/watt and GB/s/watt — critical for data center TCO

NVIDIA H100 80GB

FP16 TFLOPS989 TFLOPS
TDP700 W
TFLOPS/watt1.413 T/W
GB/s per watt4.79 GB/s/W

NVIDIA L40S

FP16 TFLOPS362 TFLOPS
TDP350 W
TFLOPS/watt1.034 T/W
GB/s per watt2.47 GB/s/W

TFLOPS/watt measures compute efficiency — how much AI throughput you get per watt of power consumed. For data centers with PUE of 1.2–1.5, a 10% improvement in TFLOPS/watt translates directly to lower electricity costs and cooling requirements. GB/s/watt measures memory bandwidth efficiency, which is the binding constraint for memory-bound LLM inference workloads.

When to choose each GPU

Choose NVIDIA H100 80GB for:

  • LLM training
  • Large-scale inference
  • HPC workloads

Choose NVIDIA L40S for:

  • Inference
  • Video AI
  • Multi-modal workloads