Compute Comparison
vs
All comparisons →

Ampere

NVIDIA A100 40GB

wins

4

The A100 40GB is the more affordable sibling of the A100 80GB.

ML trainingInferenceCost-effective HPC

Volta

NVIDIA V100 32GB

wins

1

The V100 was NVIDIA's flagship AI GPU before the A100.

Legacy ML trainingBudget HPCMature framework support

Performance comparison

Visual bars — winner highlighted

NVIDIA A100 40GB

Metric

NVIDIA V100 32GB

FP16 TFLOPS

312 T
125 T

Bandwidth

1.6k GB/s
900 GB/s

VRAM

402 GB
322 GB

TDP (lower=better)

400 W
300 W

TFLOPS/watt

0.78
0.42

GB/s per watt

3.89
3.00

Full specification table

SpecNVIDIA A100 40GBNVIDIA V100 32GB
ArchitectureAmpereVolta
VRAM40 GB HBM2e32 GB HBM2
VRAM typeHBM2eHBM2
Memory bandwidth1,555 GB/s900 GB/s
FP16 TFLOPS312 TFLOPS125 TFLOPS
TDP400 W300 W
TFLOPS/watt0.78 T/W0.42 T/W
GB/s per watt3.89 GB/s/W3.00 GB/s/W
NVLinkYesYes
Release year20202018

Live cloud pricing

On-demand $/hr across providers — updated in real time

Loading prices…

Power efficiency analysis

TFLOPS/watt and GB/s/watt — critical for data center TCO

NVIDIA A100 40GB

FP16 TFLOPS312 TFLOPS
TDP400 W
TFLOPS/watt0.780 T/W
GB/s per watt3.89 GB/s/W

NVIDIA V100 32GB

FP16 TFLOPS125 TFLOPS
TDP300 W
TFLOPS/watt0.417 T/W
GB/s per watt3.00 GB/s/W

TFLOPS/watt measures compute efficiency — how much AI throughput you get per watt of power consumed. For data centers with PUE of 1.2–1.5, a 10% improvement in TFLOPS/watt translates directly to lower electricity costs and cooling requirements. GB/s/watt measures memory bandwidth efficiency, which is the binding constraint for memory-bound LLM inference workloads.

When to choose each GPU

Choose NVIDIA A100 40GB for:

  • ML training
  • Inference
  • Cost-effective HPC

Choose NVIDIA V100 32GB for:

  • Legacy ML training
  • Budget HPC
  • Mature framework support