Compute Comparison
vs
All comparisons →

Turing

NVIDIA T4

wins

2

The NVIDIA T4 is the most widely deployed inference GPU in the cloud.

Low-cost inferenceEdge AIMulti-tenant serving

Ada Lovelace

NVIDIA L4

wins

3

The NVIDIA L4 is the most power-efficient Ada Lovelace GPU at just 72W TDP.

Efficient inferenceVideo transcodingEdge AI

Performance comparison

Visual bars — winner highlighted

NVIDIA T4

Metric

NVIDIA L4

FP16 TFLOPS

65 T
242 T

Bandwidth

320 GB/s
300 GB/s

VRAM

166 GB
246 GB

TDP (lower=better)

70 W
72 W

TFLOPS/watt

0.93
3.36

GB/s per watt

4.57
4.17

Full specification table

SpecNVIDIA T4NVIDIA L4
ArchitectureTuringAda Lovelace
VRAM16 GB GDDR624 GB GDDR6
VRAM typeGDDR6GDDR6
Memory bandwidth320 GB/s300 GB/s
FP16 TFLOPS65 TFLOPS242 TFLOPS
TDP70 W72 W
TFLOPS/watt0.93 T/W3.36 T/W
GB/s per watt4.57 GB/s/W4.17 GB/s/W
NVLinkNoNo
Release year20182023

Live cloud pricing

On-demand $/hr across providers — updated in real time

Loading prices…

Power efficiency analysis

TFLOPS/watt and GB/s/watt — critical for data center TCO

NVIDIA T4

FP16 TFLOPS65 TFLOPS
TDP70 W
TFLOPS/watt0.929 T/W
GB/s per watt4.57 GB/s/W

NVIDIA L4

FP16 TFLOPS242 TFLOPS
TDP72 W
TFLOPS/watt3.361 T/W
GB/s per watt4.17 GB/s/W

TFLOPS/watt measures compute efficiency — how much AI throughput you get per watt of power consumed. For data centers with PUE of 1.2–1.5, a 10% improvement in TFLOPS/watt translates directly to lower electricity costs and cooling requirements. GB/s/watt measures memory bandwidth efficiency, which is the binding constraint for memory-bound LLM inference workloads.

When to choose each GPU

Choose NVIDIA T4 for:

  • Low-cost inference
  • Edge AI
  • Multi-tenant serving

Choose NVIDIA L4 for:

  • Efficient inference
  • Video transcoding
  • Edge AI