Ada Lovelace
NVIDIA L40
wins
1
The NVIDIA L40 is the predecessor to the L40S, offering 48GB GDDR6 and Ada Lovelace architecture.
Ada Lovelace
NVIDIA L40S
wins
2
The L40S is NVIDIA's inference-optimized Ada Lovelace GPU. 48GB GDDR6 and 362 TFLOPS FP16 make it a strong choice for production inference at lower cost than H100.
Performance comparison
Visual bars — winner highlighted
NVIDIA L40
Metric
NVIDIA L40S
FP16 TFLOPS
Bandwidth
VRAM
TDP (lower=better)
TFLOPS/watt
GB/s per watt
Full specification table
| Spec | NVIDIA L40 | NVIDIA L40S |
|---|---|---|
| Architecture | Ada Lovelace | Ada Lovelace |
| VRAM | 48 GB GDDR6 | 48 GB GDDR6 |
| VRAM type | GDDR6 | GDDR6 |
| Memory bandwidth | 864 GB/s | 864 GB/s |
| FP16 TFLOPS | 181 TFLOPS | 362 TFLOPS |
| TDP | 300 W | 350 W |
| TFLOPS/watt | 0.60 T/W | 1.03 T/W |
| GB/s per watt | 2.88 GB/s/W | 2.47 GB/s/W |
| NVLink | No | No |
| Release year | 2022 | 2023 |
Live cloud pricing
On-demand $/hr across providers — updated in real time
Power efficiency analysis
TFLOPS/watt and GB/s/watt — critical for data center TCO
NVIDIA L40
NVIDIA L40S
TFLOPS/watt measures compute efficiency — how much AI throughput you get per watt of power consumed. For data centers with PUE of 1.2–1.5, a 10% improvement in TFLOPS/watt translates directly to lower electricity costs and cooling requirements. GB/s/watt measures memory bandwidth efficiency, which is the binding constraint for memory-bound LLM inference workloads.
When to choose each GPU
Choose NVIDIA L40 for:
- Professional inference
- Rendering
- Visualization AI
Choose NVIDIA L40S for:
- Inference
- Video AI
- Multi-modal workloads
Popular GPU comparisons