A40
48GB GDDR6 at a lower price than A100. Good for workloads needing large VRAM without HBM cost.
A40 Overview
The NVIDIA A40 is an Ampere data-center GPU that combines 48GB of GDDR6 with a 300W thermal envelope and broad professional graphics support. Its 74.8 TFLOPS of FP16/BF16 compute makes it useful for mixed visualization and AI work, while its 48GB capacity is the feature that separates it from many lower-cost inference cards.
Its 696 GB/s GDDR6 bandwidth is far below A100-class HBM throughput, so it is better at workloads that need capacity or graphics features than at bandwidth-intensive autoregressive generation. The 48GB allocation gives useful room for medium-sized models, quantized larger models, and rendering assets, but there is no NVLink to combine memory across cards.
The A40 fits virtual workstations, rendering plus inference, and mid-size model serving where 48GB is more important than peak tokens per second. It is an older Ampere option with no FP8 support and is not a direct substitute for an HBM accelerator in large-model training.
Memory
Compute Performance
Hardware Specifications
Chip
Processors
Clocks
Memory
Power
Relative Performance
Relative to highest-spec GPU in database
Limitations
Live Cloud PricingOn-demand hourly rates
Compare A40 vs…
Use Case Guidance
LLM Model Size Guidance
Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.
Related Guides
LLM APIs Running on This GPU Class
Providers that serve frontier LLM inference on Ampere-class hardware.
A40 vs Alternatives — Spec Comparison
| Spec | A40 this | RTX A6000 | RTX 3090 | RTX 5060 Ti |
|---|---|---|---|---|
| VRAM | 48GB GDDR6▲ | 48GB GDDR6▲ | 24GB GDDR6X | 16GB GDDR7 |
| Memory Bandwidth | 696 GB/s | 768 GB/s | 936 GB/s▲ | 672 GB/s |
| FP16 TFLOPS | 74.8 | 77.4▲ | 71 | 70.4 |
| BF16 TFLOPS | 74.8 | 77.4▲ | 71 | 70.4 |
| FP8 TFLOPS | — | — | — | — |
| INT8 TOPS | 149.7 | 309▲ | 142 | 140.8 |
| TDP | 300W | 300W | 350W▲ | 180W |
| Process Node | Samsung 8nm▲ | Samsung 8nm▲ | Samsung 8nm▲ | TSMC 4NP |
| Architecture | GA102 | GA102 | GA102 | GB206▲ |
| Release Year | 2020 | 2020 | 2020 | 2025▲ |
| Max model (FP16) | ~24B params▲ | ~24B params▲ | ~12B params | ~8B params |
| Max model (INT4) | ~96B params▲ | ~96B params▲ | ~48B params | ~32B params |
Related GPUs
Frequently Asked Questions
How much VRAM does the A40 have?
The A40 has 48GB of GDDR6 memory with 696 GB/s bandwidth. This enables running models up to approximately 96B parameters at INT4 precision, 48B at INT8, or 24B at FP16.
What is the FP16 performance of the A40?
The A40 delivers 74.8 TFLOPS of FP16 performance and 74.8 TFLOPS BF16. INT8 throughput is 149.7 TOPS. For transformer inference, memory bandwidth (696 GB/s) is often the binding constraint rather than raw TFLOPS.
What is the A40 best used for?
The A40 is best suited for: Visualization + compute, Mid-size inference, Virtual workstations. 48GB GDDR6 at a lower price than A100. Good for workloads needing large VRAM without HBM cost.
What interconnect does the A40 use?
The A40 uses PCIe 4.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.
What LLM model sizes can the A40 run?
With 48GB of GDDR6, the A40 can run models up to approximately 24B parameters at FP16 (2 bytes/param), 48B at INT8 (1 byte/param), or 96B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.
How does the A40 compare to the A100 for LLM inference?
The A40 has 74.8 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 696 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the A40's lower cost.
What is the power consumption of the A40?
The A40 has a TDP (Thermal Design Power) of 300W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 300W, the A40 is in the mid-range tier — compatible with standard data center power infrastructure.
Ready to rent?
Compare A40 prices across 102+ providers
Live on-demand & spot rates · monthly cost estimates · availability status