Compute Comparison
NVIDIAAmpere2020

RTX A6000

Ampere flagship professional GPU. 48GB GDDR6 with NVLink 3.0. Predecessor to RTX 6000 Ada. Widely available at lower cost.

VRAM
48GB
GDDR6
FP16
77.4
TFLOPS
Bandwidth
768.0
GB/s
TDP
300W
power
Best for:Large model inferenceMulti-GPU NVLink setupsProfessional rendering + AI

RTX A6000 Overview

The RTX A6000 is an Ampere professional GPU with 48GB of ECC GDDR6, 77.4 TFLOPS of FP16/BF16 performance, and a 300W power envelope. As the predecessor to RTX 6000 Ada, it trades newer architecture performance for mature professional drivers, large memory, and broad availability.

Its 768 GB/s memory bandwidth is adequate for many workstation and inference workloads, while 48GB gives much more room than 24GB consumer cards for model weights, assets, and batch state. NVLink 3.0 allows paired professional cards to work together, but GDDR6 bandwidth remains a constraint compared with A100-class HBM platforms.

The RTX A6000 works well for professional rendering plus AI, larger-model inference, and cost-conscious multi-GPU workstation setups. It lacks FP8 support and does not match modern Ada or Hopper cards for pure inference throughput, so it is best chosen for capacity and professional features rather than speed alone.

Memory

VRAM48 GB
Memory TypeGDDR6
Bandwidth768 GB/s
NVLink BW112 GB/s

Compute Performance

FP3238.7 TFLOPS
FP1677.4 TFLOPS
BF1677.4 TFLOPS
INT8309 TOPS

Hardware Specifications

Chip

ArchitectureGA102
GenerationAmpere
Process NodeSamsung 8nm
Transistors28.3B
Die Size628 mm²
Release DateNovember 17, 2020
Launch MSRP$4,650

Processors

CUDA / Shader Cores10,752
Tensor Cores336
RT Cores84

Clocks

Base Clock1,455 MHz
Boost Clock1,860 MHz

Memory

VRAM48 GB
Memory TypeGDDR6
Memory Bus384-bit
Bandwidth768 GB/s
L2 Cache6 MB
NVLink BW112 GB/s

Power

TDP300 W
InterconnectNVLink 3.0 / PCIe 4.0

Relative Performance

FP16 Compute1%
VRAM Capacity17%
Mem Bandwidth5%

Relative to highest-spec GPU in database

Limitations

GDDR6 memory bandwidth far below HBM alternatives
No NVLink on A4000/A2000 — single-card VRAM ceiling
Professional but older Ampere architecture — no FP8 support

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare RTX A6000 vs…

Use Case Guidance

Large model inference
Multi-GPU NVLink setups
Professional rendering + AI

LLM Model Size Guidance

Max model (FP16)~24Bparameters at FP16 precision
Max model (INT8)~48Bparameters at INT8 precision
Max model (INT4)~96Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Ampere-class hardware.

Browse all 42 LLM models

RTX A6000 vs Alternatives — Spec Comparison

SpecRTX A6000 thisA40RTX 4070 TiRTX 3090
VRAM48GB GDDR648GB GDDR612GB GDDR6X24GB GDDR6X
Memory Bandwidth768 GB/s696 GB/s504 GB/s936 GB/s
FP16 TFLOPS77.474.880.271
BF16 TFLOPS77.474.880.271
FP8 TFLOPS
INT8 TOPS309149.7160142
TDP300W300W285W350W
Process NodeSamsung 8nmSamsung 8nmTSMC 4NSamsung 8nm
ArchitectureGA102GA102AD104GA102
Release Year2020202020232020
Max model (FP16)~24B params~24B params~6B params~12B params
Max model (INT4)~96B params~96B params~24B params~48B params
▲ indicates best value in row · FP16/BF16 TFLOPS at full precision · Max model estimates at 2 bytes/param (FP16) and 0.5 bytes/param (INT4)Full side-by-side comparison

Related GPUs

Frequently Asked Questions

How much VRAM does the RTX A6000 have?

The RTX A6000 has 48GB of GDDR6 memory with 768 GB/s bandwidth. This enables running models up to approximately 96B parameters at INT4 precision, 48B at INT8, or 24B at FP16.

What is the FP16 performance of the RTX A6000?

The RTX A6000 delivers 77.4 TFLOPS of FP16 performance and 77.4 TFLOPS BF16. INT8 throughput is 309 TOPS. For transformer inference, memory bandwidth (768 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the RTX A6000 best used for?

The RTX A6000 is best suited for: Large model inference, Multi-GPU NVLink setups, Professional rendering + AI. Ampere flagship professional GPU. 48GB GDDR6 with NVLink 3.0. Predecessor to RTX 6000 Ada. Widely available at lower cost.

What interconnect does the RTX A6000 use?

The RTX A6000 uses NVLink 3.0 / PCIe 4.0 with 112 GB/s NVLink bandwidth for multi-GPU configurations. NVLink enables near-linear tensor-parallel scaling across multiple cards for models that exceed single-card VRAM.

What LLM model sizes can the RTX A6000 run?

With 48GB of GDDR6, the RTX A6000 can run models up to approximately 24B parameters at FP16 (2 bytes/param), 48B at INT8 (1 byte/param), or 96B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the RTX A6000 compare to the A100 for LLM inference?

The RTX A6000 has 77.4 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 768 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX A6000's lower cost.

What is the power consumption of the RTX A6000?

The RTX A6000 has a TDP (Thermal Design Power) of 300W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 300W, the RTX A6000 is in the mid-range tier — compatible with standard data center power infrastructure.

Ready to rent?

Compare RTX A6000 prices across 102+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices