Compute Comparison
NVIDIAAmpere2020

RTX 3090

Previous-gen consumer GPU. Cheapest option for 24GB VRAM workloads. Limited BF16 support.

VRAM
24GB
GDDR6X
FP16
71.0
TFLOPS
Bandwidth
936.0
GB/s
TDP
350W
power
Best for:Budget inferenceSmall model fine-tuningExperimentation

RTX 3090 Overview

The RTX 3090 is an Ampere consumer GPU that remains relevant because it offers 24GB of GDDR6X at a comparatively low price. It provides 71 TFLOPS of FP16 performance and 936 GB/s of bandwidth, giving it considerably more memory headroom than many newer midrange consumer cards.

The 24GB pool can support smaller full-precision models and larger models after quantization, while 936 GB/s is respectable for its generation. Its shortcomings are equally important: it lacks ECC, does not provide NVLink memory pooling, and does not have native BF16 acceleration or Hopper-era FP8 support.

The card is a useful budget option for experimentation, small-model fine-tuning, and inference where 24GB matters more than newest-generation features. It is not intended for reliability-sensitive data-center deployments or for fast modern transformer training that relies on BF16 or FP8 paths.

Memory

VRAM24 GB
Memory TypeGDDR6X
Bandwidth936 GB/s

Compute Performance

FP3235.6 TFLOPS
FP1671 TFLOPS
BF1671 TFLOPS
INT8142 TOPS

Hardware Specifications

Chip

ArchitectureGA102
GenerationAmpere
Process NodeSamsung 8nm
Transistors28.3B
Die Size628 mm²
Release DateSeptember 24, 2020
Launch MSRP$1,499

Processors

CUDA / Shader Cores10,496
Tensor Cores328
RT Cores82

Clocks

Base Clock1,395 MHz
Boost Clock1,695 MHz

Memory

VRAM24 GB
Memory TypeGDDR6X
Memory Bus384-bit
Bandwidth936 GB/s
L2 Cache6 MB

Power

TDP350 W
InterconnectPCIe 4.0

Relative Performance

FP16 Compute1%
VRAM Capacity8%
Mem Bandwidth6%

Relative to highest-spec GPU in database

Limitations

No BF16 hardware acceleration — software emulation only
No NVLink — single-card VRAM ceiling
Older Ampere architecture — no FP8 support

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare RTX 3090 vs…

Use Case Guidance

Budget inference
Small model fine-tuning
Experimentation

LLM Model Size Guidance

Max model (FP16)~12Bparameters at FP16 precision
Max model (INT8)~24Bparameters at INT8 precision
Max model (INT4)~48Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Ampere-class hardware.

Browse all 42 LLM models

RTX 3090 vs Alternatives — Spec Comparison

SpecRTX 3090 thisRTX 5060 TiRTX 3080 TiRTX A5500 24GB
VRAM24GB GDDR6X16GB GDDR712GB GDDR6X24GB GDDR6
Memory Bandwidth936 GB/s672 GB/s912 GB/s768 GB/s
FP16 TFLOPS7170.468.268.2
BF16 TFLOPS7170.468.268.2
FP8 TFLOPS
INT8 TOPS142140.8136272.8
TDP350W180W350W230W
Process NodeSamsung 8nmTSMC 4NPSamsung 8nmSamsung 8nm
ArchitectureGA102GB206GA102GA102
Release Year2020202520212021
Max model (FP16)~12B params~8B params~6B params~12B params
Max model (INT4)~48B params~32B params~24B params~48B params
▲ indicates best value in row · FP16/BF16 TFLOPS at full precision · Max model estimates at 2 bytes/param (FP16) and 0.5 bytes/param (INT4)Full side-by-side comparison

Related GPUs

Frequently Asked Questions

How much VRAM does the RTX 3090 have?

The RTX 3090 has 24GB of GDDR6X memory with 936 GB/s bandwidth. This enables running models up to approximately 48B parameters at INT4 precision, 24B at INT8, or 12B at FP16.

What is the FP16 performance of the RTX 3090?

The RTX 3090 delivers 71 TFLOPS of FP16 performance and 71 TFLOPS BF16. INT8 throughput is 142 TOPS. For transformer inference, memory bandwidth (936 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the RTX 3090 best used for?

The RTX 3090 is best suited for: Budget inference, Small model fine-tuning, Experimentation. Previous-gen consumer GPU. Cheapest option for 24GB VRAM workloads. Limited BF16 support.

What interconnect does the RTX 3090 use?

The RTX 3090 uses PCIe 4.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.

What LLM model sizes can the RTX 3090 run?

With 24GB of GDDR6X, the RTX 3090 can run models up to approximately 12B parameters at FP16 (2 bytes/param), 24B at INT8 (1 byte/param), or 48B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the RTX 3090 compare to the A100 for LLM inference?

The RTX 3090 has 71 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 936 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX 3090's lower cost.

What is the power consumption of the RTX 3090?

The RTX 3090 has a TDP (Thermal Design Power) of 350W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 350W, the RTX 3090 is in the mid-range tier — compatible with standard data center power infrastructure.

Ready to rent?

Compare RTX 3090 prices across 102+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices