All 10 cards, side by side

Full spec sheet, single-GPU performance, and multi-GPU scaling for the entire RTX 30 Series desktop lineup, as they behave under llama.cpp.

Q4_K_M estimates 2-GPU & 4-GPU builds NVLink matters

Full specification sheet

CardDieCUDA coresBoost VRAMBusBandwidth TDPNVLink
RTX 3050GA10725601.78 GHz8 GB GDDR6128-bit224 GB/s130 W—
RTX 3060 8GGA10635841.78 GHz8 GB GDDR6128-bit256 GB/s170 W—
RTX 3060 12GGA10635841.78 GHz12 GB GDDR6192-bit360 GB/s170 W—
RTX 3070GA10458881.73 GHz8 GB GDDR6256-bit448 GB/s220 W—
RTX 3070 TiGA10461441.77 GHz8 GB GDDR6X256-bit912 GB/s290 W—
RTX 3080 10GGA10287041.71 GHz10 GB GDDR6X320-bit760 GB/s320 W—
RTX 3080 12GGA10289601.71 GHz12 GB GDDR6X384-bit912 GB/s350 W—
RTX 3080 TiGA102102401.67 GHz12 GB GDDR6X384-bit960 GB/s350 W—
RTX 3090GA102104961.70 GHz24 GB GDDR6X384-bit936 GB/s350 W112.5 GB/s
RTX 3090 TiGA102107521.86 GHz24 GB GDDR6X384-bit1008 GB/s450 W112.5 GB/s

Green rows: the two cards this site recommends most often. 3080 10G shown at launch spec (21 Gbps refresh ≈ 832 GB/s); 3080 12G bandwidth is derived (384-bit × 19 Gbps). All cards are PCIe 4.0 x16; all are GA10x (Ampere) and fully supported by current llama.cpp CUDA builds.

What the two numbers buy you

Memory bandwidth → token generation speed

3050
224 GB/s
3060 8G
256 GB/s
3060 12G
360 GB/s
3070
448 GB/s
3080 10G
760 GB/s
3070 Ti
912 GB/s
3080 12G
912 GB/s
3090
936 GB/s
3080 Ti
960 GB/s
3090 Ti
1008 GB/s

VRAM → which models fit (Q4_K_M + ~2 GB KV)

8 GB (3050/3060/3070/Ti)
→ 8B
10 GB (3080)
→ 13B tight
12 GB (3060 12G / 3080 12G / 3080 Ti)
→ 13B comfortable
24 GB (3090 / 3090 Ti)
→ 30B, MoE 30B

Single-GPU performance (Q4_K_M, full offload)

Card8B Q413B Q430B Q4Prompt proc (13B)
RTX 3050~22 t/s~12 t/s (Q3)—~200 t/s
RTX 3060 8G~25 t/s~13 t/s (Q3)—~350 t/s
RTX 3060 12G~35 t/s~20 t/s—~350 t/s
RTX 3070~45 t/s~23 t/s (Q3)—~500 t/s
RTX 3070 Ti~90 t/s~50 t/s (Q3)—~900 t/s
RTX 3080 10G~60 t/s~30 t/s (tight ctx)—~700 t/s
RTX 3080 12G~70 t/s~40 t/s—~900 t/s
RTX 3080 Ti~74 t/s~42 t/s—~900 t/s
RTX 3090~75 t/s~40 t/s~15 t/s~1400 t/s
RTX 3090 Ti~80 t/s~44 t/s~16 t/s~1800 t/s

"—" = doesn't fit at that quantization. 13B "Q3" rows: model doesn't fit at Q4_K_M, shown at Q3_K_M for reference. Estimates from typical community benchmarks, ~8K context.

How to read this table Watch what happens at each VRAM step: 8 GB → 12 GB unlocks 13B at full quality (a capability jump), while 448 → 912 GB/s doubles the 13B speed (a comfort jump). The 3090's 24 GB adds an entirely new row (30B). Buying for LLMs, VRAM steps are worth more than bandwidth steps — until you're already at 24 GB.

Multi-GPU builds

llama.cpp splits transformer layers across cards (-sm layer). Token generation speed ≈ (sum of bandwidths) minus an inter-GPU tax: ~10–15% with NVLink, ~25–35% over PCIe 4.0. Prompt processing scales nearly linearly in all cases.

BuildGPU powerInterconnectVRAM13B Q430B Q470B Q4_K_M
2× 3060 12G340 WPCIe 4.024 GB~35 t/s~11 t/s—
3× 3060 12G510 WPCIe 4.036 GB~46 t/s27B Q6: ~11 t/s—
2× 3070440 WPCIe 4.016 GB~40 t/s——
3× 3070660 WPCIe 4.024 GB~40 t/s27B Q4: ~16 t/s—
2× 3080 12G700 WPCIe 4.024 GB~61 t/s~25 t/s—
2× 3080 Ti700 WPCIe 4.024 GB~65 t/s~26 t/s—
3× 3080 Ti1050 WPCIe 4.036 GB~78 t/s27B Q6: ~26 t/s—
4× 3080 Ti1400 WPCIe 4.048 GB~100 t/s27B Q8: ~24 t/s~21 t/s
3090 + 3080 Ti700 WPCIe 4.036 GB~60 t/s~22 t/s~6 t/s (Q2)
2× 3090700 WNVLink48 GB~70 t/s~30 t/s~16 t/s
3090 + 3090 Ti800 WNVLink48 GB~72 t/s~31 t/s~17 t/s
2× 3090 Ti900 WNVLink48 GB~78 t/s~33 t/s~18 t/s
3× 30901050 WNVLink + PCIe72 GB—27B Q8: ~25 t/s~20 t/s (Q4), Q5 at 32K+
3× 3090 Ti1350 WNVLink + PCIe72 GB—27B Q8: ~27 t/s~22 t/s (Q4)
4× 3060 12G680 WPCIe 4.048 GB~60 t/s~20 t/s~10 t/s
4× 30901400 W2× NVLink96 GB—27B Q8: ~32 t/s~28 t/s (Q5)
4× 3090 Ti1800 W2× NVLink96 GB—27B Q8: ~34 t/s~27 t/s (Q4) · Q8: ~17
Best 70B build 2× 3090 (NVLink) — 48 GB, ~16 t/s, 700 W. The price/performance point of the table; 2× 3090 Ti buys ~12% more speed for more money and 200 W.
Best budget multi 2× 3060 12G — 24 GB for a fraction of the price. 30B Q4 at ~11 t/s is a legitimate "big model" experience and 340 W total.
Worst value build 2× 3070 Ti (not shown) — 580 W, no NVLink, 16 GB. A single 3090 beats it on capability, speed and power. The small cards only tile well as 3060 12G.

CPU offloading — when a model outgrows the VRAM

llama.cpp's -ngl N puts the last N transformer layers on the CPU and the rest on the GPUs. The number to internalize: offloaded layers run at CPU speed (~1–3 t/s for a 27B model on a modern 8-core), so the penalty is not proportional:

ScenarioOn CPUResult
27B Q4 in 24 GB: last ~4 layers on CPU to fit~6%~85% of full-GPU speed
27B Q4 in 16 GB: ~half the layers offloaded~50%~3–5 t/s — a different machine
70B Q4 in 24 GB: ~2/3 offloaded~67%~2–3 t/s — demo territory
Offload example llama-server -m 27b-q4_k_m.gguf -ngl 60 -sm layer -c 8192 --tensor-split 1,1 -ot output=CPU 27B Q4 across 2 GPUs, last 4 layers + output embedding on CPU

Multi-GPU commands

Equal cards, split by layer:
llama-server -m model-q4_k_m.gguf -ngl 99 -sm layer --tensor-split 1,1

Unequal pair (weight toward the bigger card):
llama-server -m model.gguf -ngl 99 -sm layer --tensor-split 60,40

Keep N layers on the fast card, rest on the second (hybrid split, good when VRAM differs a lot):
llama-server -m model.gguf -ngl 99 -sm row --n-cpu-moe 0 # row split for very asymmetric pairs

27B (Qwen class) quantization matrix

Qwen-class 27–32B models (Qwen3-32B, Qwen2.5-32B) are the most common "one size up" target. GGUF sizes: Q4_K_M ≈ 19 GB · Q5_K_M ≈ 22 GB · Q6_K ≈ 25 GB · Q8_0 ≈ 33 GB — plus KV cache. That puts hard lines on the table: Q4 needs ~20 GB (i.e. a 24 GB card), Q5 needs ~23 GB, Q6 needs 36 GB+, Q8 wants 48 GB.

BuildVRAM27B Q4_K_M27B Q5_K_M27B Q6_K27B Q8_0
1× 3050 / 3060 8G / 3070 / 3070 Ti8 GB❌ (partial offload ~2–6 t/s)❌❌❌
1× 3060 12G / 3080 12G / 3080 Ti12 GB❌ (partial offload ~5–7 t/s)❌❌❌
1× 3080 10G10 GB❌ (partial offload ~6–8 t/s)❌❌❌
2× 3060 12G24 GB✅ ~10–12 t/s⚠️ ~9 t/s, 4–8K ctx❌❌
2× 308020 GB⚠️ ~18–20 t/s, tight ctx❌❌❌
2× 3080 12G24 GB✅ ~24–26 t/s⚠️ ~19–21 t/s, 4–8K ctx❌❌
2× 3080 Ti24 GB✅ ~25–28 t/s⚠️ ~20–22 t/s, 4–8K ctx❌❌
1× 309024 GB✅ ~13–16 t/s⚠️ ~12–14 t/s, 8K ctx❌ (~1 GB over — IQ6_XS fits)❌
1× 3090 Ti24 GB✅ ~15–17 t/s⚠️ ~13–15 t/s, 8K ctx❌ (~1 GB over — IQ6_XS fits)❌
3090 + 3080 Ti36 GB✅ ~30 t/s✅ ~24 t/s✅ ~22 t/s⚠️ ~13–14 t/s, 8K ctx
2× 3090 (NVLink)48 GB✅ (overkill)✅ ~24–27 t/s✅ ~22–25 t/s✅ ~18–21 t/s
2× 3090 Ti (NVLink)48 GB✅ (overkill)✅ ~26–29 t/s✅ ~24–27 t/s✅ ~19–22 t/s
3× 3060 12G36 GB✅ ~14 t/s✅ ~16 t/s✅ ~10–12 t/s❌
4× 3060 12G48 GB✅ (overkill)✅ ~14 t/s✅ ~12 t/s✅ ~9 t/s
3× 3080 Ti36 GB✅ ~33 t/s✅ ~28 t/s✅ ~26 t/s⚠️ ~19 t/s, 8K ctx
3× 309072 GB✅ (overkill)✅ (overkill)✅ ~27 t/s✅ ~25 t/s, 32K ctx
4× 3090 (NVLink)96 GB✅ everything, long context: Q8 at ~30 t/s with 32K+ ctx
Reading the matrix 24 GB is the 27B line: one 3090/3090 Ti or any 24 GB pair runs Q4, and Q5 if you keep context at 8K. 36 GB unlocks Q6 — 3× 3060 12G is the budget point, 3× 3080 Ti and 3090+3080 Ti the fast ones (or fit IQ6_XS in 24 GB). 48 GB unlocks Q8 — the dual 3090 is the value pick, 4× 3080 Ti the PCIe-only alternative. 72 GB (3× 3090) is the long-context Q8 / 70B-Q5 point. Below 20 GB, 27B only exists as slow partial CPU offload; plan your VRAM budget around 24 GB if Qwen-27B is a core workload.

Buyer's summary

If you want…BuyBecause
8B models, minimum spend3060 8G / 3050Fast enough, cheap, cool
13B at full quality, budget3060 12G12 GB for the price of an 8GB card
13B, fast3080 Ti / 3080 12G960 / 912 GB/s + 12 GB — the Ti when prices are close, the 12G when it's cheaper
27B (Qwen) at Q41× 3090 / 2× 3060 12G24 GB is the line: ~15 t/s single, ~11 t/s budget pair
27B at Q6/Q82× 309048 GB + NVLink — Q6 ~23 t/s, Q8 ~20 t/s
Best all-round single card309024 GB, NVLink, 350 W — the reference card
Maximum single-card speed3090 Ti1008 GB/s, if the premium is small
70B LLMs2× 309048 GB + NVLink = the 70B build
Budget big model (30B)2× 3060 12G24 GB, 340 W, ~11 t/s on 30B
← home first card: RTX 3050 →