Every desktop card the site tracks — RTX 30, 40 and 50 Series — on the two numbers that matter for llama.cpp: VRAM capacity and memory bandwidth, plus the multi-GPU interconnect story.
| Card | Series | Die | VRAM | Bus | Bandwidth | TDP | NVLink |
|---|---|---|---|---|---|---|---|
| RTX 3050 | 30 | GA107 | 8 GB GDDR6 | 128-bit | 224 GB/s | 130 W | — |
| RTX 3060 8G | 30 | GA106 | 8 GB GDDR6 | 128-bit | 256 GB/s | 170 W | — |
| RTX 3060 12G | 30 | GA106 | 12 GB GDDR6 | 192-bit | 360 GB/s | 170 W | — |
| RTX 3070 | 30 | GA104 | 8 GB GDDR6 | 256-bit | 448 GB/s | 220 W | — |
| RTX 3070 Ti | 30 | GA104 | 8 GB GDDR6X | 256-bit | 912 GB/s | 290 W | — |
| RTX 3080 10G | 30 | GA102 | 10 GB GDDR6X | 320-bit | 760 GB/s | 320 W | — |
| RTX 3080 12G | 30 | GA102 | 12 GB GDDR6X | 384-bit | 912 GB/s | 350 W | — |
| RTX 3080 Ti | 30 | GA102 | 12 GB GDDR6X | 384-bit | 960 GB/s | 350 W | — |
| RTX 3090 | 30 | GA102 | 24 GB GDDR6X | 384-bit | 936 GB/s | 350 W | 112.5 GB/s |
| RTX 3090 Ti | 30 | GA102 | 24 GB GDDR6X | 384-bit | 1008 GB/s | 450 W | 112.5 GB/s |
| RTX 4050† | 40 | AD107 | 8 GB GDDR6 | 128-bit | 224 GB/s | 135 W | — |
| RTX 4060 8G | 40 | AD107 | 8 GB GDDR6 | 128-bit | 272 GB/s | 115 W | — |
| RTX 4060 Ti 8G | 40 | AD106 | 8 GB GDDR6 | 128-bit | 288 GB/s | 160 W | — |
| RTX 4060 Ti 16G | 40 | AD106 | 16 GB GDDR6 | 128-bit | 288 GB/s | 160 W | — |
| RTX 4070 12G | 40 | AD103 | 12 GB GDDR6X | 192-bit | 504 GB/s | 200 W | — |
| RTX 4070 Super 12G | 40 | AD103 | 12 GB GDDR6X | 192-bit | 504 GB/s | 220 W | — |
| RTX 4070 Ti 12G | 40 | AD103 | 12 GB GDDR6X | 192-bit | 672 GB/s | 285 W | — |
| RTX 4070 Ti Super 16G | 40 | AD103 | 16 GB GDDR6X | 256-bit | 672 GB/s | 285 W | — |
| RTX 4080 16G | 40 | AD102 | 16 GB GDDR6X | 256-bit | 717 GB/s | 320 W | — |
| RTX 4090 24G | 40 | AD102 | 24 GB GDDR6X | 384-bit | 1008 GB/s | 450 W | — |
| RTX 5050 8G | 50 | GB207 | 8 GB GDDR6 | 128-bit | 320 GB/s | 130 W | — |
| RTX 5060 8G | 50 | GB206 | 8 GB GDDR7 | 128-bit | 448 GB/s | 145 W | — |
| RTX 5060 Ti 16G | 50 | GB206 | 16 GB GDDR7 | 128-bit | 448 GB/s | 180 W | — |
| RTX 5070 12G | 50 | GB205 | 12 GB GDDR7 | 192-bit | 672 GB/s | 250 W | — |
| RTX 5070 Super 12G† | 50 | GB205 | 12 GB GDDR7 | 192-bit | 672 GB/s | 250 W | — |
| RTX 5070 Ti 16G | 50 | GB203 | 16 GB GDDR7 | 256-bit | 896 GB/s | 300 W | — |
| RTX 5080 16G | 50 | GB203 | 16 GB GDDR7 | 256-bit | 960 GB/s | 360 W | — |
| RTX 5090 32G | 50 | GB202 | 32 GB GDDR7 | 512-bit | 1792 GB/s | 575 W | — |
Green rows: the site's most-recommended cards per tier. All cards are PCIe x16 (4.0 on the RTX 30/40 Series, 5.0 on the RTX 50 Series) and supported by current llama.cpp CUDA builds. RTX 30 Series card names link to full guides. † Not listed on NVIDIA’s official comparison page (fetched 2026-09-24) — specs on that card’s page are flagged as supplementary.
On the same model, token-generation speed tracks this list almost linearly. The RTX 5090 alone covers the entire previous two generations' spread in one step above them.
| VRAM tier | RTX 30 Series | RTX 40 Series | RTX 50 Series | Runs (Q4_K_M + KV) |
|---|---|---|---|---|
| 8 GB | 3050, 3060 8G, 3070, 3070 Ti | 4050, 4060, 4060 Ti 8G | 5050, 5060 | 8B Q8, 13B Q3–Q4 tight |
| 10–12 GB | 3080 10G, 3080 12G, 3060 12G, 3080 Ti | 4070, 4070 S, 4070 Ti | 5070, 5070 S | 13B Q4 comfortable |
| 16 GB | — | 4060 Ti 16G, 4070 Ti S, 4080 | 5060 Ti 16G, 5070 Ti, 5080 | 13B Q8, MoE ~30B, 27B Q3–Q4 |
| 24 GB | 3090, 3090 Ti | 4090 | — | 27B Q4–Q5, 30B Q4, MoE ~30B |
| 32 GB | — | — | 5090 | 27B Q8, 30B Q5+, 70B Q2–IQ3 |
llama.cpp splits transformer layers across cards (-sm layer). Token generation
speed ≈ (sum of bandwidths) minus an inter-GPU tax:
| Interconnect | Available on | Link speed | Split tax |
|---|---|---|---|
| NVLink | RTX 3090 / 3090 Ti only (2-way) | 112.5 GB/s | ~10–15% |
| PCIe 4.0 x16 | all other RTX 30 / 40 Series cards | ~31 GB/s per dir. | ~25–35% |
| PCIe 5.0 x16 | all RTX 50 Series cards | ~63 GB/s per dir. | ~20–30% (est.) |
--tensor-split lets you weight layers toward the faster card. The PCIe
tax applies to traffic between different cards regardless of generation.Full mechanics, CPU offloading and exact flags: multi-GPU guide →. RTX 30 Series build-by-build scaling tables: 30 Series comparison →.
| If you want… | Buy | Because |
|---|---|---|
| 8B models, minimum spend | 3060 8G / 4060 / 5060 | All fine; used 3060 is cheapest per token |
| 13B at full quality, budget | 3060 12G / 4070 / 5070 | 12 GB tier; pick by price |
| 13B Q8 + MoE ~30B, single card | 4060 Ti 16G / 4070 Ti S / 5060 Ti 16G | The 16 GB tier is new since Ampere — the RTX 30 Series never had it |
| 27B (Qwen) at Q4–Q5 | 3090 / 4090 / 2× 16 GB pair | 24 GB is the line (or 32 GB combined from two 16s) |
| 27B at Q8, single card | 5090 | Only consumer card with 32 GB |
| 70B LLMs, best value | 2× 3090 (NVLink) | 48 GB, ~10–15% split tax, cheapest $/GB of the big builds |
| 70B LLMs, new hardware | 2× 5090 (PCIe 5.0) | 64 GB — Q4 with headroom, Q5/Q6 in long context |
| Maximum single-card speed | 5090 | 1792 GB/s — ~1.8× a 3090 Ti on token generation |
| Best used value | 3090 | 24 GB + NVLink at a used price; the reference card |