Llama 3.3 70B quantized: the hardware you need
The VRAM maths for running Meta's Llama 3.3 70B locally, and the refurbished workstation build that hits the price to performance sweet spot in the UK.

Llama 3.3 70B hits roughly GPT-4 class quality on many benchmarks, yet runs on consumer hardware if you understand the quantisation trade off. The catch is that 70 billion parameters do not fit on a single 16GB card at native precision. You have to shrink the weights, and the shrinkage you pick decides the hardware bill.
The VRAM maths, honestly
A 70B model at native FP16 needs roughly 140GB of memory for the weights alone, and 160GB plus once you add KV cache, context and activation memory. That is data centre territory.
Quantisation is the escape hatch. The three tiers that matter for Llama 3.3 70B:
- Q8_0 (8 bit): 75GB to 80GB working set. Near lossless, indistinguishable from FP16 in blind tests.
- Q5_K_M (5 bit): 55GB to 58GB working set. Very small quality drop, still highly capable.
- Q4_K_M (4 bit): 46GB to 50GB working set. The popular default and the tier most home rigs target.
Go smaller than Q4 and you get repetition and reasoning breakdowns. Q4_K_M is the practical floor.
Single GPU is a fantasy for 70B
Nothing in the consumer stack has enough VRAM for even Q4 on its own. The RTX 5090 tops out at 32GB, the RTX 5080 sits at 16GB. Refurbished workstation cards like the RTX A6000 or RTX 6000 Ada (48GB each) can just about fit Q4 with a short context, but they cost between £3,800 and £6,500 used and still cannot fit Q5 or Q8. If someone tells you they run 70B on a single 24GB card, they are running Q2 and the outputs show it.
Dual GPU strategies that work
Two cards is the honest answer for 70B at home. llama.cpp, ExLlamaV2 and vLLM all support tensor splitting across GPUs, and the performance penalty is modest as long as the cards sit on decent PCIe lanes. The three configurations that make sense in mid 2026:
- Two RTX 3090 (48GB total): the value king. Handles Q4_K_M with a 16k to 32k context. Used cards sit at £600 to £750 each. Draws 700W under load, so plan for a 1200W PSU.
- Two RTX 4090 (48GB total): faster prompt processing, but used prices stay stubborn at £1,300 to £1,600 per card. Real gain only if you also fine tune.
- One RTX 5090 plus one RTX 5070 Ti (48GB total): the balanced new build. The 5090 holds most layers, the 5070 Ti takes the tail. Comfortable Q4, tight Q5. Around £2,400 to £2,700 new, warranty on both.
For Q5_K_M you want closer to 60GB total (dual RTX A6000 or 5090 plus 5090). For Q8 you are into three card territory or a single H100.
The BAV build: Ryzen 7 5700X plus RTX 5070 Ti
For most UK buyers who want Llama 3.3 70B at Q4 without workstation prices, the sensible starting point is a refurbished Ryzen 7 5700X tower paired with an RTX 5070 Ti as the first GPU. Add a second card later, and you have a dual GPU 70B rig without buying everything at once.
The base spec: AMD Ryzen 7 5700X (8 cores, 16 threads), 32GB or 64GB DDR4 3200, a B550 or X570 board with two proper x16 slots, NVMe boot drive and a Gold rated PSU sized for two cards. The 5700X matches the 5800X in inference workloads, runs cooler, and costs less refurbished.
Configurations range from single RTX 5070 Ti setups for smaller models, through to dual card variants ready for Llama 3.3 70B on day one.
See the current Birmingham AV Ryzen 7 5700X plus RTX 5070 Ti build on eBay
Why refurbished is the right call
An LLM workstation runs at moderate CPU load with GPUs doing the heavy lifting. It is the workload a refurbished business tower handles well, because the chassis, PSU and board were engineered for sustained duty. You are buying a serviced platform with a modern GPU bolted in, not a worn out gaming PC.
A new Ryzen tower with equivalent PSU headroom sits at £900 to £1,200 before the GPU goes in. A refurbished build lands well below that, freeing budget for a second GPU. Given the GPU is 70 to 80 percent of inference performance, that is money in the right place.
Power, cooling and PSU sizing
Two 24GB cards can pull 600W to 750W under load. Add 120W for the CPU and 50W for the rest, and you want a Gold or Platinum rated PSU at 1000W to 1200W. Cheap PSUs are the single most common failure point in DIY LLM rigs. A three slot gap between GPUs is comfortable, two slots is workable with careful fan curves.
FAQ
How many tokens per second on dual 3090 at Q4?
Realistic numbers are 12 to 18 tokens per second for generation on dual RTX 3090 with Q4_K_M and a 16k context. Prompt processing runs 300 to 600 tokens per second. That is comfortably faster than reading speed.
Is Q4 quality good enough for coding tasks?
Yes for most workflows. Q4_K_M on Llama 3.3 70B holds up well for code completion, refactoring and explanation. It will occasionally miss an edge case that Q8 catches, but the gap is smaller than the gap between 70B and 32B models.
Can I run 70B on a laptop?
Not at usable quality. No consumer laptop ships with enough VRAM for Q4, and CPU only inference on 70B is measured in seconds per token. Run a 7B or 14B model locally instead, and access your 70B rig remotely over Tailscale.
What about an M4 Ultra Mac Studio?
The M4 Ultra with 128GB or 256GB unified memory runs 70B at Q4 or Q5 at roughly 8 to 12 tokens per second, and draws far less power. The trade off is price (£4,500 plus) and no upgrade path. For flexibility and value, a refurbished tower plus dual GPUs still wins on a pounds per token basis.
About Birmingham AV
Birmingham AV Ltd has sold over 87,000 items on eBay UK since 2017, with 24,756 buyer feedbacks at 98.9 percent positive. Every build ships with a twelve month warranty and is bench tested before dispatch. We are Companies House 12383651, VAT GB 348755066, based in Bromsgrove, Worcestershire, and we are one of the highest volume refurbished PC operations on eBay UK. If you are speccing a local LLM rig and want a sanity check on the parts list before buying, our listings pages carry direct contact for pre sales questions.