Measured on our own hardware
The 16 GB Local AI Playbook
Every number here is measured on one real card — an NVIDIA GeForce RTX 5060 Ti with 16 GB of memory — not estimated from a spec sheet. We own exactly one consumer GPU, so instead of guessing at hardware we don't have, we go deep on the tier most people actually run: what fits in 16 GB, at what speed, at the quant a buyer would use.
Last measured August 4, 2026 · driver 580.173.02 · Ollama 0.23.2
What fits your card?
Speed is measured on the 16 GB card above. Whether a modelfits is mostly its memory footprint, which travels between cards far better than its speed does — so filter by the VRAM you have and see which of these measured models fit inside it.
The measured ladder
| Model | Params | Quant | VRAM used | On GPU | Generation | Prompt | Cold load |
|---|---|---|---|---|---|---|---|
Gemma 3 · 1Bollama pull gemma3:1b | 999.89M | Q4_K_M | 1.3 GB | 100% GPU | 265 tok/s | 7996 tok/s | 1.73s |
Qwen2.5 · 3B Instructollama pull qwen2.5:3b-instruct | 3.1B | Q4_K_M | 3.1 GB | 100% GPU | 176 tok/s | 9761 tok/s | 1.59s |
Gemma 3 · 4Bollama pull gemma3:4b | 4.3B | Q4_K_M | 4.6 GB | 100% GPU | 120 tok/s | 3886 tok/s | 3.04s |
Llama 3.2 · 3Bollama pull llama3.2:3b | 3.2B | Q4_K_M | 4.8 GB | 100% GPU | 175 tok/s | 9020 tok/s | 1.78s |
Qwen2.5 · 7B Instructollama pull qwen2.5:7b-instruct | 7.6B | Q4_K_M | 6.3 GB | 100% GPU | 88 tok/s | 4956 tok/s | 2.14s |
Qwen2.5 · 14Bollama pull qwen2.5:14b | 14.8B | Q4_K_M | 13 GB | 100% GPU | 45 tok/s | 2549 tok/s | 3.69s |
No measured model fits that budget on its own.
How these were measured
- One NVIDIA GeForce RTX 5060 Ti (16 GB). Generation rate is eval tokens ÷ eval time from the Ollama API, taken as the median of 3 timed runs of 200 tokens after a warm-up.
- VRAM used and “on GPU” are what
ollama psreports for the resident model — the model's real footprint, and whether all of it sits on the GPU or spills to the CPU (which is the moment speed falls off a cliff). - Params, quantisation and architecture are read from
ollama show— the model file's own metadata, not typed by us. Everything shown is Q4_K_M-class weights, the quant most people actually run. - Speed is specific to this card; your tokens/sec will differ on other hardware. Footprint and fit travel much better — that's why the filter above keys off memory, not speed.