Model fit
Narrow down what a machine can plausibly run. Fit guidance and measured evidence are shown separately, with the evidence source on every row.

Tested here Measured profile: NVIDIA GeForce RTX 3060 Laptop GPU · 6 GB · 12th Gen Intel(R) Core(TM) i7-12700H · 64 GB RAM
| Model & quant | Fit | Speed | Evidence | Next |
|---|---|---|---|---|
| Qwen3.6-35B-A3B · UD-Q4_K_XL | GPU + experts on CPU Parameters: -ngl 99 -ncmoe 40 | 21.49 t/s MOE-005 | Tested here | Read review · |
Speeds come from llama-bench (-p 512 -n 128) and exclude tokenisation and sampling time — not real chat speed.
What this version covers
This page lists only configs measured here, each row carrying its evidence source and a link to the report. llmfit’s fit estimation is not wired up yet — device tiers without a profile show “Not tested” rather than a guess.
17 distinct configs · 23 valid measurements
“Configs” counts distinct configurations: repeating the same config several times counts once, so measurement count is never passed off as config count.