Two rules of thumb that hold on both pi and DeepSeek Harness: Fast tier — pick the best value in your stack: strong quality per dollar, because this tier also powers the Judge (deepseek-v4-flash is the default pick — 0731 quality near Opus 5/GLM-5.2 territory at the low end of the price table). Smart tier — keep it on a frontier cloud model; local Smart is impractical under ~80 GB of VRAM, and even at 96 GB a flat-fee token plan usually wins.
01
One subscription, many models. Use a token gateway (one key proxying every major provider) or reuse a coding-tool subscription you already pay for — most expose an OpenAI-compatible endpoint.
🦾 Fast
- deepseek-v4-flash
- gpt-5.6-luna (OpenCode Zen / Go)
- qwen3.6-flash
🧠 Smart
- gpt-5.6-sol (OpenCode Zen / Go)
- qwen3.8-max
- kimi-k3
Reusable subscriptions: OpenCode Zen (OpenAI-compatible endpoint), GitHub Copilot (BYOK — OpenAI + Anthropic), Cursor (Pro / Pro+ / Ultra pools), OpenAI Codex (bundled with ChatGPT Plus / Pro / Business).
Best for
Zero extra setup — one bill, no new account.
02
Fast tier on quantized local (q4-k-m / NVFP4 / MXFP4 / AWQ-int4 / 1–2 bit ternary). Once you have 64+ GB of VRAM or unified memory, Smart can run locally too. VRAM ≈ params × 0.6 at Q4_K_M.
| Setup | Fast 🦾 | Smart 🧠 |
| ≤ 32 GB | LFM2.5-8B-A1B (8.5B), granite-4.1-8b (8.8B), Qwen3.6-27B (dense, q4 ≈ 14 GB), gemma-4-26b-a4b-it (q4 ≈ 13 GB), Qwen3.6-35B-A3B (MoE 36B/3B, q4 ≈ 18 GB), Laguna-XS-2.1 (MoE 33B/3B, q4 ≈ 17 GB), Ternary-Bonsai-27B (1.58-bit ≈ 7 GB, laptop/phone) | Cloud frontier model |
| 32–128 GB | Fable-Fusion-711-NEO-MAX-MTP GGUF (27B, 2026-07, best post-trained), gemma-4-31B-it (q4 ≈ 16 GB), Laguna-S-2.1 (117.6B, q4 ≈ 59 GB — needs 64 GB+) | Fable-Fusion-711 base weights, or cloud frontier |
| ≥ 128 GB | Fable-Fusion-711-NEO-MAX-MTP GGUF (same pick — post-training beats raw size) | DeepSeek-V4-Flash (284 B MoE / 13 B active, UD-Q4_K_XL ≈ 155 GB — needs 192 GB+ unified memory; on 128 GB-class use 1–2 bit ternary) |
Sizes are quant (production, not fp16). The AxxB suffix on a MoE model means active parameters per token — it affects compute speed, not disk size; a GGUF/q4 stores every expert weight. Qwen3.6-27B scores 77.2% on SWE-bench Verified — the strongest open-weight model that runs on a 24 GB card. Below ~256 GB of unified memory, local Smart rarely beats a flat-fee token plan on quality-per-dollar; run local Smart only for privacy or air-gapped work.
Best for
Cost-sensitive or air-gapped work; cloud only sees the hard turns.
03
One provider, two tiers — simplest setup. Use this when you already pay for one vendor and don't want to juggle keys.
| Setup | Fast 🦾 | Smart 🧠 |
| Anthropic | sonnet-5 | opus-5 / fable-5 |
| OpenAI | gpt-5.6-luna | gpt-5.6-sol |
| Google | gemini-3.5-flash-lite | gemini-3.6-flash |
| Qwen (AliBaba) | qwen3.7-plus | qwen3.8-max |
| DeepSeek | deepseek-v4-flash | deepseek-v4-pro |
| Z.AI (GLM) | glm-5 / glm-5-turbo | glm-5.2 |
| xAI (Grok) | grok-4.5-fast | grok-4.5 |
Best for
Simplest path from zero to routing — just two model IDs.
04
Best-of-breed regardless of vendor. Default: deepseek-v4-flash for Fast, claude-opus-5 for Smart. Add fallbacks for resilience.
| Setup | Fast 🦾 | Smart 🧠 |
| DeepSeek Harness (DSH) | deepseek-v4-flash | deepseek-v4-pro |
| Lowest cost | deepseek-v4-flash | claude-opus-5 |
| Multi-provider fallback | deepseek-v4-flash + glm-5.2 | claude-opus-5 + gpt-5.6-sol + kimi-k3 |
| Flat-fee | opencode-go/deepseek-v4-flash | opencode-go/glm-5.2 |
| 1 M context, long repo | deepseek-v4-flash | gemini-3.6-pro / kimi-k3 |
| Multimodal | deepseek-v4-flash | claude-opus-5 (vision) / gemini-3.6-pro |
| Multilingual / Chinese-first | deepseek-v4-flash | qwen3.8-max |
| GDPR / Europe | deepseek-v4-flash (via OpenRouter) | mistral-medium-2604 |
Best for
Strongest model in each tier, regardless of who sells it.
Pricing and model availability change monthly — refresh with `curl -s https://models.dev/api.json | jq` to see the latest. Full per-provider catalog lives at models.dev.