One router for pi and DeepSeek Harness

Every task, the right model.

It's a CTO for the work that matters, an engineer for the workload.

Everyday coding runs on a cheap, fast model. Hard problems get the smart one. When a provider stumbles, the router switches — automatically, mid-turn. And when a task is genuinely complex, the Smart tier plans and delegates it to Fast subagents — on pi and DeepSeek Harness alike. Your work never stops.

Multi-provider auto-switch keeps your task running No extra servers runs inside your Pi process Open source MIT zero runtime deps
pi-shift-router — in pi
status bar · toasts · auto
router: MiniMax-M3
tier: fast
mode: auto
judge 🧭 judging…
window 4/5 fast
🦾 [MiniMax-M3] → fix the failing test
🧭 judging…
🧠 [kimi-k3] ← upgraded for the architecture question
⚠️ MiniMax-M3 429 → switching to deepseek-v4-flash — retry in 1m
🦾 [deepseek-v4-flash] ← same-tier failover (v0.6.0)
See how
The family

One routing idea, two editions

Same LLM-judge routing, same fallback chains, same task-level orchestration — adapted to the agent you actually run.

pi-shift-router

v1.3.0
for pi-coding-agent

The original. Installs into pi in one command, routes every turn between Fast and Smart, and orchestrates complex tasks by delegating them to Fast subagents.

  • Zero runtime deps · ~409 kB
  • Task-level orchestration on by default
  • 369 unit tests

dsh-shift-router

v0.5.0
for DeepSeek Harness

The DSH edition. Plugs into the harness via cordis.patch.yml, routes through the harness's own agent pipeline, and configures live from the GUI "Shift-Router" card or /router config.

  • Rides the harness's agent/request pipeline
  • GUI config card + live commands
  • 109 tests · credential-free e2e

One codebase to read in an evening, two editions. Pick the one that matches your agent.

Quick Start

Up and running in three steps

Install it and nothing changes — routing only starts once you pick your models. Fully reversible, no lock-in.

  1. 01 / 03

    Install

    One command. pi registers the extension in your settings and loads it on the next launch — no rebuild, no config files to touch.

    Why it's safe to install
    • Zero dependencies — pure TypeScript — nothing new in node_modules, loads in milliseconds
    • Zero telemetry, zero backend — no data leaves your machine; everything runs inside your pi session
    • Open source, no lock-in — entire codebase readable in an evening; uninstall anytime and pi falls back to your default model
    pi — zsh
    $ pi install npm:pi-shift-router
    Resolving pi-shift-router@1.3.0...
    ✔ Downloaded and registered
    ✔ Added to ~/.pi/agent/settings.json
    ℹ Loaded on next pi launch — nothing else to do.
  2. 02 / 03

    Pick your models

    Run /router config and pick a Fast model and a Smart model — the pair you already use. Save to user or project scope.

    How to pick
    • Fast — the best value-to-quality model in your stack, not the cheapest (the Judge runs on it, so its judgment matters too)
    • Smart — a frontier model for the work that matters
    • Fallback chain — add 2–3 models per tier and they form a fallback chain for 429/5xx
    See recommended pairings
    pi-shift-router — Configuration
    / router config
    pi-shift-router — Configuration
    › 🦾 Fast — 0 model(s) (engineer: execution, daily coding)
    🧠 Smart — 0 model(s) (CTO: architecture, review, planning)
    🪄 Orchestration — auto (complex → Smart CTO delegates to Fast subagents)
    🎨 UX settings
    💾 Save & exit
    🚫 Discard & exit
    ↑↓ navigate · enter to select
  3. 03 / 03

    Watch it work

    /router status shows your setup; the next turn runs the first classification. From there it's automatic — tune later with /router quiet and /route-force when you want control.

    What the output means
    • Spend — how much each tier cost, and the baseline: what it would have cost without the router
    • Turns / upgrades / downgrades — how often the router escalated or settled
    • Cooldowns — models resting after a 429/5xx; they resume after backoff
    • Window — the last 5 classifications driving the downgrade gate
    pi — /router status
    / router status
    Tiers:
    🦾 Fast · deepseek-v4-flash (engineer mode)
    🧠 Smart · claude-opus-5 (CTO mode)
    Session:
    Mode: AUTO Enabled: ✅ Quiet: 🔊
    Orchestration: 🪄 auto (idle)
    Last audit: ✓ clean
    Turns: 12 (↑upgrade 1 · ↓downgrade 0)
    Manual: ✗ None
    Cooldowns: none
    Stats:
    Spend: fast $0.045 (9 calls) · smart $0.42 (3 calls) · total $0.465
    baseline: all-turns-on-smart (deepseek-v4-flash) → $3.21 · saved $2.74
    speed current=28 avg=24 tok/s
    Detail:
    Window: 5 entries (confidence: high=3 mid=2 low=0 none=0)
    Judge: 🧭 deepseek-v4-flash
    Tokens: total 12,847
    Config: ~/.pi/agent/pi-shift-router.json
How it works

One simple, elegant idea

Every task has a difficulty. Every model has a price. Shift Router matches the two — automatically.

Think of it like a team: the engineer handles the day-to-day fast and cheap; the CTO takes over the whole turn when the work matters. Same task, two minds — the router decides which one you need before you finish typing. And when the CTO's job is big enough, it doesn't just do it alone: it delegates to Fast subagents and reviews their work.

Routine work stays cheap

The judge only upgrades when a task actually needs depth — everyday coding keeps running on your budget model, and expensive flagships are reserved for what matters.

The deciding is nearly free

One tiny classification call — a few thousand tokens at your fast-tier price, 200ms–2s. The savings from skipping unnecessary smart turns dwarf it.

Keeps working when a provider hiccups

If a model 429s or times out, it cools down and the next healthy model in the same tier takes over — mid-turn, no action needed from you.

Fast

Engineer

Your everyday engineer. Fast on routine work — writing code, running tests, fixing bugs — so you don't waste a flagship model on a rename or a refactor.

  • Day-to-day coding and fixes
  • Tests and mechanical changes
  • Routine, low-stakes work

Smart

CTO

The CTO — sets direction, corrects course, reviews results, and takes on the hard problems itself: architecture, design review, security, multi-step plans, irreversible changes. High-stakes turns don't get dropped.

  • Architecture and design decisions
  • Reviews that need real judgment
  • Risky, ambiguous, or deep tasks
LLM Judge

A tiny call decides

One small classification — by the fast-tier model itself — reads your request and picks fast or smart. No heavy reasoning, no delay you'll notice.

judgeTimeout: 5000
Downgrade gate

No flip-flopping

Upgrades are instant when you need depth. Downgrades wait for a clear trend — fast votes in 5 of the last 5 classified turns — so the router never flickers between models mid-flow.

window: { size: 5, threshold: 0.6 }
Runtime failover

Failures, handled

When a provider rate-limits or errors, that model cools down with exponential backoff (5xx from 1m; 429/quota from 16m, capped at 6h) and the next healthy one in the same tier takes over — you keep working through the outage.

cooldown: 1m → 4m → 16m → 1h → 4h → 6h
Task-level orchestration

The CTO doesn't work alone

When the Judge says the task is complex and orchestration is in auto mode (default), the Smart tier runs as a CTO: it plans the work, delegates implementation to Fast engineer subagents, reviews each result, and iterates until the work is clean. Simple tasks never orchestrate — they stay on the plain router.

  • Plan → delegate → review → iterate → accept
  • Post-turn acceptance audit — grounding · goal alignment · delivered quality
  • Fresh-context workers: small, fast, cheap (~$0.004 vs ~$0.06 per fork)
  • Hard caps: maxRounds 3 · escalationThreshold 2 · Last audit in /router status
  • Requires pi-subagents; degrades gracefully without it
Use cases

Three setups that cover most teams

Pick the one closest to your situation — everything is reversible.

01

Multi-provider resilience

Run 2–3 providers per tier. When one rate-limits or stalls, the next healthy model in the same tier takes over mid-turn.

02

Single-provider tier ladder

One provider, two tiers. One bill, one rate-limit pool. Simplest setup — pick a budget model for routine work and the same vendor's flagship for hard turns.

03

Local + cloud hybrid

Fast tier on a quantized local model. Smart tier in the cloud. Privacy on the routine turns; frontier quality when it matters.

FAQ

Questions, answered

Straight from the README.

What if I don't configure any models?

Nothing changes. Both tiers start empty — the router does nothing and pi keeps using your default model. Routing only begins once you run /router config and pick your models.

Does this actually save money?

Yes. Routine turns stay on your fast, value-priced model instead of burning a flagship on every request. The judge itself costs a few thousand tokens at your fast-tier price — the savings dwarf it.

Is it safe to try?

Completely. The router does nothing until you configure it, it is pure TypeScript with zero runtime dependencies, and /router off disables it instantly. No lock-in.

What exactly does the smart tier do?

It's the CTO role — it sets direction, corrects course, reviews results, and takes on the hard problems itself: architecture, design review, security, multi-step plans, irreversible changes. High-stakes turns don't get dropped. It doesn't just advise or review; it writes the code and does the work.

Does the Judge add noticeable latency?

No. The classification call is a few thousand tokens and takes about 200ms–2s. You'll see 🧭 judging… in the status bar during the call — most users never notice it.

What if my primary model 429s or times out?

You keep working. The failed model enters an exponential-backoff cooldown — 5xx starts at 1m (1m → 4m → 16m → 1h → 4h… capped at 6h), while a failover-worthy 4xx (429 rate limit / quota) skips the first tiers and starts at 16m, because client-side limits usually outlive server blips. The next healthy model in the same tier takes over automatically — even mid-turn. A successful response clears the cooldown.

What is task-level orchestration?

When the Judge says a task is complex, the Smart tier runs as a CTO: it plans the work, delegates implementation to Fast engineer subagents (fresh context, model pinned from your Fast tier), reviews each result, iterates with concrete feedback, and does a final acceptance pass. Simple tasks never orchestrate — they stay on the plain router. Orchestration is on by default (auto mode); /router orchestrate off disables it. It needs the pi-subagents extension — without it, complex tasks run directly on the Smart tier, exactly as before.

Does orchestration cost more?

No — that's the point. The Smart tier plans and reviews; Fast subagents implement, each in a fresh context that's small and cheap: ~$0.004 for a narrow task vs ~$0.06 for an inherited 176k-token fork. Hard caps keep it bounded: maxRounds 3 and escalationThreshold 2. If a task turns out simple, the machinery never engages.

What keeps an orchestrated turn honest?

Two safety nets. The convergence protocol requires every re-delegation to carry a structured failure report — what failed, where, the exact acceptance test to re-run — and re-sending identical feedback is forbidden (the CTO takes the phase over). On top of that, a post-turn acceptance audit (v1.3.0) runs after every orchestrated turn: deterministic checks always run (every worker reported back, the final message carries a CTO summary, whether a hard cap was hit), plus an optional small fast-tier LLM audit that checks grounding, goal alignment, and delivered quality against your original request. Findings never block the finished turn — they surface as console.warn + toast and in /router status as "Last audit". Toggle with orchestration.audit.enabled (default on).

Can I force a specific model for one turn?

Yes. /route-force <tier> pins Smart or Fast for the next turn; /route-force <provider>/<model> pins an exact model. /route-force auto clears the override.

Does it work with models from different providers?

Yes — mix freely. Each tier is an ordered list of {provider, model, priority} pairs, so your fast tier can be DeepSeek and your smart tier Kimi, all from one config.

Can I monitor how it's performing?

Yes — run /router stats to see per-tier spend and savings, window size and confidence distribution (high / mid / low / none), cumulative upgrade and downgrade counts, and tokens-per-second. It reports cost telemetry: how much each tier spent and a baseline — what the session would have cost if every turn ran on your Smart model (i.e. no router). Spend: fast $0.045 (9 calls) · smart $0.42 (3 calls) · total $0.465, baseline $3.21, saved $2.74. The status bar also shows live throughput: [🧠 kimi-k3 • 23 tok/s].

How do I tune the downgrade behavior?

Two thresholds: `threshold` (default 0.6) controls how many of the last 5 classified turns must weigh toward fast before a downgrade happens; `minConfidence` (default 0.5) sets the floor below which votes are ignored entirely. Raise `threshold` to 0.8 to stay on Smart longer. Upgrades are always instant — only downgrades wait.

How is it different from pi-bifrost and pi-smart-router?

Both are pi routers with the same goal. @tenchi4u/pi-bifrost decides with 7-step rules + history tricks across 4 tiers — more cases, but harder to reason about; it saves subscription quota rather than dollars, and its circuit breaker can switch tiers on failure. pi-smart-router runs a 12-step local pipeline (no LLM) with 3 tiers including a local one, estimates cost with a formula, and needs a local DB + ML model (~2.5 MB plus downloads) and 15+ env vars. Shift Router keeps the decision to one readable LLM prompt (auditable plain text, JSON-mode enforced), two tiers with one mental model, orchestration (the Smart CTO delegates to Fast subagents), dollars-saved cost telemetry, same-tier failover with a cooldown shared with the judge, and ~409 KB with zero deps. CC-Switch is a different category entirely: a desktop app for manually switching provider configs between sessions — it coexists fine.

Is there a version for DeepSeek Harness?

Yes — dsh-shift-router (v0.5.0), a DSH adaptation of the same router. It plugs into the harness via cordis.patch.yml, routes through the harness's own agent/request pipeline, and configures live two ways: a full GUI card (Settings → Plugins → Plugin configuration — the "Shift-Router" card, with both tier chains editable) and an interactive /router config editor (get/set/unset/diff). Same capabilities as pi: LLM Judge, cache-aware routing, runtime failover, task-level orchestration, cost telemetry — with orchestration caps enforced by the plugin, not just prompted. 109 unit tests plus a credential-free headless e2e. Install: git clone https://github.com/green-dalii/dsh-shift-router && npm install && npm run build, then dsh plugin --profile web add <path>.

On DeepSeek Harness, are subagents routed too?

No. Subagents spawned by the subagent tool carry origin === 'subagent' and keep their pinned model — the router only drives top-level agents. During orchestration, the CTO delegates to Fast workers whose model is pinned by deployment configuration (tool-subagent agentOptions), so workers always run at the Fast tier's price.