Model routing
An agent spends most of its tokens reading: searching files, re-reading context, running routine edits. Those don't need your best model. Larik lets the main agent plan and review on a strong model while subagents do the volume work on cheap ones, across providers.
Run /routing to set it up. The wizard lists the models of every provider you've connected, with prices, and offers presets:
- Balanced: your current model stays in charge; the cheapest models it finds become
workerandexplore. - Cheapest: the lowest-priced capable model for every subagent.
- Local and plan first: Ollama, LM Studio or your ChatGPT plan before paid APIs.
Then you can adjust each role, add fallbacks and set a budget. The result is saved as plain settings, for example for a student with a small Anthropic budget, a free Gemini key and Ollama:
{
"model": "anthropic/claude-sonnet-5",
"roles": {
"worker": "ollama/qwen3-coder",
"explore": "gemini/gemini-3.8-flash"
},
"fallbacks": {
"worker": ["gemini/gemini-3.8-flash"],
"explore": ["anthropic/claude-haiku-4-5"]
},
"role_options": {
"worker": { "isolation": "worktree", "max_turns": 40 },
"explore": { "max_turns": 30 }
},
"budget": { "session_usd": 2.0, "warn_at": 0.8 }
}
-
Roles:
workerruns the built-ingeneral-purposesubagent,explorethe read-onlyexploresubagent.smartis a strong model the main agent can hand hard subproblems to, andcompactsummarizes the conversation when the context fills. An unset role uses the main model, so nothing changes until you set one. You can add roles of your own and use them in agent definitions (model: reviewer), in--modeland in/model. -
Per task: once a role has a model, the
tasktool gets amodelinput listing the roles with their prices. The main agent is told to send well-specified, mechanical work to cheap roles and keep design decisions, ambiguous debugging and final review for itself. Subagent rows show the model they ran on. -
Fallbacks: when a model fails before answering with a rate limit, an exhausted quota, an auth problem or an outage, Larik switches to the next model in its list and says so. It never switches once output has started. A model that failed is left alone for a minute (rate limits, server errors) or ten (a missing model, a bad key, no credit, a stopped server), so later requests don't pay for the same failure. Useful with free tiers that run out mid-session. Fallbacks can be keyed by role or by
provider/model. -
Keeping cheap models on a leash: small and mid-size models sometimes lose the thread. They create stray files, "clean up" by deleting things, or repeat one command forever. Three safeguards cover this, and the presets turn them on:
role_options.<role>.isolation: "worktree"runs that role's subagents in their own git worktree (see Subagents). Their edits come back as a branch for the main agent to review and merge, so a bad run never touches your checkout. It applies unless the task or agent definition chooses otherwise; outside a git repository the subagent works in place, with a notice.role_options.<role>.max_turnscaps the role's subagents (default 100).- Every subagent is stopped when it repeats the same tool call with the same result 4 times within 8 turns. The main agent is told the subagent got stuck and to check its changes. The main agent itself isn't stopped this way, since you can see it and interrupt it.
In the wizard,
wtoggles the worktree,ctoggles a minimal prompt (below) and+/-change the turn cap. From the command line:/routing worker.isolation=worktree worker.max_turns=30. -
A smaller prompt for cheap models:
role_options.<role>.context: "minimal"drops your global instruction file (~/.config/larik/AGENTS.md) and the skills index from that role's subagents, keeping only this project's ownAGENTS.md/CLAUDE.md. The presets set it forworkerandexplore: it's a smaller prompt with less to misread, and it stops a small model from loading a skill that has nothing to do with the task. From the command line:/routing worker.context=minimal(or=empty for the full prompt). -
Budget: Larik warns once at
warn_at(default 80%) and stops before the next request once the session, subagents included, has spentsession_usd. Raise it with/routing budget=5. Local and plan-included models count as free. -
Comparing models before you trust one:
larik bench --models worker,anthropic/claude-haiku-4-5,ollama/qwen3-coderruns a few small, self-checking coding tasks (fix a failing test, implement a stub, rename a symbol across files) against each model in its own throwaway directory, then reports pass/fail, cost and time — no separate judge model,go testis the check. Use it to see whether a role you're about to add is actually good enough, not just cheap. Seelarik bench -h. -
Compaction: the
compactrole is usually best left unset. With prompt caching, the main model re-reads the conversation at the cache price (for Opus, $0.50 per million tokens), which can cost less than a cheap model reading it all uncached. -
Quick edits:
/routing worker=groq/llama-4-scout,/routing explore=(back to the main model),/routing budget=(no cap).