v0.3.0 · written in Go · one binary
One terminal agent.
Every model.
Larik plans with your strongest model, hands the searching, reading and routine edits to cheap or local ones, and runs shell commands inside an OS sandbox by default.
go build -o larik ./cmd/larik && ./larik
Anthropic, OpenAI, Gemini, your ChatGPT plan, Ollama and any OpenAI-compatible endpoint, in one session.
- Anthropic
- OpenAI
- Google Gemini
- ChatGPT plan
- NVIDIA NIM
- Ollama
- LM Studio
- OpenRouter
- Groq
- DeepSeek
- xAI
- Mistral
- Together
Model routing
Pay frontier prices
only for the thinking.
- anthropic/claude-sonnet-5main · plans, decides, reviews$0.031
- gemini/gemini-3.8-flashexplore · 9 tool calls, search and read$0.002
- ollama/qwen3-coderworker · edits and tests, in a worktreefree
- anthropic/claude-sonnet-5main · reviews the diff, merges$0.007
- fallback readyrate limit → next model, before output starts
A budget that holds.
Warn at 80%, stop before the request that would cross the cap. Local and plan models count as free.
{
"model": "anthropic/claude-sonnet-5",
"roles": {
"worker": "ollama/qwen3-coder",
"explore": "gemini/gemini-3.8-flash"
},
"fallbacks": { "worker": ["gemini/gemini-3.8-flash"] },
"budget": { "session_usd": 2.00, "warn_at": 0.8 }
}Presets in one command.
/routing lists every connected model with prices and offers Balanced, Cheapest, or Local and plan first.
- Fallbacks that don't waste a turn
- A leash for small models: worktrees, turn caps, loop detection
larik benchmeasures pass/fail, cost and time per model
The agent loop
Read. Edit. Verify.
In a sandbox, by default.
✻ larik v0.3.0 · ~/code/api · seatbelt sandbox
/ commands · ? shortcuts · alt+p model · shift+tab mode
› fix the failing auth tests and tell me what broke
✻ Thought for 6s · ctrl+o to show
● todos(0/3 done)
● task › explore(find where session tokens are validated)
└ gemini/gemini-3.8-flash · 9 tool calls · $0.002
● read(internal/auth/token.go)
read 184 lines
● edit(internal/auth/token.go)
- if exp.Before(now) {
+ if !exp.After(now) {
● bash(go test ./internal/auth/...)
ok api/internal/auth 0.41s
A token expiring this very second still counted as
valid. exp == now is now expired; all 23 tests pass.
sonnet-5default 18% context · $0.04
OS sandbox, on by default.
Seatbelt on macOS, bubblewrap on Linux: write only to the project, no network but localhost, no way to rewrite git hooks. Sandboxed commands don't need to ask.
Claude Code–compatible.
Your existing hooks, Agent Skills, custom commands, subagent definitions, CLAUDE.md and .mcp.json work unchanged. Bring your setup; switch models.
Everything a serious coding agent needs,
built in, not bolted on.
Subagents in git worktrees
Parallel subagents with fresh contexts. Isolated ones work on their own branch for the main agent to review and merge.
Branch, rewind, undo
/fork a conversation, /rewind to before any prompt, /undo the last turn's file changes, resume any of them later.
MCP, with approval
stdio, HTTP or SSE servers with OAuth sign-in. Servers from shared project files wait for your approval.
Language servers
Diagnostics after every edit, so the model sees the type error it just introduced. Jump to definitions, apply code actions.
Web fetch and search
Pages as clean Markdown. Domain rules, blocked metadata addresses, cross-host redirects reported rather than followed.
Permissions you control
Four modes, from plan to yolo, on shift+tab. Allow and deny rules with globs; deny always wins.
Plan, then approve
In plan mode the model reads, then puts its plan to you. Longer work runs from a checklist in the sidebar.
A project-aware prompt
@ attaches files and images, ! runs a shell command, ctrl+r searches history.
Debug traces
--debug records every request as sent. larik trace opens a timeline that diffs each prompt against the last.
Make it yours
Your own status line, rebindable keys, vim mode for the prompt, light or dark themes.
Your models
Your machine
Your terminal
The first-run wizard finds a running Ollama or LM Studio server and any API keys in your environment, tests the connection and can download a recommended model with a progress bar.
- A usable context window. Larik talks to Ollama natively and asks for a 32K window instead of the 4K default, then reads back what actually loaded.
- Honest warnings. It tells you when a request fills the window or the model can't call tools.
- Your ChatGPT plan.
/connect codexsigns in with the same OAuth flow as the Codex CLI.
✻ Welcome to larik · let's connect a model
› Ollama ● running · 6 models
LM Studio not running
Anthropic ● ANTHROPIC_API_KEY found
pulling qwen3-coder 62%
Designed around the loop
One static binary.
Small choices that keep it fast.
Cache-stable prompts
The system prompt and tool set stay fixed within a context and the transcript is append-only, so every request after the first reads from the prompt cache.
Reasoning kept per model
Switch from Claude to GPT to a local model mid-session. Each model gets its own thinking replayed.
Clean compaction
When context fills, the history becomes one summary that names the transcript file, so the model can grep what it left out.
No runtime to install
Build it with go build, copy it anywhere, start it in milliseconds.
Three front ends, one loop
Terminal, scripts
or your editor.
All three read the same event stream from the same agent loop, so they behave the same way.
Terminal UI
Command palette, model picker, @ mentions, nested subagent activity, a session sidebar on F2 and per-model cost.
larik
larik -c # continue here
larik -c --fork # branch itHeadless
One-shot prompts for CI and scripts, as plain text or one JSON event per line.
larik -p "summarize this repo"
larik -p --output json \
--mode yolo "run tests"HTTP + SSE server
Drive sessions from an editor or web UI. Loopback only, token auth, replayable event stream.
larik serve
# {"url": "http://127.0.0.1:4096",
# "token": "…"}