larik v0.3.0

v0.3.0 · written in Go · one binary

One terminal agent.
Every model.

Larik plans with your strongest model, hands the searching, reading and routine edits to cheap or local ones, and runs shell commands inside an OS sandbox by default.

go build -o larik ./cmd/larik && ./larik

Anthropic, OpenAI, Gemini, your ChatGPT plan, Ollama and any OpenAI-compatible endpoint, in one session.

  • Anthropic
  • OpenAI
  • Google Gemini
  • ChatGPT plan
  • NVIDIA NIM
  • Ollama
  • LM Studio
  • OpenRouter
  • Groq
  • DeepSeek
  • xAI
  • Mistral
  • Together

Model routing

Pay frontier prices
only for the thinking.

Routing guide
  1. anthropic/claude-sonnet-5main · plans, decides, reviews
    $0.031
  2. gemini/gemini-3.8-flashexplore · 9 tool calls, search and read
    $0.002
  3. ollama/qwen3-coderworker · edits and tests, in a worktree
    free
  4. anthropic/claude-sonnet-5main · reviews the diff, merges
    $0.007
  5. fallback readyrate limit → next model, before output starts

A budget that holds.

Warn at 80%, stop before the request that would cross the cap. Local and plan models count as free.

json
{
  "model": "anthropic/claude-sonnet-5",
  "roles": {
    "worker": "ollama/qwen3-coder",
    "explore": "gemini/gemini-3.8-flash"
  },
  "fallbacks": { "worker": ["gemini/gemini-3.8-flash"] },
  "budget": { "session_usd": 2.00, "warn_at": 0.8 }
}

Presets in one command.

/routing lists every connected model with prices and offers Balanced, Cheapest, or Local and plan first.

  • Fallbacks that don't waste a turn
  • A leash for small models: worktrees, turn caps, loop detection
  • larik bench measures pass/fail, cost and time per model

The agent loop

Read. Edit. Verify.
In a sandbox, by default.

Your models
Your machine
Your terminal

The first-run wizard finds a running Ollama or LM Studio server and any API keys in your environment, tests the connection and can download a recommended model with a progress bar.

  • A usable context window. Larik talks to Ollama natively and asks for a 32K window instead of the 4K default, then reads back what actually loaded.
  • Honest warnings. It tells you when a request fills the window or the model can't call tools.
  • Your ChatGPT plan. /connect codex signs in with the same OAuth flow as the Codex CLI.

✻ Welcome to larik · let's connect a model

› Ollama ● running · 6 models

LM Studio not running

Anthropic ● ANTHROPIC_API_KEY found

pulling qwen3-coder 62%

Set up a provider →

Designed around the loop

One static binary.
Small choices that keep it fast.

Cache-stable prompts

The system prompt and tool set stay fixed within a context and the transcript is append-only, so every request after the first reads from the prompt cache.

Reasoning kept per model

Switch from Claude to GPT to a local model mid-session. Each model gets its own thinking replayed.

Clean compaction

When context fills, the history becomes one summary that names the transcript file, so the model can grep what it left out.

No runtime to install

Build it with go build, copy it anywhere, start it in milliseconds.

See how Larik compares with other harnesses →

Three front ends, one loop

Terminal, scripts
or your editor.

All three read the same event stream from the same agent loop, so they behave the same way.

Read the docs

Terminal UI

Command palette, model picker, @ mentions, nested subagent activity, a session sidebar on F2 and per-model cost.

shell
larik
larik -c            # continue here
larik -c --fork     # branch it

Headless

One-shot prompts for CI and scripts, as plain text or one JSON event per line.

shell
larik -p "summarize this repo"
larik -p --output json \
  --mode yolo "run tests"

HTTP + SSE server

Drive sessions from an editor or web UI. Loopback only, token auth, replayable event stream.

shell
larik serve
# {"url": "http://127.0.0.1:4096",
#  "token": "…"}

Bring your models.
Keep your setup.

Start with any key you already have, or none at all.