larik v0.3.0

Guide

Getting started

To build and install larik into /usr/local/bin:

shell
./scripts/install.sh

The installer uses sudo when needed. Set PREFIX or BINDIR to install elsewhere, for example:

shell
PREFIX="$HOME/.local" ./scripts/install.sh

To build without installing:

shell
go build -o larik ./cmd/larik
shell
export ANTHROPIC_API_KEY=...   # or OPENAI_API_KEY / GEMINI_API_KEY
./larik

After installation, run larik from any directory.

The first time you run larik in a terminal with nothing configured, a setup wizard walks through connecting a model. It finds a running Ollama or LM Studio server and any API keys in your environment, tests the connection, lists the provider's models (with Ollama, it can also download a recommended model, or any model you name, with a progress bar), and saves your choice to ~/.config/larik/config.json or private project settings under ~/.config/larik/projects/. Run /connect at any time to add another provider.

With no --model and no saved default, Larik picks the default model of the first provider whose key is set:

Key Default model
ANTHROPIC_API_KEY anthropic/claude-opus-5
OPENAI_API_KEY openai/gpt-5.5
GEMINI_API_KEY gemini/gemini-3.8-flash
shell
./larik --model openai/gpt-5.5
./larik --model ollama/qwen3-coder
./larik --model openrouter/anthropic/claude-sonnet-5
./larik -p "summarize this repo"                  # headless, text output
./larik -p --output json --mode yolo "run tests"  # one JSON event per line
./larik -c                                        # continue the last session here
./larik --resume 20260926-2358                    # resume by id prefix
./larik -c --fork                                 # branch the last session instead of appending
./larik serve                                     # HTTP + SSE API (see Server mode)
./larik bench --models worker,anthropic/claude-haiku-4-5  # compare models on self-checking tasks (see Mixing cheap and strong models)

ChatGPT plan (Codex models)

To use Codex models on your ChatGPT plan instead of an API key, run /connect codex (or pick "ChatGPT (Codex)" in the first-run wizard). larik opens the ChatGPT sign-in page in your browser and receives the result on localhost:1455, the same OAuth flow the Codex CLI uses. It keeps its own sign-in in ~/.config/larik/chatgpt-auth.json, readable only by you, refreshes it before it expires, and never touches the Codex CLI's login. Sign out from /providers (select codex, press d twice).

For NVIDIA's hosted model catalog, run /connect nvidia-nim and enter a key, or set NVIDIA_API_KEY; for example, select an available nvidia-nim/<model-id> after connection. NVIDIA's developer access is for prototyping, research, development, and testing. Self-hosted NIM can use a custom OpenAI-compatible provider instead. Gemini and Claude in Larik use API keys, not Gemini CLI or Claude Code subscription credentials.

shell
./larik --model codex/gpt-6-luna   # the fast, affordable one; good for testing

Local models with Ollama

Ollama loads models with a small context window by default (4096 tokens), and Larik's system prompt and tool definitions alone use most of that. Larik talks to Ollama through its native chat API and asks for a 32768-token window (num_ctx), or the model's own maximum if that is smaller, so no server setting is needed.

A larger window uses more memory for the model's KV cache. To change it, set context_length for the provider in ~/.config/larik/config.json:

json
{ "providers": { "ollama": { "context_length": 65536 } } }

Larik reads back the window Ollama actually loaded, uses it for the context percentage and compaction, and warns when a request fills it. It also warns if the chosen model can't call tools. Small models such as qwen3:4b handle simple read/edit/test loops but can be unreliable with delegation.

Generated from README.md · section “Quick start”. Edit that file to change this page.