larik v0.3.0

Internals

Extensibility

Larik has six extension mechanisms. All of them plug into the same two seams: they either add tools to the registry, or run at points in the loop.

Mechanism Seam Configured by Package
Subagents the task tool .larik/agents/*.md, .claude/agents/*.md subagent, worktree
MCP servers tools named mcp__server__tool mcp_servers, .mcp.json mcp
Skills system-prompt index + the skill tool + /name expansion SKILL.md folders skills
Hooks lifecycle points in the loop hooks in settings hooks
Language servers edit/write results + the lsp tool lsp in settings, or on PATH lsp
Web web_fetch, web_search tools env keys, web in settings web

Subagents

The task tool (task.go) lets the model delegate. A subagent is a full Agent, created with parent.Spawn (spawn.go:50).

What the child gets:

  • A fresh context (it can't see the conversation, so the prompt must be self-contained).
  • A system prompt: the definition's body + a footer asking for a complete final report + the shared env/instructions/skills context.
  • The parent's tools filtered by the definition's tools list, minus task, task_wait and task_stop (no nesting).
  • A model chosen by the task's model input, else the definition's model, else the parent's (Tool.model). Either can name a role (worker, explore, …) or a provider/model. The built-ins default to the worker and explore roles, and an unset role inherits. The task tool reads roles through Tool.Roles on every request, so /routing changes apply without a restart, and offers the model input only once some role has a model. Unmapped opus/sonnet/haiku aliases resolve only on Anthropic; elsewhere the child falls back to the parent's model with a notice.
  • At most 100 model turns, or the role's max_turns. When the cap is hit, or the loop guard (loop.go) sees the same calls with the same results 4 times in 8 turns, the parent gets an error result telling it to check the child's changes, followed by the child's last message.
  • Its own worktree when its role has isolation: "worktree" and neither the task nor the definition sets isolation. Only inside a git repository; elsewhere it works in place, with a notice.
  • A smaller shared context when its role has context: "minimal": Tool.MinimalContextFunc (app.go) in place of ContextFunc, built from agent.MinimalContextSections (prompt.go) — the env block and this project's own AGENTS.md/CLAUDE.md, without the user's global instruction file or the skills index.

What it shares with the parent: the permission checker (same mode and rules), hooks (with SubagentStop instead of Stop), the checkpoint store, language servers and the sandbox. Its usage is added to the parent's totals, priced for the model that served each request (UsageInfo.Model). A transcript is saved under <session>-agents/. The parent's session budget covers it too: checkBudget (budget.go) walks up to the root agent, whose cost already includes every subagent's spend.

What comes back: only the child's final text message, truncated to the tool output cap. Tool calls and permission prompts are forwarded as events labelled Agent: "explore: find auth", so the UI can nest them, but they never enter the parent's context.

sequenceDiagram
    autonumber
    participant P as Parent Agent
    participant T as task tool
    participant W as worktree
    participant C as Child Agent
    participant F as Front end
    P->>T: tool_use task{subagent_type, prompt, isolation?}
    T->>T: agent.FromContext(ctx) → parent, emit
    opt isolation = worktree
        T->>W: Create(repo): git worktree add -b larik/task-xxxxxx at HEAD
        W-->>T: path, branch, base
        Note over T: cwd, perms.WithCwd, sandbox.ForWorktree,<br/>no checkpoints, no lsp tool
    end
    T->>C: parent.Spawn(opts).Run(ctx, prompt)
    loop child turn
        C-->>T: EvToolStart / EvToolEnd / EvPermission / EvNotice
        T-->>F: forwarded with Agent label
        C-->>T: EvUsage
        T->>P: AddUsage(model, usage)
    end
    C-->>T: EvAssistant (final text), EvDone
    opt worktree
        T->>W: Finish: commit leftovers, then remove if no commits or keep
        W-->>T: report: branch, files, how to diff and merge
    end
    T-->>P: tool_result = final text (+ worktree report)

task reports itself as read-only. The call itself changes nothing (every tool call inside the child is checked on its own), and read-only tools run in parallel, so several task calls in one response run as parallel subagents.

Worktree isolation

With isolation: "worktree" (or isolation: worktree in the definition), the child works in its own checkout (worktree.go):

  • Created under ~/.local/share/larik/worktrees/ from the current HEAD, on branch larik/task-xxxxxx. Uncommitted changes in your checkout are not included.
  • The child's cwd, relative paths and path rules point into the worktree. Writes elsewhere ask for permission like any path outside cwd.
  • The sandbox is re-derived with ForWorktree: writable are the worktree and the repo's shared .git (so commits work), while .git/hooks, .git/config and the other git files that choose which code git runs stay read-only.
  • On finish, leftover changes are committed (with --no-verify, falling back to a Larik identity). No commits → worktree and branch are deleted. Otherwise both are kept and the result tells the parent how to review (git diff base...branch) and merge.
  • Git operations on the shared repo are serialized with a per-repo lock, so parallel tasks can start and finish safely.

Background tasks

With run_in_background: true, task calls parent.StartBackground (background.go:78) and returns bg-N immediately. At most 8 run at once.

sequenceDiagram
    participant M as Model
    participant A as Agent
    participant B as background hub
    participant F as Front end
    M->>A: task{run_in_background: true}
    A->>B: StartBackground(label, runChild)
    A-->>M: "Started background task bg-1…"
    Note over B: child runs on its own context,<br/>events go to Agent.Background()
    B-->>F: EvToolStart / EvPermission (labelled)
    B-->>F: EvTaskDone{bg-1, status, result}
    alt agent mid-turn
        A->>A: next loop iteration: takeNotifications()
        A->>M: task-notification appended to the next request
    else agent idle
        F->>A: RunNotifications(ctx)
        A->>M: new turn: task-notification + "Background task results arrived…"
    end

The model can also block with task_wait (all tasks or specific ids) or cancel with task_stop. Background tasks survive Esc on the foreground turn; /tasks stop or quitting stops them. In -p mode Larik exits only after every task has finished and been delivered.

Definitions

Files are <name>.md with YAML frontmatter (name, description, tools, model, isolation) and a body that becomes the system prompt, the same format as Claude Code. Discovery (defs.go), lowest to highest precedence: ~/.claude/agents, ~/.config/larik/agents, then .claude/agents and .larik/agents in each directory from the repo root down to cwd. Built-ins: general-purpose (all tools) and explore (read, grep, glob, lsp).

MCP servers

mcp.go uses the official Go SDK over stdio, streamable HTTP or SSE.

  • NewManager(cfg).Start() in app.Setup connects every enabled, approved server in the background, so startup isn't blocked.
  • Manager.Registry is the agent's LoadTools: at the first prompt of a context it waits (bounded by the connect timeout) and returns built-in tools plus all connected servers' tools, sorted by name for a stable prompt prefix. Servers still starting, failed, or unapproved are reported as notices.
  • Each server tool becomes a tools.Tool named mcp__<server>__<tool> (sanitized, de-duplicated). It is read-only only if the server declares readOnlyHint: true and openWorldHint: false: an open-world tool could leak data through its inputs.
  • Text results pass through. Images (PNG, JPEG, GIF or WebP, up to 5 MB and 4 per call) go to the model with the result; other images, audio and binary resources are summarized.
  • Project-scoped servers need /mcp approve, pinned to MCPServer.Hash().
  • Resources (content.go): at connect, open notes whether a server has the resources capability; Registry then adds list_mcp_resources and read_mcp_resource (read-only). The agent attaches @server:uri mentions through the agent.MCPContent interface (IsServer, ReadResource), so the loop doesn't import the MCP package. A mention whose server isn't connected stays prose.
  • Prompts: listed at connect and run as /mcp__<server>__<prompt>, expanded in runWith by MCPContent.ExpandPrompt (arguments in order, the last takes the rest). A failing expansion ends the turn with an error before any request.
  • OAuth (oauth.go): streamable HTTP transports get an SDK AuthorizationCodeHandler. The SDK's SSE transport has no OAuth hook, so for sse servers the same handler sits in the HTTP client (oauthTransport): it adds the bearer token to the event stream's GET and each message POST, and on a 401 or 403 runs Authorize and sends the request once more (message bodies are buffered for the retry). Larik supplies the redirect (a local listener on 127.0.0.1), the browser step and token storage (mcp-auth/<server>-<url hash>.json, 0600, rewritten when a token refreshes). Background connections use a fetcher that fails with errNeedsLogin, giving StateNeedsAuth rather than a surprise browser window; Manager.Login reconnects with an interactive one. Without oauth.client_id the client registers dynamically. oauth is part of MCPServer.Hash(), so changing it needs re-approval for project servers.

Skills

skills.go implements Agent Skills with progressive disclosure:

  1. At startup, Discover scans the skill roots (user dirs, then .agents/skills, .claude/skills, .larik/skills from repo root to cwd; later overrides earlier).
  2. Only name: description lines go into the system prompt (Set.Index). The index is rebuilt at each fresh context (App.SystemPrompt calls Set.Reload), so a new or edited skill is picked up after /clear.
  3. When a task matches, the model calls the read-only skill tool to load the full body, then reads bundled files with normal tools.
  4. The user can run a skill directly with /name args; Set.Expand replaces the prompt with the body ($ARGUMENTS substituted).

Frontmatter flags: disable-model-invocation removes a skill from the index; user-invocable: false hides it from /.

Custom commands (commands/<name>.md under ~/.claude, the config dir, and .claude or .larik from repo root to cwd) are loaded into the same Set by parseCommand, with Skill.Command set. Their roots come first in Roots, so a skill overrides a command of the same name. Frontmatter is optional (the first body line stands in for a missing description); only commands with a description are model-invocable. argument-hint is read as a raw line, because Claude Code's [a] [b] style isn't valid YAML. When the user runs one, the agent replaces each !`cmd` with the bash tool's output, but only for commands permission.Decide allows without asking (inline.go); the rest are left out with a note. allowed-tools isn't honored, because a repository's file must not grant itself permissions.

Hooks

hooks.go runs shell commands at lifecycle points, in Claude Code's format, so existing scripts work.

flowchart LR
    SS["SessionStart"] --> UPS["UserPromptSubmit"] --> PC["PreCompact<br/>(when compacting)"] --> REQ(("model<br/>request"))
    REQ --> PTU["PreToolUse"] --> PERM{{"permission"}} --> RUN(("tool")) --> POST["PostToolUse"] --> REQ
    PERM -.-> NOTE["Notification"]
    REQ --> STOP["Stop / SubagentStop"]
    STOP -->|"block: continue"| REQ
    STOP --> END(("turn ends"))
    END -.-> SE["SessionEnd<br/>(on exit)"]

How a hook runs (hooks.go:191):

  • Matchers are case-insensitive regexes over the target (tool name, or startup/resume/clear, or compaction trigger). All matching hooks for an event run in parallel; results are combined.
  • Stdin is a JSON payload (session_id, transcript_path, cwd, hook_event_name, tool_name, tool_input, tool_response, permission_mode, …). LARIK_PROJECT_DIR and CLAUDE_PROJECT_DIR are set. Each hook has its own process group and a timeout (60 s default).
  • Exit 0: success; stdout may be JSON with decision, reason, continue: false + stopReason, systemMessage, or hookSpecificOutput (permissionDecision, updatedInput, additionalContext).
  • Exit 2: block; stderr is the reason and goes to the model (for PreToolUse it's a deny).
  • Other codes: non-blocking error shown to the user.

Prompt hooks ("type": "prompt", prompt.go) send the hook's prompt, with $ARGUMENTS replaced by the input JSON, to a model through the runner's Evaluator, which app.hookEvaluator supplies per session: the hook's model, else the explore role, else the session's model, with spend added to the agent's totals. The hooks package itself has no model code. The reply is parsed for {"ok": …, "reason": …} (prose or a code fence around it is tolerated, as is the older decision: approve|block), and ok: false becomes the same Result as exit code 2. Unreadable answers, timeouts (30 s default) and errors fail open with a message.

Inside the agent (agent/hooks.go), hook output enters the conversation wrapped in <hook-context source="…"> or <hook-feedback source="…"> tags, so the model can tell it apart from the user.

Language servers

lsp/manager.go starts built-in servers (gopls, typescript-language-server, pyright, rust-analyzer, clangd) when their binary is on PATH, lazily on the first matching file, once per project root.

The key integration is in tools.Env: after edit, multi_edit or write succeeds, Env.Diagnostics syncs the file to its server and waits (up to 3 s; 15 s while the server is loading) for fresh diagnostics, then appends new errors and warnings to the tool result. The model sees ERROR 5:19 undefined: gret in the same response that confirmed its edit, without running a build. read calls Env.Touch to warm up the server in the background.

Servers that support pull diagnostics (a diagnosticProvider in initialize, or a dynamic registration for textDocument/diagnostic) are asked with textDocument/diagnostic after the sync; a failed or unsupported pull falls back to waiting for publishDiagnostics.

The read-only lsp tool offers definition, references, hover, symbols, workspace symbols, diagnostics and code actions (actions.go). code_actions refreshes the file's diagnostics and passes those overlapping the line, so quick fixes come back. apply_code_action (not read-only; rules and modes treat it like edit) re-requests the actions on the current text, picks one by title, resolves it with codeAction/resolve if the edit is deferred, applies its WorkspaceEdit (changes or documentChanges; file creates, renames and deletes are refused) through tools.Env.RewriteFile (project boundary, protected paths, undo snapshot, atomic replace), then runs its command. The client answers the server's workspace/applyEdit only while that command runs; any other time it refuses, so a server can't change files on its own.

Web

web_fetch (fetch.go) converts HTML to Markdown (stripping scripts, nav, headers, footers), resolves links, pages long content with start/max_length, and caches for 15 minutes. web_search picks a backend (Brave, Tavily, SearXNG) from personal config or environment keys. Both ask before running; see security for the network protections.

Browser

The browser_* tools (browsercdp) drive Chrome over the DevTools Protocol with chromedp: pure Go, no CGo, no webview. browsercdp.Session launches Chrome on the first call (from context.Background, so a tool call ending never closes the browser) with a persistent profile in the data directory, and starts over if the user closed the window.

  • Snapshots. snapshot.js walks the visible DOM, including open shadow roots, and outlines headings, text, and interactive elements. Each interactive element gets a data-larik-ref attribute that later snapshots of the same document reuse; window.__larikFind looks refs up across shadow roots. Snapshots cap at 20k characters.

  • Input is real. Clicks scroll the element into view and dispatch mouse events at its center; typing focuses the field, clears it through the native value setter (so React and similar notice), and sends key events. After an input action, act waits up to 300 ms for a navigation to begin, and the snapshot then waits for readyState == "complete", retrying while the old document is torn down.

  • Tabs. Each tab is a chromedp context. Before each call syncTabs drops tabs the user closed and adopts ones the page opened, making a new one active. The first tab's context is the browser's, so closing it closes the page instead of cancelling the context.

  • History goes through Page.navigateToHistoryEntry rather than chromedp.NavigateBack, which waits for a load event that back/forward-cached pages never fire.

  • Dialogs would block the page and every later call, so a listener answers them (alerts accepted, the rest dismissed) and logs them with console messages and exceptions in a 200-line buffer per tab.

  • Screenshots. browser_screenshot captures the viewport, the full page (cut at 5000 CSS pixels, past which models shrink text beyond reading) or one element by ref (padded 8 px) as JPEG, and returns it in tools.Result.Images. The capture clip is in document coordinates with scale = 1/devicePixelRatio, so a Retina window gives the same image size as headless Chrome: one image pixel per CSS pixel. With labels, it takes a snapshot to assign refs, draws a tag and dashed outline for each ref in view, captures, and removes the overlay.

Generated from docs/extensibility.md. Edit that file to change this page.