A supervised coding agent for research software, built to run on your models, on your machines, with a receipt for everything it did.
Warning
Experimental v0.3.0 release. Clio Coder's behavior and interfaces may break or change without notice. Use version control, review proposed changes, and keep backups when operating on important repositories.
Clio Coder is a terminal coding agent for people who work on real scientific and HPC codebases: simulation kernels, data pipelines, numerical libraries, build systems that take twenty minutes and break in ways no cloud model has ever seen.
You bring the model. A local llama.cpp, Ollama, LM Studio, vLLM, or SGLang server; a cloud API; your ChatGPT or Claude subscription; or an Argonne Leadership Computing Facility inference gateway. Clio brings the harness around it: a terminal UI, twenty typed tools instead of an unrestricted shell, a fleet of bounded worker agents that can run across your whole cluster over SSH, durable sessions, and an integrity-sealed receipt for every run.
CLIO stands for Context Layer for Input/Output. Clio Coder is the interactive coding agent in IOWarp's ecosystem of agentic science, named for the Greek muse of history and built by the Gnosis Research Center at Illinois Tech.
| You are | Start here | |
|---|---|---|
| 🔬 | A researcher or developer who wants to use it | Install → Five-minute start → Bring your own model |
| 🤖 | An AI agent that just landed in this repository | For agents |
| 🛠️ | A developer who wants to contribute | For contributors |
Most coding agents ask you to trust a remote model with a shell. Clio makes a different bet: the harness should be strong enough that a 20B model running on your own GPU is useful, and honest enough that you can reconstruct every decision afterward.
The model never gets a shell by default. The tool surface is twenty typed tools organized into seven policy planes. Bash is default-deny, filtered through damage-control rules and per-project policy. Reads are bounded, writes are queued and reviewable, and every privileged call passes through one admission path that cannot be widened by the model asking nicely.
Local models are the design target, not a fallback. llama.cpp and similar
servers expose a single prefix-cache slot. Clio keeps the compiled prompt and
provider tool schemas byte-stable so that slot stays hot across turns and
sessions, bounds every tool result so one grep cannot blow the window, and
records a per-call cache verdict (hot, partial, cold, small) in the
session ledger so you can see when and why the cache went cold.
Work is delegated to bounded workers, not to one long context. The orchestrator dispatches focused agents with explicit tool profiles, call budgets, cost ceilings, and typed result contracts. A worker that cannot produce a conforming answer fails loudly instead of returning confident prose.
Your cluster is the runtime. Declare your nodes and the same worker protocol tunnels over SSH. A remote worker gets the same prompts, the same safety matrix, the same receipts. Placement is deterministic and pinnable, and capacity is governed by durable expiring leases that survive process death.
Everything is auditable. Each run seals a receipt covering token usage,
priced cost, tool activity, safety decisions, routing intent, the resolved
route, worker attestation, and result-contract conformance. clio-coder evidence
and /view verify check them; nothing in the audit trail is reconstructed
from prose.
Science is a first-class domain. clio-kit contributes MCP servers for HDF5, Slurm, ParaView, Pandas, NetCDF, FITS, Zarr, and ArXiv, and the shipped skills catalog includes scientific debugging and experiment-protocol guides.
- Node.js
>=22.19.0and npm - Linux or macOS. Windows is best effort until a stable release.
- At least one model target: a local OpenAI-compatible server, Ollama, LM
Studio, llama.cpp, vLLM, SGLang, a cloud API key, a ChatGPT or Claude
subscription login, an ALCF Globus account, or an installed
claudecommand
From npm:
npm install -g @iowarp/clio-coder
clio-coder --versionFrom source, pinned to this release:
git clone --branch v0.3.0 https://github.com/iowarp/clio-coder.git
cd clio-coder
npm run install:local
export PATH="$HOME/.local/bin:$PATH"
hash -r
"$HOME/.local/bin/clio-coder" --versionThe clone is pinned to v0.3.0, the release these instructions describe.
Without --branch you get the default branch, which is ahead of the release and
is not what the rest of this page documents.
npm run install:local verifies dependencies, builds the CLI, installs a
symlink at ${CLIO_CODER_BIN_DIR:-$HOME/.local/bin}/clio-coder, and runs the installed
CLI's structure repair so a fresh install passes plain clio-coder doctor with no
manual steps. It warns if the bin directory is not on your PATH and prints the
export line above; the line is a no-op for the shell that already has it, and
it is what makes a bare clio-coder resolve in the shell that does not. Put it in
your shell profile to keep it across sessions. The symlink executes
dist/cli/index.js, so re-run npm run build after editing TypeScript sources.
The last line runs the launcher by its full path on purpose. A bare clio-coder may
resolve to an older install earlier on your PATH, so it verifies whichever one
that is rather than the one you just installed; the installer warns when it
finds that shadowing, and names the other path.
Before you switch to the bare name, ask which file it reaches:
command -v clio-coder # expect $HOME/.local/bin/clio-coderComparing clio-coder --version against "$HOME/.local/bin/clio-coder" --version does not
answer that. Two installs of the same release print the same version, so the
versions agree while the name still resolves to the other one. The path is the
question.
To remove it, preview first:
clio-coder uninstall --dry-run
clio-coder uninstall --remove-binary --force
hash -rFor selective wipes that keep settings or credentials, use clio-coder reset. Full
details live in docs/installation-and-lifecycle.md.
Run Clio from the repository you want to work on and point one target at a
running model server. This example uses LM Studio; other local runtime ids
include ollama-native, llamacpp, vllm, and sglang.
cd /path/to/your/repo
clio-coder configure \
--id local-lmstudio \
--runtime lmstudio-native \
--url http://localhost:1234 \
--model your-model-id \
--set-orchestrator \
--set-fleet-default
clio-coder targets use local-lmstudio
clio-coder targets --probeOnce the target probes healthy, teach Clio about your project, try a headless turn, then open the TUI:
clio-coder context init # bootstraps local generated CLIO-CODER.md context from your real source tree
clio-coder run "Summarize this repository layout and identify the main entry points."
clio-coder # interactive terminal UIInside the TUI, /targets, /agents, /fleet, and /skill confirm what the
session can see. /help opens the interactive help center.
Clio treats models as named targets. A target is a runtime plus an endpoint plus a model plus credentials, and you can route interactive chat and fleet dispatch through different targets independently.
| Runtime id | Server |
|---|---|
llamacpp, llamacpp-anthropic, llamacpp-completion |
llama.cpp and llama-swap routers |
lmstudio-native |
LM Studio |
ollama-native |
Ollama |
vllm, sglang |
vLLM and SGLang |
lemonade, lemonade-anthropic |
Lemonade |
openai-compat, anthropic-compat |
Any OpenAI- or Anthropic-shaped endpoint |
openai, anthropic, google, groq, mistral, deepseek, openrouter,
bedrock, and alcf for Argonne's Sophia and Metis gateways over Globus
OAuth. See docs/alcf-provider.md for the HPC path.
You can drive Clio from a ChatGPT Plus/Pro or Claude Pro/Max subscription instead of an API key.
clio-coder auth login anthropic-max # Claude Pro/Max OAuth
clio-coder auth login openai-codex # ChatGPT Plus/Pro OAuth
clio-coder configure --id claude-sub --runtime anthropic-max --model claude-sonnet-5 --set-orchestrator
clio-coder configure --id chatgpt-sub --runtime openai-codex --model gpt-5.4 --set-orchestratorPick model ids from clio-coder models --target <id> after login.
Note
Connecting a Claude Pro/Max subscription over OAuth uses the same path as Claude Code. Using subscription credentials outside a vendor's first-party apps may not align with their terms of service. Enable at your own discretion.
Clio can also drive other coding agents as workers while keeping its own permission gating in front of them.
claude auth login # authenticate the official Claude CLI first
# Claude Code SDK worker, with enforced per-tool safety
clio-coder configure --id claude-sdk-worker --runtime claude-sdk --model sonnet --set-fleet-default
# claude -p subprocess worker, advisory permission-mode gating only
clio-coder configure --id claude-code-worker --runtime claude-code --model sonnet
# Google Antigravity subprocess worker, under your existing agy login
clio-coder configure --id agy-worker --runtime antigravity-code --model "Gemini 3.5 Flash (High)"The interesting configuration is a strong orchestrator with cheap local muscle, or the reverse.
clio-coder configure --id chatgpt-orch --runtime openai-codex --model gpt-5.4 --set-orchestrator
clio-coder configure --id claude-worker --runtime claude-sdk --model sonnet
clio-coder configure --id local-fleet --runtime lmstudio-native --url http://localhost:1234 \
--model qwen-7b --set-fleet-default
clio-coder targets profile claude-sdk claude-worker --model sonnet
clio-coder run --agent coder "Refactor src/engine/parser.ts"Full reference: docs/configuration-and-targets.md.
Fast scout agents work best when a small scout model is already loaded beside your main coding model on a local router. This is only safe when the combined weights, KV caches, context windows, and parallel slots fit in GPU memory. If the router spills into CPU RAM, both scout calls and main turns get slow.
Load both models manually on the target host, then point the orchestrator and
the scout worker profile at them. On llama.cpp routers, keep max_instances
at least as high as the number of models you want resident. Clio can see which
router instances are loaded and the router's instance limit, but current
llama.cpp router responses do not expose free VRAM, so confirming the loaded
set fits remains the operator's job. Workers on other nodes are unaffected.
There is one tool surface and one admission path. What changes is the autonomy
level, set in /settings or overridden for a single run with --autonomy.
| Level | Behavior |
|---|---|
read-only |
Inspection only. Every mutation and execution is denied. |
suggest |
Mutations are proposed and parked for your approval. |
auto-edit |
File edits proceed; execution and dispatch still gate. |
full-auto |
Approved classes proceed unattended, still inside damage-control rules. |
Notices name their mechanism so you always know who stopped a call:
[safety-net] for level-independent blocks, [approval] for parked calls,
[autonomy] for read-only denials, and [middleware] for hook diagnostics.
Workers can never exceed the orchestrator's authority. A dispatch request can
only narrow it, and reviewers and judges always run read-only. A && chain is
judged at its most restrictive recognized member and refused whole if any
member is unrecognized, and /tmp, /var/tmp, and /var/folders are scratch
while the rest of the system roots stay protected. Details:
docs/safety-model.md.
Clio's orchestrator delegates work to bounded workers. With a fleet declared, those workers run on other machines over SSH while every guarantee holds: one admission path, one autonomy matrix, one receipt chain.
flowchart LR
U["you"] --> O["orchestrator TUI"]
O --> P["execution plan<br/>hashed DAG, capacity waves"]
P --> A["admission<br/>leases, queue, cost ceiling"]
A --> L["local worker"]
A --> S1["ssh node: blade"]
A --> S2["ssh node: dragon"]
L --> R["receipts and evidence"]
S1 --> R
S2 --> R
R --> O
Declare nodes in settings.yaml and the implicit local node is always
present:
fleet:
nodes:
- id: blade
host: blade.example.net
maxWorkers: 2
residency: observe
- id: dragon
host: dragon.example.net
maxWorkers: 1Then clio-coder doctor runs a per-node preflight, clio-coder fleet list|run|status
drives and observes contracts, and clio-coder fleet drain|resume closes or reopens
durable dispatch admission. A drain preserves running work, rejects every new
execution start, and expires after one hour unless renewed. The /fleet
overlay shows nodes, profiles, bindings, and live runs. Nodes must share the
project filesystem at the same absolute path; hosts that do not fail admission
with a clear reason. Target URLs resolve on the node the worker runs on, so
localhost means that node's own inference server and there is no central
proxy.
Everything you need to reproduce it end to end, including a recorded multi-node demo script, is in docs/fleet-dispatch.md and docs/fleet-demo-runbook.md.
Routing is measured but conservative. Every dispatch records a joint decision over agent, target, model, runtime, and node, while shadow mode leaves the explicit route unchanged. Operators can activate only named read-only and quality roles, and only after the exact tuple has enough integrity-valid quality, reliability, cost, freshness, and decision-latency evidence:
routing:
activeRoles: [researcher, verifier, reviewer, judge]
activePostures: [quality, balanced]
agentAutomation:
activeAgentRoles: [] # stays advisory until exact agent/role pairs are namedManual pins and failover: none remain exact. Active mode fails closed when
no route is ready. agent: auto is separately bounded by recipe audience,
authority, tools, skills, result contract, locality, and approved governance;
changing from a read-only Scout phase to workspace editing requires an
authenticated plan approval or authority already granted by full-auto policy.
Clio loads a local CLIO-CODER.md as generated project context on every session.
clio-coder context init grounds a draft in the actual source tree, preserves an
existing handbook until an explicit replacement action, and can adopt existing
CLAUDE.md, AGENTS.md, GEMINI.md, Cursor, and Copilot context with
provenance and conflict reporting. CLIO-CODER.md is a gitignored runtime artifact,
not canonical repository documentation.
Alongside it, clio-coder context index builds a structural codewiki that the
code_nav tool navigates, so a model can find a symbol without reading half
the repository into its window.
Skills are reusable SKILL.md guides the model loads on demand. Clio
discovers them from per-user and per-project roots, including .clio-coder/skills
and cross-harness layouts such as .claude/skills and .codex/skills. A
skill's allowed-tools declaration is enforced at tool admission, and a skill
can ship executable RED-GREEN evals that clio-coder skills eval <name> runs
instead of trusting the prose.
This repository ships a curated catalog under skills/ with
provenance frontmatter, evals, and content hashes pinned in
skills/registry.yaml, so an installed copy verifies against its audited
source at activation. Nothing auto-loads.
clio-coder skills install context-handoff # copy into .clio-coder/skills
clio-coder skills list # confirm Clio sees itThe catalog includes find-skills, which routes
discovery through clio-coder skills search and clio-coder skills install. Install it
with clio-coder skills install find-skills --user so it outranks the community
skill of the same name that other installers drop into compat roots.
Long agentic runs decay: a requirement or a failed attempt is still in the transcript but no longer influences the next action. Clio's proactive task memory watches tool and lifecycle hooks, keeps a session task bank, and surfaces visible advisory reminders at trigger boundaries.
The default tier is rules-only and makes no model calls. An LLM memory tier is
opt-in through an independent background route. The action agent's prompt and
tool surface never change, /memory inspects the bank, and disabling
memory.intervention.enabled removes the whole mechanism. Durable lessons are
separate, scoped, evidence-linked, and managed through clio-coder memory list|propose|approve|reject|prune. Design notes:
docs/proactive-memory.md.
Clio Coder is experimental software in a soft beta. The current release is
v0.3.0, installable from npm as
@iowarp/clio-coder or
from source. Interfaces may still move between minor versions, and
model-specific behavior varies by target.
Release notes live in the CHANGELOG, the implementation detail
behind each entry lives in the commit history, and every release is gated by
the deterministic npm run ci:release suite.
| Problem | Try this |
|---|---|
clio-coder: command not found |
Run npm run install:local, then hash -r; confirm ${CLIO_CODER_BIN_DIR:-$HOME/.local/bin} is on PATH. |
| No model target is available | Run clio-coder configure, then clio-coder targets --probe. |
| Local model does not respond | Confirm the local runtime is running and the target URL is correct. |
| Cloud model auth fails | Check clio-coder auth status <target> and verify the API key or login flow. |
| A fleet node never gets work | Run clio-coder doctor; per-node preflight reports filesystem parity and target facts. |
| Source changes do not appear | Re-run npm run build; the linked CLI points at dist/. |
| State appears corrupted | Run clio-coder doctor, then clio-coder doctor --fix. |
When filing an issue, include the output of clio-coder --version, node --version, clio-coder doctor, and clio-coder targets. Redact secrets, private
prompts, logs, and proprietary code.
If you are an AI agent operating inside this repository or driving Clio as a tool, this section is the orientation you need.
Do not read broadly. Start from the codewiki, which indexes 927 source files.
Use code_nav in entries, path, or symbol mode before any wide read.
The indexed entry points are src/cli/index.ts, src/domains/agents/index.ts,
src/domains/components/index.ts, src/domains/config/index.ts,
src/domains/context/bootstrap.ts, src/domains/context/index.ts,
src/domains/dispatch/index.ts, and src/domains/eval/index.ts. When present,
read the local generated CLIO-CODER.md after that index-led orientation; it carries
the project-specific invariants and traps that are not obvious from the source.
Twenty tools in seven planes. Each plane is one policy unit covering action class, size posture, result schema, and concurrency rule.
| Plane | Tools | Posture |
|---|---|---|
| OBSERVE | read, grep, find, ls, code_nav, context, credential_present |
Read class, parallel, bounded by a truncation envelope |
| MUTATE | write, edit |
Write class, sequential, queued through the file-mutation queue |
| EXECUTE | bash, git, verify |
Containment posture; bash is default-deny, git is read-only inspection |
| ORCHESTRATE | dispatch, monitor, steer, tasks, ledger |
Dispatch class, sequential except read-only monitor |
| RETRIEVE | web_fetch |
Network read, parallel |
| INTERACT | ask_user |
Host-owned operator interview |
| ARTIFACT | artifact |
Plans, reviews, and reports as durable artifacts |
Every observation carries a truncation envelope with offload paths and next hints, so a large result is bounded rather than silently cut. Full parameter and payload reference: docs/tool-usage.md.
All topologies go through the same tool, admission chain, and autonomy matrix.
| Topology | Invocation | Semantics |
|---|---|---|
| Singular | task: "..." |
One assignment, with an optional separate briefing. |
| Parallel | tasks: [...] |
Fan out, wait for all, one summary. |
| Sequential | mode: "sequential" |
One at a time, stop reporting on timeout or abort. |
| Pipeline | mode: "pipeline" |
Each step receives the previous step's output as data. |
| Detached | detach: true |
Return assignment ids immediately and collect later. |
| Review gate | review: {reviewer?, max_cycles?} |
Builder, read-only verifier verdict, bounded revise loop. |
| Compete | mode: "compete", candidates: 2..4 |
N candidates in scratch worktrees, read-only judge, winner applied. |
Detached batches are durable, so collection survives session exit. Use
monitor with mode="collect" as the authoritative terminal barrier over a
batch, and collect every detached batch before final synthesis; mode="tools"
reports what a run actually executed. Alt+S sends a running attached dispatch
to the background as a detached batch, which review gates, compete, pipelines,
and time-boxed calls refuse with a reason.
Admission refuses a pairing it can prove impossible, such as a pinned read-only recipe aimed at mutation work, and flags the receipt where it is unsure rather than blocking. Claimed file changes and validation commands are checked against the run's own tool events, so a path the run never wrote cannot seal as done.
Every dispatch carries an ExecutionRole: builder, reviewer, judge,
researcher, verifier, or recovery. The role is typed on every request,
ledger envelope, receipt, route candidate, plan task, and route decision.
Route statistics never mix roles, and any attempt after the first is
recovery.
Workers answer typed terminal contracts, not trailing prose. A scout-report
carries findings as {claim, path, line}, and grounding is structural rather
than a regex over prose: a cited line must fall inside a span this run
actually read, so an estimated line number cannot pass as an observation. A
worker validates its own result and spends a bounded number of repair rounds
before failing the run; the orchestrator's sealed validation is the authority.
Review and compete gates default to the builtin verifier and never fall back
to the builder agent. A gate decider's postcondition is the gate result
contract, not its own recipe contract.
architect, coder, tester, verifier, debugger, documenter, scout,
researcher, provenance, and git-master. Each is a versioned frontmatter
recipe with an explicit tool profile, call and cost budget, and result
contract. Malformed custom recipes are quarantined with a diagnostic; malformed
builtins fail startup. Reference:
docs/built-in-agents.md.
This roster is compiled into the session prompt whenever the dispatch tool is
available, so pin an id from it. agent: "auto" baselines from task shape and
is a fallback, not a router.
clio-coder run "<task>" --json # one headless turn, JSONL events
clio-coder run "<task>" --agent coder # one explicit fleet agent, writes a receipt
clio-coder acp # serve ACP v1 over stdio for ACP frontends
clio-coder fleet run <contract> # run a fleet DAG contract
clio-coder fleet drain # pause new execution starts for up to one hour
clio-coder fleet resume # reopen durable dispatch admission
clio-coder evidence build|inspect|list # deterministic evidence artifacts
clio-coder eval validate|run|report|compare|gateDispatch can also delegate to external ACP agents while Clio mediates permissions. Clio implements the Agent Client Protocol so the engine stays decoupled from IDE frontends.
Every run seals a receipt. Receipt integrity is at v15 and covers normalized routing intent, the resolved route, worker attestation, priced cost, phase timing, tool activity, safety decisions, and result-contract conformance. Gate decisions are v2 artifacts that seal route correlation across agent, target, model family, runtime, and node, and they cross a staged durable boundary rather than being written directly.
A worker attests its protocol version, pid, process-group id, host, settings fingerprint, WorkerSpec digest, runtime, target, endpoint identity hash, wire model, effective tool signature, and bounded resource facts before any model call. Any drift from the approved identity kills the worker.
clio-coder trace reads the same store and now records interactive turns beside
dispatched runs, one event per tool call with its verdict.
Verify from the TUI with /view verify <runId>, or from the shell with clio-coder evidence inspect. See docs/observability.md.
Contributions are welcome, and the fastest way in is to fix something you hit while using it on your own research code.
flowchart TB
CLI["src/cli"] --> ENG["src/engine"]
TUI["src/interactive"] --> ENG
ENG --> TOOLS["src/tools<br/>20 typed tools, 7 planes"]
ENG --> DOM["src/domains"]
DOM --> DISP["dispatch<br/>plans, leases, routing, receipts"]
DOM --> CTX["context<br/>codewiki, compaction, CLIO-CODER.md"]
DOM --> PROV["providers<br/>runtimes, auth, catalog"]
DOM --> SAFE["safety<br/>damage control, policy"]
DISP --> WORK["src/worker<br/>bounded worker runtime"]
The largest indexed areas are src/domains (392 files), tests/contracts
(236), src/interactive (83), src/cli (48), src/tools (42), src/engine
(40), and src/core (35). Compile-time boundaries between domains are
enforced by a test suite, not by convention. Read
docs/architecture.md before adding a cross-domain
import.
npm ci
npm run dev # tsup watch build
npm run ci # the full local gateTargeted checks when the risk is narrower:
| Check | Command |
|---|---|
| Types | npm run typecheck |
| Style | npm run lint |
| Contracts | npm run test:contracts |
| Smoke flows | npm run test:smoke |
| Domain boundaries | npm run check:boundaries |
| Everything | npm run test |
Conventions worth knowing before your first PR: local imports end in .js,
tests use node:test, and any needs a tracking issue.
npm run ci:releaseThat runs typecheck, Biome, the skills pin check, the production build, the
contract, smoke, and boundary suites, and the check-release dist and package
audit. Live model validation is separate, manual, and opt-in, because no
deterministic suite can promise that every local model behaves identically:
CLIO_CODER_LIVE_SMOKE=1 \
CLIO_CODER_LIVE_TARGET=openai-compat \
CLIO_CODER_LIVE_RUNTIME=openai-compat \
CLIO_CODER_LIVE_MODEL=your-model \
CLIO_CODER_LIVE_BASE_URL=http://localhost:8080/v1 \
npm run test:live
CLIO_CODER_LIVE_SMOKE=1 npm run test:live -- --delegation # needs local opencode and copilot
npm run test:live-eval:fleet-dispatch # multi-node dispatch regressionBenchmarks against public suites live under benchmarks/:
npm run bench:swe # SWE-bench Lite
npm run bench:scicode # SciCodeRead CONTRIBUTING.md for setup, architecture invariants, branch and commit conventions, and the review rubric. Good first areas: provider adapters for a runtime you use (cookbook), skills for a scientific domain you know (catalog), and documentation gaps you hit during onboarding.
Security reports go through SECURITY.md, not public issues.
The full set lives under docs/, and clio-coder docs serves it
locally with interactive blueprints.
| Topic | Guide |
|---|---|
| Commands, slash commands, operating posture, keybindings, dispatch, verification, troubleshooting | commands-and-modes.md |
| Multi-node fleet dispatch: SSH transport, doctor preflight, placement, topologies, receipts | fleet-dispatch.md |
| Executable multi-node demo with a reviewer gate and receipt provenance walkthrough | fleet-demo-runbook.md |
| NDJSON parent-child protocols, watchdog timers, and exit status mapping | worker-dispatch-mechanics.md |
| Built-in agent recipes, discovery roots, frontmatter schema, dispatch admission | built-in-agents.md |
| Context window resolution, probe capabilities, token accounting, compaction, priming | context-engine.md |
| Proactive task memory, session task bank, intervention rules, handoff carrying | proactive-memory.md |
| Runtime targets, local model configuration, fleet profiles, auth | configuration-and-targets.md |
| Argonne ALCF Sophia and Metis inference targets over Globus OAuth | alcf-provider.md |
| Safety posture, default-deny Bash, project policy, damage-control rules, typed validation | safety-model.md |
| Source layout, compile-time boundaries, domain loading, runtime data flow | architecture.md |
| Reference for all 19 worker tools: parameters, payloads, error examples | tool-usage.md |
| Prompt envelope reuse, provider tool delivery, bounded tool results | prompt-envelope-and-tools.md |
| Implementing custom model runtimes and inference server integrations | provider-adapter-cookbook.md |
| Artifact browsing, receipt verification, dispatch diagnostics, observability routing | observability.md |
| Evidence directory structures, findings, operator-approved memory retrieval | evidence-and-memory.md |
| Local YAML eval suites, reports, comparisons, command evidence | eval-runner.md |
| Installation, upgrade, reset, uninstallation, configuration folders, permissions | installation-and-lifecycle.md |
| Every environment variable the runtime reads | environment-variables.md |
| Prompt and skill resources, extension manifests, portable share archives | extensions-and-sharing.md |
| Skills Hub marketplace discovery, install actions, publishing | skills-marketplace.md |
| Runtime model refresh, catalog sources, local and cloud model quirks | model-catalog.md |
| Active component snapshots and the experimental middleware hook contract | middleware-and-components.md |
| Advisory validation-contract patterns for scientific artifacts and HPC assumptions | scientific-validation.md |
Falsifiable Change Manifest templates, auditability, and clio-coder evolve |
evolution.md |
| Interface layout, palette, Unicode vocabulary, drawing choreography | tui-design.md |
| Source-first docs workflow, mapping matrix, alpha wording guidance | documentation-guide.md |
| Private context index determinism and target smoke matrices (internal) | evals-internal.md |
| Point-in-time inventory of legacy environment variables (historical) | config-knobs-audit.md |
llama.cpp and similar backends often expose a single prefix-cache slot. When
dispatch traffic or compaction invalidates it, the next turn records the
expected-cold reasons and shows one dim notice. Per-call cache verdicts
(hot, partial, cold, small) are persisted with timing and prompt-cache
counters in each session's context-snapshots.jsonl, so a slow session can be
diagnosed from the ledger alone.
clio-coder usage report --days 7 # cost and token facts with cited run idsInside the TUI, /cost shows session totals and /context opens the
context-window ledger. See docs/context-engine.md
for how the context engine measures and protects the prompt prefix.
Clio Coder is developed under the IOWarp project by the Gnosis Research Center at the Illinois Institute of Technology in collaboration with the University of Utah.
IOWarp and the CLIO (Context Layer for Input/Output) architecture are funded by the National Science Foundation under Award #2411318 for 2024 through 2029. Principal Investigator: Dr. Xian-He Sun. Co-Principal Investigators: Dr. Anthony Kougkas, Dr. Jake Hochhalter, and Dr. Vivek Srikumar.
Clio Coder is the interactive coding orchestrator in a larger ecosystem:
- clio-core is the foundational storage layer using Chimaera-based tiered data and context storage.
- clio-kit is a suite of 15+ Model Context Protocol servers exposing 150+ tools for scientific computing domains including HDF5, Slurm, ParaView, Pandas, ArXiv, NetCDF, FITS, and Zarr.
- Pi Agent Framework from Earendil Works: the @earendil-works/pi-ai execution engine, @earendil-works/pi-tui terminal rendering, and @earendil-works/pi-agent-core subagent orchestration.
- Anthropic Claude Agent SDK through @anthropic-ai/claude-agent-sdk for Claude Code worker runs under Pro/Max subscriptions.
- Agent Client Protocol for decoupling the engine from IDE frontends.
- Globus Auth for authenticating against ALCF's Sophia and Metis inference gateways.
Subagents and prompt techniques are evaluated against SWE-bench and SciCode. Every subagent run produces structured execution evidence, matched against baseline and candidate evaluations to catch silent regressions.
Licensed under Apache-2.0. See LICENSE and NOTICE.
Built for the people who maintain the code that science runs on.