A local-first agent harness for the Web and terminal, with durable context, agent delegation, and explicit control.
Install · Quick start · How it works · Roadmap
Note
Hames is ready for everyday use on Linux and now has early macOS support. It has been used extensively in real work, with good results so far. Expect continued refinement as more people use it. See the changelog for 0.3.0.
Hames brings coding agents and long-running work into one local workspace. Work in the browser or terminal, review plans before execution, and follow delegated work through its own conversations. A shared gateway keeps sessions, context, tools, and execution history durable across clients, with controls for approvals and account connections.
- Web and terminal — a browser workspace with live chat, agent editing, settings, event inspection, and memory, skills, and plugin management; a Ratatui TUI and classic REPL share the same gateway and durable sessions.
- Local and cloud models — connect llama.cpp, Ollama, OpenAI API, Codex, DeepSeek, Z.ai API or Coding Plan, Xiaomi MiMo API or Token Plan, Grok API, and Grok Build. Choose provider, model, and reasoning effort, with connection controls in Web and TUI.
- Plan, build, and review — approve or revise a plan before execution, delegate to agents with their own model defaults, and follow implementation and review chats.
- Explicit safety modes — Manual, Auto, and Plan behavior is enforced by the gateway, with exact one-shot approval for high-risk operations.
- Durable sessions and goals — resume or branch conversations, run autonomous goals across bounded turns, recover context as conversations grow, and keep foreground chat responsive while work continues.
- Auditable context and memory — inspect compiled context, token use, immutable events, request snapshots, layered memories, corrections, and Markdown/JSONL exports.
- Extensible tools — built-in filesystem and shell tools, private web search, external MCP servers, portable agents, delegation, isolated plugins, evolving Skills, and configurable personal or project slash commands.
Hames supports Linux and macOS. Mac support is newer and has been exercised by the maintainer; see macOS setup and limitations for its current scope.
Both platforms require:
- Python 3.12 or newer
- uv
- Rust 1.85 or newer
- Git
Install the latest stable version tag:
curl -fsSL https://raw.githubusercontent.com/saltnpepper97/hames/main/install.sh | bashThe installer uses locked dependencies, never invokes sudo, and
installs hames to uv's user tool bin directory (normally
). It keeps the source and Python environment under
/.local/bin/.local/share/hames/source so the Rust client can launch its matching
gateway. Running the command again installs the latest stable version tag.
The script is fetched from main, but the installed source comes exclusively from a tag.
Set HAMES_VERSION=0.3.0 to pin this release; branches and commit hashes are rejected.
To build from a reviewed local checkout instead:
git clone https://github.com/saltnpepper97/hames.git
cd hames
HAMES_INSTALL_LOCAL=1 ./install.shSet HAMES_VERSION, HAMES_BIN_DIR, or
HAMES_INSTALL_ROOT to customize a remote installation.
HAMES_REF remains a compatibility alias for a tag name. Local source builds
require HAMES_INSTALL_LOCAL=1; the default always installs tagged source.
Run the guided setup and verify the environment:
hames setup
hames doctorThen choose your interface:
hames web # Open the Web UI
hames # Open the terminal UI
hames repl # Open the classic REPLSetup can configure llama.cpp, Ollama, OpenAI, Grok API, Grok Build, or Codex. Hames defaults to a local
llama.cpp endpoint at http://127.0.0.1:8080; the gateway starts on
demand and stays available after the client exits. In Web, use Settings to
manage provider connections and Agents to configure your agents.
Useful terminal commands:
/model choose a reachable provider, model, and reasoning level
/mode auto switch between manual, auto, and plan
/sessions resume work in the current directory
/goal <objective> start durable autonomous work
/context inspect what entered the model request
/usage inspect estimated and provider-reported usage
/memory browse durable memories
/skills browse available procedures
/gateway inspect gateway health and active work
/help show the complete command list
Run hames repl for the classic line-oriented client. Piped or
redirected input selects it automatically. Run hames web for the
local browser interface; it starts or verifies the persistent gateway, requests
a one-time authenticated launch URL, opens it, and exits. Use
hames web --no-open to print the URL instead. The site remains
available for as long as the gateway is running.
Hames separates presentation from authority:
- The Web UI and terminal clients present conversations and controls. The SolidJS Web UI is served locally; the Rust launcher opens authenticated browser sessions and provides the TUI and classic REPL.
- The Python gateway owns sessions, providers, context compilation, tools, policy decisions, memory, and background work.
- The event ledger records durable, integrity-checked provenance and content-addressed payloads.
All clients talk to the same gateway and session model. Closing a client does not
discard an active goal or background terminal. Reopen a chat from the Web sidebar
or use /sessions in the terminal.
| Mode | Behavior |
|---|---|
| Manual | Confirms state-changing work with allow-once, allow-for-session, or deny. |
| Auto | Proceeds with ordinary trusted work and confirms dangerous operations. |
| Plan | Allows inspection and tests while preventing writes. |
Trusted workspace roots permit normal project work without repetitive prompts. Known secrets, credential stores, raw devices, Hames's private state, and deterministic high-risk shell signatures remain protected.
Sessions preserve provider, model, reasoning effort, interaction mode, ancestry,
and transcript state. Web provides chat and agent controls in the sidebar and
composer. In the terminal, use /new for a fresh conversation,
/sessions to resume, and /fork to branch after an answer.
/goal <objective> starts independently bounded agent turns under
a durable supervisor. Foreground messages take priority, and explicit
evidence-backed reports advance or complete the goal. Portable
AGENT.md capsules define an agent's role and authority separately
from session settings; delegation creates a bounded child session rather than
silently copying the parent conversation.
Select a coordinator with the chat agent picker and describe your task directly. Its configured agents can handle delegated work, with activity and results recorded in the normal conversation transcript.
Each run defaults to a 30-minute active-work budget. This includes model and foreground tool execution; waiting for delegated workers or human input does not use the coordinator's budget. Every worker has its own budget.
For long builds and team workflows, set the following in ~/.hames/config.toml
under the existing [runtime] section, then restart the gateway once it is idle:
[runtime]
max_active_seconds_per_run = 0Zero disables the active-time cutoff for both coordinators and workers. Positive values retain a cutoff in seconds. Model-turn and tool-call limits, individual tool/provider timeouts, and user cancellation still apply. Running tasks keep the budget they started with; this setting takes effect after gateway reload.
Stop affects only the selected agent. Its existing workers continue running and return their results to the parent chat. Stopping a worker reports the interruption to its parent without stopping siblings or automatically restarting the worker. Steer replaces only that agent's current turn; if a worker is steered, its parent waits for the replacement turn's result.
The agent_control tool lets a lead inspect and wait for existing assignments.
Stopping a named worker or the whole team requires an explicit user instruction;
the lead must quote that instruction from its current user turn. Stop/Steer on
the lead itself never authorizes a team stop. Gateway shutdown still ends all
running work.
Web exposes conversation activity in the Events tab, with dedicated pages for
memory, Skills, and Scars. Every model request passes through a deterministic,
budgeted context compiler. In the terminal,
/context explains selected, compacted, and omitted sources and links
them to the exact request snapshot. /inspect, /events,
and /export expose the corresponding activity and provenance.
Relationship, semantic, and episodic memory provide durable continuity.
Background extraction proposes bounded facts after settled turns; explicit
/remember captures, review controls, typed memory tools, immutable
corrections, and permanent forgetting keep that state manageable.
Corrections can open evidence-backed Scars that connect a failure to expected behavior and a repair. Hames routes the repair to the narrowest suitable layer, evaluates it, and guards later runs for healing or regression.
Hames discovers Skills from project .agents/skills, global
~/.agents/skills, and its built-in catalog. Repeated successful
workflows can become evaluated, versioned procedures; Skill scripts run offline
inside Bubblewrap on Linux or sandbox-exec on macOS, with a read-only project
and disposable writable scratch.
Optional private web search runs through a digest-pinned, loopback-only SearXNG container and a bundled MCP server. Hames can also connect to user-configured stdio or Streamable HTTP MCP servers:
hames mcp add filesystem --cwd "$PWD" -- npx -y @modelcontextprotocol/server-filesystem "$PWD"
hames mcp inspect filesystem
hames mcp enable filesystem
hames mcp listConfigured servers begin disabled. Environment and header mappings store only the source variable name, never its secret value. See MCP architecture, web search, and Skills.
For DeepSeek, Z.ai, and Xiaomi MiMo API/subscription setup, see cloud provider connections.
State is private by default under . A minimal
/.hames/.hames/config.toml looks like:
[runtime]
default_provider = "llama_cpp"
default_model = "qwen3.8-27b"
default_reasoning_effort = "medium"
default_interaction_mode = "auto"
[providers.llama_cpp]
adapter = "llama_cpp"
base_url = "http://127.0.0.1:8080"
model = "qwen3.8-27b"
reasoning_effort = "medium"
supported_reasoning_efforts = ["low", "medium", "xhigh"]
context_window_tokens = 131072Profile names are arbitrary, and multiple profiles may use the same adapter.
Nested environment variables override settings, for example
HAMES_RUNTIME__DEFAULT_MODEL=qwen3.8-27b. Set
HAMES_HOME to relocate all persistent state.
Fresh sessions use current runtime defaults. Resumed sessions restore their own durable provider, model, reasoning effort, agent, workspace, and mode.
The client starts the gateway when needed. To control it directly:
hames gateway status
hames gateway stop
hames gateway restartAn optional systemd user unit is available at
contrib/systemd/hames.service. Installing or
enabling it is intentionally left to the user.
On macOS, hames gateway service install enables an optional per-user
LaunchAgent for login startup; see macOS setup.
Install locked dependencies and run Hames from the checkout:
uv sync --locked
cargo build --locked
target/debug/hames doctor
target/debug/hamesRun the complete checks:
uv run ruff format --check .
uv run ruff check .
uv run pyright
uv run pytest
cargo fmt --all --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspaceThe web source lives under web/ and uses Node 22+ with pnpm. Its
production output is committed under src/hames/web_dist/ and
packaged with the gateway, so end-user installations do not require Node:
pnpm --dir web install --frozen-lockfile
pnpm --dir web check
pnpm --dir web test
pnpm --dir web build
cargo run -p hames-repl -- web --no-openThe backend diagnostic command is also available as
uv run hamesd doctor. The user-facing executable is built by
crates/hames-repl.
The implementation plan is the source of truth for completed and upcoming milestones:
See Release notes and custom command configuration for current capabilities.
Hames is available under the MIT License.