Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

execkit

Stateful, structured, safe shell sessions for AI agents, on real infrastructure.

When you give an AI agent a raw shell, you get a black box: one-shot commands with no memory between them, a wall of mixed stdout and stderr the agent has to guess its way through, secrets landing in plaintext, no record of what ran, and no way to undo a mistake. execkit replaces that with a session abstraction built for agents.

  • Stateful sessions. cd and environment persist across calls, like a real terminal, instead of every command starting fresh from home.
  • Structured results. Each command returns split stdout/stderr, an exit code, duration, and cwd as data the agent can act on, not a blob to parse.
  • Safe by default. Output is ANSI-stripped, secret-redacted, and bounded so a noisy build cannot blow the agent’s context window or leak credentials.
  • Real transports. Local shell, SSH, and Docker, with host-key verification, a sandboxed key directory, and host aliases from your ssh config.
  • Timeouts that keep the session. A command that runs too long is interrupted and reported with timed_out: true; the session keeps its cwd and env.
  • Undo. Git-backed workspace checkpoints let you snapshot before a risky change and restore on demand (files only, not side effects).
  • Observability. An append-only audit log, a live read-only viewer, and live MCP notifications so you can watch what the agent does in the shell.

Where it fits

execkit complements your agent’s built-in shell or sandbox rather than replacing it. Reach for it when the agent needs to work on a remote host or inside a container, when you want a record of what ran, or when you want to undo file changes on a remote workspace.

Two ways to use it

  • As an MCP server (execkit-mcp): a stdio Model Context Protocol server that any MCP-capable agent (Claude Code, Cursor, Gemini CLI, Codex, VS Code, Windsurf, and others) can drive directly. uvx execkit-mcp runs it without installing anything. Start at Installation.
  • As a Rust library (execkit): embed sessions in your own program. See the Rust library page. A Python SDK wraps the same core.

A note on safety

The agent driving these tools can be prompt-injected, so execkit treats every tool argument as untrusted. Anything dangerous to the host is controlled by the operator at startup (environment variables), never by a per-call agent argument. The command allow/deny fence is advisory defense-in-depth, not a sandbox: the real boundary is running the agent as a least-privilege user, in a container, or on a scoped SSH account. See the Security model.

execkit is Apache-2.0 licensed. Source: https://github.com/blinkingbit-oss/execkit.

Installation

execkit-mcp ships as a prebuilt binary. Pick whichever fits your toolchain.

Zero-install with uv. Nothing to install up front: the MCP client starts execkit through uvx, which fetches the PyPI package execkit-mcp on first use. See Wiring into an agent for the config block. To try it from a terminal:

uvx execkit-mcp --version

Installed:

pip install execkit-mcp      # a wheel; no Rust toolchain needed
cargo install execkit-mcp    # ...or via cargo

A prebuilt-binary installer script for Linux and macOS is attached to each GitHub release.

Building from source instead:

cargo build -p execkit-mcp --release   # binary at target/release/execkit-mcp

Verify the install

execkit-mcp --version
execkit-mcp doctor

doctor reports what is configured and what is missing before you ever wire an agent in: whether an audit destination is set and writable, where the SSH key directory and known_hosts resolve to, whether the Docker daemon is reachable, and whether an operator policy file is loaded. A typical run:

execkit-mcp 0.9.0
[ -- ] binary: /home/you/.cargo/bin/execkit-mcp

[ -- ] audit: off (set EXECKIT_MCP_AUDIT or EXECKIT_MCP_AUDIT_DIR to record + watch activity)
[ ok ] ssh key dir: /home/you/.ssh (override: EXECKIT_MCP_KEY_DIR)
[ -- ] known_hosts: /home/you/.execkit/known_hosts (absent; created on first SSH connect via TOFU)
[ ok ] docker: daemon reachable
[ -- ] policy: off (set EXECKIT_MCP_POLICY_FILE to enable)

Each [warn] or [ -- ] line tells you what to set. None of these are required to start, they just enable optional features (auditing, SSH, Docker).

Requirements on the target

Whatever the transport, the machine or container the agent works on needs a POSIX shell and base64 (part of coreutils and busybox). Checkpoints on SSH and Docker sessions also need git there.

Where the binary lives

cargo install puts it at ~/.cargo/bin/execkit-mcp. If that is not on your MCP client’s PATH, use the full path when you register it (the next page shows how, and execkit-mcp setup fills the absolute path in for you).

Next: Wiring into an agent.

Wiring into an agent

execkit-mcp is a stdio MCP server. You register the installed binary with your client once, and the agent gains session_create, session_exec, and the rest.

The fastest path is to let execkit print the exact config with the binary’s absolute path already filled in:

execkit-mcp setup claude     # or: cursor | gemini | codex | vscode | windsurf

It prints a ready-to-use block (and, for Claude Code, the one-line command). It deliberately does not edit your client’s config file for you, so it can never corrupt one; you paste the block into the right place.

Zero-install with uv

If you have uv, skip the install and let the client start execkit through uvx. This block works in any client that uses the mcpServers format:

{
  "mcpServers": {
    "execkit": { "command": "uvx", "args": ["execkit-mcp"] }
  }
}

Don’t run setup through uvx: the path it prints points into uv’s cache.

Claude Code

One command:

claude mcp add execkit -- execkit-mcp        # add `-s user` to enable it everywhere

Cursor, Gemini CLI and Windsurf

Cursor reads ~/.cursor/mcp.json (or .cursor/mcp.json in a project), Gemini CLI reads ~/.gemini/settings.json, and Windsurf reads ~/.codeium/windsurf/mcp_config.json. Add the same block to any of them:

{
  "mcpServers": {
    "execkit": { "command": "execkit-mcp" }
  }
}

If the binary is not on the client’s PATH, use the absolute path that execkit-mcp setup printed.

Codex CLI and VS Code

These use different formats. Codex reads TOML from ~/.codex/config.toml:

[mcp_servers.execkit]
command = "execkit-mcp"

VS Code reads .vscode/mcp.json in the workspace, with a servers key:

{
  "servers": {
    "execkit": { "type": "stdio", "command": "execkit-mcp" }
  }
}

execkit-mcp setup codex and execkit-mcp setup vscode print these with the absolute path filled in.

Turning on operator settings

Anything that affects the host (auditing, SSH key location, session limits) is configured by you, the operator, through environment variables in the client config, not by the agent. Add an env block:

{
  "mcpServers": {
    "execkit": {
      "command": "execkit-mcp",
      "env": { "EXECKIT_MCP_AUDIT": "/var/log/execkit.jsonl" }
    }
  }
}

See the Security model for the full list of settings and why they live with the operator. Once wired, the agent calls session_create -> session_exec -> session_destroy; see Sessions.

FAQ

Does execkit replace the agent’s built-in shell?

No. execkit is an MCP server: it adds tools (session_create, session_exec, session_destroy, …) to whatever the agent already has. It does not hook, wrap, or intercept the client’s native shell tool (Claude Code’s Bash, for example). After you wire it in, the agent’s tool list is “native tools plus execkit’s tools.” Both are live at once.

Should I use execkit instead of my agent’s built-in sandbox?

Use both. A built-in shell or sandbox is the right tool for quick local commands in the project the agent is working on. execkit adds what those usually don’t have: sessions on remote hosts over SSH and inside Docker containers, a JSONL audit trail with a live viewer, redacted and budgeted output, and checkpoints you can restore on remote workspaces. It is not a sandbox itself; see Security model.

How does the agent decide to use execkit instead of running commands locally?

It is the model choosing from its tool list; there is no automatic rerouting. The choice is driven by:

  • Tool descriptions and the server’s instructions. execkit ships an instructions string (“call session_create to get a session, session_exec to run commands…”) that tells the model what the tools are for.
  • Your project instructions (for example CLAUDE.md).
  • The task. For anything the native shell cannot do (SSH to a remote host, exec inside a Docker container, a session with persistent cwd/env, workspace checkpoints), execkit is the only tool that can, so the model reaches for it. For a quick local command, the native shell is the path of least resistance unless you steer the model.

So out of the box, plugging in execkit makes it available, not mandatory.

How do I make the agent always use execkit?

Three levels, weakest to strongest:

1. Instruct it. Add a line to CLAUDE.md (or the client’s system/project prompt): “Run all shell commands through the execkit session tools, not the built-in shell.” This steers the model; it does not guarantee.

2. Remove the alternative. Disable the client’s built-in shell tool so execkit is the only shell path the model has. In Claude Code, deny the Bash tool in settings.json:

{
  "permissions": {
    "deny": ["Bash"],
    "allow": ["mcp__execkit__*"]
  }
}

A bare tool name like "Bash" removes the tool from the model’s context entirely: the model never sees it and never attempts it. (This is different from a scoped rule like "Bash(rm *)", which leaves Bash available and only blocks matching calls at execution time.) The allow line explicitly permits every execkit tool so the model reaches for those instead. On Windows, also deny "PowerShell", the other built-in shell surface:

{ "permissions": { "deny": ["Bash", "PowerShell"], "allow": ["mcp__execkit__*"] } }

For a one-session override instead of durable settings, use the CLI flag: claude --disallowedTools Bash.

3. Isolate at deployment. Run the agent where it has no local shell to the machine that matters, and only execkit’s SSH/Docker transport reaches it. Now execkit is not a preference. It is the only door.

What do execkit’s safety features actually cover?

Only commands that go through execkit. The audit log, secret redaction, the allow/deny fence, and checkpoints apply to session_exec calls. If the agent runs something via its own native shell, execkit never sees it. This is why “make execkit the only path” (options 2 and 3 above) matters: enforcement comes from removing the competing path, not from execkit trapping calls.

And even for commands it does see, the command fence is advisory, not a sandbox: string matching is bypassable. The real boundary is the operating system: a least-privilege user, a container, or a scoped SSH account. See the Security model.

Is any of this specific to Claude Code?

No. “MCP servers add capabilities; they do not hijack the host’s existing tools” is true of every MCP client. Only the step that disables the native shell is client-specific; how the agent chooses a tool is the same everywhere.

My command timed out. Is the session gone?

No. A command that runs past its timeout (120 seconds by default over MCP, or timeout_secs on the call) is interrupted with Ctrl-C and returned with timed_out: true and exit code 124. Its output so far comes back (stderr keeps the last 16 KiB, followed by a timeout note). The session keeps its cwd and env. For jobs that take longer, start them in the background and poll the log:

nohup ./long-job.sh > /tmp/job.log 2>&1 &
tail -n 20 /tmp/job.log

See Sessions.

Why can’t I answer a prompt, use sudo, or open an editor?

Commands run with stdin set to /dev/null, so nothing can wait for keyboard input. A prompt gets end-of-file and usually fails at once instead of hanging the session. Use non-interactive forms: sudo -n (with a NOPASSWD rule for what the agent may run), apt-get -y, GIT_EDITOR=true. Pagers are already off: each session exports PAGER=cat (and GIT_PAGER, MANPAGER, SYSTEMD_PAGER), so git log and man print straight through. Running less or an editor directly still hangs until the timeout and closes the session; see Sessions.

The agent says “unknown session_id”. What happened?

The session was closed: the agent destroyed it, the shell exited (exit, or a failing command under set -e), a timed-out command ignored Ctrl-C, or it was idle past EXECKIT_MCP_SESSION_TTL. The agent can call session_list to see which sessions are open, and session_create to start a new one.

Upgrading to 0.9

0.9 changes some defaults and a few library types. Check this list before you upgrade a server or bump the crate.

Behaviour

  • SSH host keys live in ~/.execkit/known_hosts. Before 0.9 execkit pinned keys in ~/.ssh/known_hosts. Those pins are not read any more. By default, the first connection to each host after upgrading records its key again (trust on first use), so connect once from a network you trust. To keep your old pins, copy execkit’s lines (they look like host SHA256:..., unlike OpenSSH’s own entries) into the new file:

    mkdir -p ~/.execkit && chmod 700 ~/.execkit
    grep -E '^[^ ]+ SHA256:' ~/.ssh/known_hosts >> ~/.execkit/known_hosts
    

    Old pins were keyed by bare host whatever the port. A line for a host you reached on a port other than 22 must be rewritten as [host]:port after copying (for example example.com SHA256:... becomes [example.com]:2222 SHA256:...); left as is, it is read as the port-22 pin, so a later port-22 connection to that host fails as a key mismatch.

    If you set EXECKIT_MCP_KNOWN_HOSTS, that file is still used. See Transports.

  • stdin is closed. Every command runs with stdin set to /dev/null. Prompts, REPLs and read get end-of-file instead of waiting. Pagers are set to cat for the session (see Sessions).

  • The target needs base64. Commands are sent base64-encoded. GNU coreutils, busybox and macOS base64 all work. A shell without it fails at session_create with a clear error.

  • Timeouts return a result. A command that runs past its timeout is interrupted with Ctrl-C and returned normally, with exit_code 124 and timed_out: true. The session stays usable. In 0.8 a timeout was an error (command still running) and closed the session. Only a command that ignores Ctrl-C still closes it.

  • Session ids changed format. They are now <run>-<n>_<transport>..., for example a3f9-1_local or a3f9-2_ssh_deploy@web1, instead of 1_local. Tools that parse ids or match audit file names need updating. See Sessions.

Rust library

  • ExecResult has a new public field, timed_out: bool. Code that builds an ExecResult with a struct literal must set it. Code that only reads results is unaffected.
  • SshConfig has a new public field, connect_timeout: Duration (default 15 seconds). Build it with SshConfig::new(...) and then set fields, rather than a struct literal.
  • Error::StillRunning now means the timeout could not be interrupted, and the session is closed. An ordinary timeout is Ok with timed_out set.

The Python ExecResult gains a timed_out attribute. Nothing was removed.

Sessions

A session is a live shell that outlives a single tool call. The agent opens one, runs as many commands as it likes against it (state carries over), and closes it.

The tools

ToolArgumentsReturns
session_createtransport ("local" / "ssh" / "docker") plus transport options (see Transports); optional allow / deny lists; optional output_budget{ "session_id": "..." }
session_execsession_id, command, optional budget, optional timeout_secsstructured ExecResult
session_listnone[{ session_id, transport, idle_secs }]
session_destroysession_id{ "destroyed": true }

Remote sessions add session_checkpoint, session_checkpoints, and session_restore; see Checkpoints.

State persists

Sessions are stateful. cd, exported variables, and shell state carry across session_exec calls, the way a real terminal works:

// session_exec {"session_id":"a3f9-1_local","command":"cd /srv/app && export ENV=prod"}
// session_exec {"session_id":"a3f9-1_local","command":"pwd"}   -> stdout: "/srv/app"
// session_exec {"session_id":"a3f9-1_local","command":"echo $ENV"} -> stdout: "prod"

This is the difference between a shell and a series of unrelated strangers: an agent that runs cd packages/api and then npm test gets the test run in packages/api, not back in the home directory.

Structured results

session_exec returns an ExecResult as JSON, not a blob:

// session_exec {"session_id":"a3f9-1_local","command":"npm run build"}
{
  "command": "npm run build",
  "stdout": "...",
  "stderr": "Error: Cannot find module 'webpack'",
  "exit_code": 1,
  "duration_ms": 3420,
  "cwd": "/home/u/app",
  "truncated": false,
  "timed_out": false
}

stdout and stderr are split, so the agent never has to guess whether output was an error. The exit code is authoritative. Output is ANSI-stripped and secret-redacted before it is returned, and bounded so one command cannot flood the agent’s context (see Output budgets). When output was cut and no budget was passed, the result also carries a hint saying a budget would shape it.

What a command can contain

Each command is sent to the shell base64-encoded and run with eval, so the text reaches the shell exactly as written. Comments, a trailing &, heredocs, !, tabs, very long commands and even syntax errors behave the way they would in a terminal; none of them hang the session. The target needs base64 on its PATH.

Commands are non-interactive. stdin is /dev/null, so a program that reads stdin (a password prompt read from stdin, a REPL, read) gets end-of-file instead of hanging. Use non-interactive flags: sudo -n, apt-get -y. Shell history is off, so commands are never written to a history file.

stdout is still a terminal, though, so tools that page their output would start less, and less reads keys from the terminal itself, not stdin, and ignores Ctrl-C. To stop git log, git diff, man, systemctl status and journalctl from hanging, every session exports these defaults when it starts:

VariableValueEffect
PAGER, GIT_PAGER, MANPAGER, SYSTEMD_PAGERcatoutput is printed straight through
LESSFRXa less that still gets started quits when the output fits on one screen

A command can override them (export PAGER=...). Don’t run a pager or any other full-screen program (less FILE, vim, top) yourself. It waits for keys until the timeout, the Ctrl-C does not stop it, and the session is closed. Use cat, head or tail, or an output budget, instead.

Timeouts

Every session_exec has a timeout: timeout_secs on the call, or the operator default EXECKIT_MCP_EXEC_TIMEOUT (120 seconds if unset). Both are clamped to 1-3600 seconds.

When a command runs past it, execkit sends Ctrl-C, waits for the shell to come back, and returns a normal result:

// session_exec {"session_id":"a3f9-1_local","command":"sleep 30","timeout_secs":2}
{
  "command": "sleep 30",
  "stdout": "",
  "stderr": "execkit: timed out after 2s; sent Ctrl-C. The shell session is intact (cwd/env kept). ...",
  "exit_code": 124,
  "duration_ms": 2052,
  "cwd": "/tmp",
  "truncated": false,
  "timed_out": true
}

The session keeps its cwd and env and accepts the next command. Only a command that ignores Ctrl-C ends the session (see below).

Long-running jobs

For builds, test suites or deploys that may outlast the timeout, start the job in the background and poll its log:

// session_exec {"session_id":"a3f9-1_local","command":"nohup make release > /tmp/release.log 2>&1 &"}
// session_exec {"session_id":"a3f9-1_local","command":"tail -n 20 /tmp/release.log"}

The first call returns at once. Later calls read the log. The job keeps running between calls because the session’s shell stays open.

When a session closes on its own

The server closes a session and frees its slot when:

  • the shell exits, because a command ran exit or a failing command hit set -e;
  • a timed-out command ignores Ctrl-C, so the shell can no longer be trusted to frame output correctly;
  • it sits idle past the TTL (below).

The session_exec that caused it returns a tool error that says the session was closed. After that, any call with the old id returns unknown session_id '...'; call session_list to see live sessions. The agent can call session_list at any time to see which sessions are open and how long each has been idle.

Session ids are self-identifying

Ids read as <run>-<n>_local, <run>-<n>_ssh_<user>@<host>[:port], or <run>-<n>_docker_<container>, for example a3f9-1_local. <run> is a short random prefix for each server process, so ids from different runs are unlikely to collide in a shared audit directory. The rest keeps logs and the watch viewer legible at a glance. Agent-provided host/user/container names are sanitized before they appear in an id or a filename.

Lifecycle and limits

Sessions are reaped when idle (default 30 minutes) to free the process and a slot against the concurrent-session cap (default 64). Both are operator-tunable; see the Security model. Always session_destroy when done.

Transports

session_create takes a transport: "local", "ssh", or "docker". Any other value is an error.

Every transport needs a POSIX shell and base64 on the target: execkit sends each command base64-encoded and decodes it there.

Local

A shell on the machine running the server.

// session_create {"transport":"local"}  -> {"session_id":"a3f9-1_local"}

The agent reaches whatever the server’s user can. Run that user with least privilege.

SSH

// session_create {
//   "transport":"ssh", "host":"web-01", "user":"deploy",
//   "key_path":"deploy_ed25519"            // or "password":"..."
// }

Required: host, plus a user and a way to authenticate, unless your ssh config supplies them (below). Optional: port (default 22) and fingerprint to pin an exact host key.

Host aliases from your ssh config

host can be a Host alias from the operator’s ssh config, read from <key dir>/config (so ~/.ssh/config by default). For an entry like this:

Host web
    HostName web-01.internal
    User deploy
    Port 2222
    IdentityFile ~/.ssh/deploy_ed25519

the agent only needs {"transport":"ssh","host":"web"}. HostName, User, Port and IdentityFile fill in whatever the call leaves out; arguments the agent passes take precedence. Only exact Host names match: wildcard patterns, Match blocks and Include are ignored.

Authentication

execkit uses the first of these that applies:

  1. password, if given.
  2. key_path, if given.
  3. The alias’s IdentityFile entries, in order, then id_ed25519, id_ecdsa and id_rsa in the key directory. The first file that exists inside the key directory is used. If the server rejects it, execkit does not try the next one; pass key_path to choose a different key.

Every key must live inside the key directory (~/.ssh by default, or EXECKIT_MCP_KEY_DIR). Out-of-bounds or traversal paths are rejected with a generic error that does not leak whether the path exists. An IdentityFile that points outside the key directory is skipped.

Host keys

Host-key handling is safe by default:

  • Verified against execkit’s own known_hosts (TOFU). The file is ~/.execkit/known_hosts unless EXECKIT_MCP_KNOWN_HOSTS overrides it. The first connection to a host records its key; a changed key is rejected as a likely man-in-the-middle. Entries are keyed host on port 22 and [host]:port otherwise, so two ports on one host are verified separately.

  • Not your OpenSSH ~/.ssh/known_hosts. Before v0.9 execkit wrote its pins to ~/.ssh/known_hosts. It now keeps its own file and does not read the old pins, so after upgrading, the first connection to each host records its key again. To carry your pins over instead, copy execkit’s lines (host SHA256:...) across:

    mkdir -p ~/.execkit && chmod 700 ~/.execkit
    grep -E '^[^ ]+ SHA256:' ~/.ssh/known_hosts >> ~/.execkit/known_hosts
    

    Old pins were keyed by bare host whatever the port. A line for a host you reached on a port other than 22 must be rewritten as [host]:port after copying (for example example.com SHA256:... becomes [example.com]:2222 SHA256:...); left as is, it is read as the port-22 pin, so a later port-22 connection to that host fails as a key mismatch.

    If EXECKIT_MCP_KNOWN_HOSTS points at an OpenSSH-format file, connecting to a new host fails with an error saying so, instead of writing to that file. See Upgrading to 0.9 for the other changes.

  • Pin a key by passing fingerprint ("SHA256:...") for an exact match.

Connect timeout

Connecting, the SSH handshake and authentication share one 15 second budget. An unreachable host or a server that stalls during auth fails session_create with a timeout error instead of hanging.

The home directory behind ~/.ssh and ~/.execkit resolves by priority ($HOME, then the system passwd database), so defaults are correct even when $HOME is unset, as in a service-launched server.

For throwaway or test hosts only, EXECKIT_MCP_INSECURE_ACCEPT_ANY_HOSTKEY=1 disables host-key verification. Never use it in production.

Docker

// session_create {"transport":"docker","container":"app-web-1"}

Runs docker exec against any container the daemon can see, so the agent reaches whatever your Docker context exposes. If the container does not exist or is not running, session_create fails with an error that names it. The container needs a POSIX /bin/sh. Grant the server Docker access only when you want that, and scope the daemon or context accordingly.

Remote workspace undo

SSH and Docker sessions support Checkpoints: a git-backed snapshot of the workspace you can restore on demand.

Output budgets

A noisy command can dump thousands of lines. That hurts twice: it costs the agent context window, and the volume itself degrades the agent’s reasoning. Output budgets shape a command’s output before it reaches the model.

Pass budget to session_exec, or output_budget to session_create for a session default:

// keep only the last 200 lines of a noisy build
{ "session_id": "a3f9-1_local", "command": "npm run build",
  "budget": { "keep": { "mode": "tail", "n": 200 } } }

// grep a 50k-line log for errors, with 2 lines of context around each
{ "session_id": "a3f9-1_local", "command": "cat big.log",
  "budget": { "grep": { "pattern": "error|fail", "context": 2 } } }

Keep modes are tail, head, and head_tail; grep is a separate filter; both honor a max_chars cap.

Shaping is line-based, applied client-side, and runs after secret redaction. It never changes the exit code or any side effect of the command, only what text comes back. When a budget is applied, the result carries a budget report so the agent knows the output was shaped:

"budget": {
  "stdout": { "mode": "tail", "lines_total": 4123, "lines_kept": 200 },
  "stderr": { "mode": "tail", "lines_total": 12, "lines_kept": 12 }
}

How much output a budget sees

With a budget, execkit holds at least 8 MiB of a command’s output in memory before shaping it, so grep finds a match anywhere in a large log and lines_total is the real count. Larger output is compacted as it arrives: the buffer may grow to 16 MiB, then execkit keeps the first 4 MiB and the most recent 4 MiB, with a [execkit: N bytes elided] line (N is the running total) where the middle was. stderr arrives after stdout and is compacted on its own the same way, with its own elision line and count. It keeps a head and a tail of at least 2 MiB each (up to 4 MiB each when stdout was small). Once that happens, the budget only sees what was kept:

  • lines_total counts the kept head and tail lines (plus the elision line), not every line the command printed. The result has truncated: true, which tells you the count is partial.
  • grep cannot match lines in the elided middle.

For very large output, filter in the command itself (grep, tail -n, wc -l), or write it to a file and read that in pieces.

Without a budget, output is capped at about 100,000 characters per stream and the result has truncated: true. Over MCP it also carries a hint suggesting a budget:

// session_exec {"session_id":"a3f9-1_local","command":"seq 1 200000"}
{ "stdout": "1\n2\n3\n...", "truncated": true, "timed_out": false,
  "hint": "output was truncated; pass budget (grep/keep/max_chars) to shape it" }

Use budgets liberally on commands you expect to be loud (builds, installs, big log reads); the agent keeps the signal without the noise.

Checkpoints

On SSH and Docker sessions, execkit can snapshot the workspace before a changing command and restore it on demand: a filesystem “undo” for agent actions.

It undoes files only, never side effects. A dropped database stays dropped, a sent email stays sent, an installed package stays installed. For “the agent mangled my source tree,” that is exactly the recovery you want.

Tools

ToolArgumentsReturns
session_checkpointsession_id, optional label{ "checkpoint_id": "..." }
session_checkpointssession_id[{ id, label, created }]
session_restoresession_id, optional checkpoint_id{ restored_to, files_changed }

Omit checkpoint_id on restore to roll back to the most recent checkpoint.

Enabling it

Two requirements:

  1. git on the remote host. Checkpoints use a shadow git repo. If git is absent, auto-snapshot disables itself and checkpoint calls return a clear “install git on the remote” error. Restore needs git 2.22 or newer (it uses git checkout --no-overlay); with an older git, session_restore fails with git’s error and changes no files.
  2. An explicit workspace on session_create. Without it, checkpoints and auto-snapshot are off. execkit will not default to the cwd or $HOME, snapshotting a home directory is slow and would capture secrets. Set workspace to the project directory you want undo for (pass $HOME explicitly only if you truly mean it).

Control it via session_create:

  • workspace (root; REQUIRED to enable checkpoints)
  • auto_snapshot (default true; effective only with a workspace)
  • paths (sub-directories under the root to track)
  • checkpoint_ignores (extra gitignore-style patterns, added to the built-in defaults: .git, node_modules, build dirs, caches, .ssh, .aws, .env, .env.*, *.pem, *.key, id_rsa*, id_ed25519*, …)

Credential-shaped files (.env, .env.*, *.pem, *.key, id_rsa*, id_ed25519*, plus .ssh, .gnupg, .aws, .netrc) are never snapshotted, regardless of checkpoint_ignores - these rules always apply last and cannot be overridden by a negation pattern.

Restore is destructive

Warning. session_restore reverts tracked files to their state at the checkpoint, removes files the checkpoint did not have, and deletes all untracked files and directories anywhere under the workspace (via git clean). Ignored paths (node_modules, build dirs, the credential files above) are left alone. This includes files created and tracked by a later checkpoint than the one you’re restoring to - restore always leaves the workspace matching the target checkpoint exactly, not a merge of it with whatever came after. Do not restore if untracked files in the workspace must be preserved.

Shadow repo cleanup

Each session’s checkpoints live in a private shadow git repo under ~/.execkit/ckpt-<token>.git on the remote host. execkit removes it automatically when the session ends (on Session drop / session_destroy), so no state persists between sessions and no per-session directories accumulate on the remote host over time.

Security model

The agent driving these tools can be prompt-injected, so execkit treats every tool argument as untrusted. Anything dangerous to the host or filesystem is controlled by the operator at startup through environment variables, never by a per-call agent argument. An injected agent cannot change where the audit log is written, which directory SSH keys come from, or the session limits.

Operator settings

Env varPurposeDefault
EXECKIT_MCP_AUDITappend a JSONL audit log of every command hereoff
EXECKIT_MCP_AUDIT_DIRone JSONL file per session in this directory (<session_id>-<open_ms>.jsonl); takes precedence over EXECKIT_MCP_AUDIToff
EXECKIT_MCP_AUDIT_RETENTION_DAYSdelete per-session log files older than N days at startup (dir mode only); 0 disables14
EXECKIT_MCP_EXEC_TIMEOUTdefault session_exec timeout in seconds (clamped 1-3600); timeout_secs overrides per call120
EXECKIT_MCP_KEY_DIRSSH keys must canonicalize to inside this dir; its config file supplies Host aliases~/.ssh
EXECKIT_MCP_KNOWN_HOSTSexeckit-managed SSH host-key verification file (TOFU; rejects changed keys)~/.execkit/known_hosts
EXECKIT_MCP_INSECURE_ACCEPT_ANY_HOSTKEYDANGEROUS disable host-key checksunset
EXECKIT_MCP_MAX_SESSIONSsoft cap on concurrent live sessions64
EXECKIT_MCP_SESSION_TTLreap sessions idle longer than N seconds; 0 disables1800
EXECKIT_MCP_POLICY_FILEJSON allow/deny (program names) + deny_patterns (regex) the agent cannot edit; advisoryoff

EXECKIT_MCP_KEY_DIR and EXECKIT_MCP_KNOWN_HOSTS default off the home directory, which resolves by priority ($HOME, then the passwd database), so the defaults are correct even when $HOME is unset. Run execkit-mcp doctor to see what each one resolves to on your machine.

What is enforced where

  • Host keys are verified by default (TOFU against known_hosts; a changed key is rejected as a likely MITM). Pin an exact key with fingerprint, or set the insecure env var only for throwaway hosts.
  • key_path is sandboxed to EXECKIT_MCP_KEY_DIR; traversal or out-of-bounds paths are rejected with a generic error that does not leak path existence.
  • The audit destination is operator-chosen, never a tool argument, so an injected agent cannot write to arbitrary files.
  • Docker sessions reach any container the daemon can see. Grant Docker access deliberately and scope the context.
  • The server speaks MCP on stdout; all diagnostics go to stderr.

Secret redaction

Output (stdout and stderr), the echoed command field, the audit log and the live notifications are all redacted before they leave execkit. Matches become [REDACTED]. Redaction runs before output budgets, so a secret cannot survive by being cut in half.

CoveredExamples
AWS access key idsAKIA...
GitHub tokensghp_, gho_, ghu_, ghs_, ghr_, github_pat_
GitLab personal access tokensglpat-...
Slack tokensxoxb-, xoxa-, xoxp-, xoxr-, xoxs-
Stripe live secret keyssk_live_...
Google API keysAIza...
Anthropic API keyssk-ant-...
OpenAI-style keyssk-..., sk-proj-... (32+ characters)
JSON Web TokenseyJ....eyJ....sig
PEM private keysthe whole -----BEGIN ... PRIVATE KEY----- block, not just the header
Passwords in URLspostgres://user:[REDACTED]@db (user and host are kept)
Bearer and Basic credentialsAuthorization: Bearer [REDACTED] (tokens of 16+ characters, up to whitespace, a quote, , or ;) and Authorization: Basic [REDACTED]; the scheme word is kept
Secret-named key=value / key: value pairspassword, passwd, secret, secret_key, token, api_key, access_key, private_key, including prefixed names like DB_PASSWORD, SECRET_KEY or AWS_SECRET_ACCESS_KEY (values of 4+ characters). A quoted value is redacted up to its matching closing quote, whatever it contains (including the other kind of quote); with no closing quote on the line it is redacted like an unquoted value. An unquoted value runs to whitespace, a quote or one of ;&|); a comma does not end it, so a trailing , is redacted with it
Values the session assigned to secret-named variablesafter export DB_PASS=hunter2hunter2, the literal hunter2hunter2 is redacted wherever it appears later in that session (names containing token, secret, passw, api_key, private_key, credential or auth; values of 6+ characters)
Not coveredWhy
Arbitrary high-entropy stringsno fixed shape; matching them would redact hashes, ids and base64 data too
Encoded, reversed or split secretsbase64, rev or cut output no longer has the shape
Secrets with no recognisable shape and no secret-named variablefor example a password printed from a file the session never assigned
Values assigned outside the sessiona variable set in a login profile or by another process is not learned
Secret-named values that look like codean unquoted value is left alone only in three code shapes, so source code stays readable: a dotted or :: path ending at a bracket (access_key = cfg.get("x")), < followed by an uppercase letter (token: Option<String>), or ( followed by a quote (token = get("abc")). Other bracketed values, such as Pass[123], Summer{2024} or abc(def), are redacted, and so are code lines like secret = vec[0]

Redacted output is not file-accurate. Redaction can also hit ordinary code: a secret-named field with a bare identifier as its value, such as TypeScript password: string, comes back as password: [REDACTED]. Never write command output back to a file (for example cat-ing a file and saving what came back); edit files in place with sed, patch or similar instead.

Redaction is a safety net for accidental leaks, not a guarantee. An agent that wants to exfiltrate a secret can encode it first. Keep secrets the agent should not see out of the environment it can reach.

The fence is advisory, not a sandbox

allow / deny command lists are defense in depth, not a jail. Matching on command strings is trivially bypassable (env rm, $(echo rm), base64, bash -c "..."). Name matching looks at the first word of each pipeline segment, so deny: ["curl"] blocks curl and /usr/bin/curl but not env curl or sudo curl. Treat the fence as a guardrail against accidents and obvious mistakes. A denial names the rule or pattern that matched, so the agent can see why a command did not run.

The real security boundary is the operating system: run the agent’s shell as a least-privilege user, in a container, or on a scoped SSH account, so that even a fully compromised agent can only reach what that account can. execkit gives you visibility and undo on top of that boundary; it does not replace it.

Operator command policy

Point EXECKIT_MCP_POLICY_FILE at a JSON file to set an allow/deny fence the agent cannot edit (unlike the per-call allow/deny, which the agent supplies):

{
  "allow": ["git", "ls", "npm"],
  "deny": ["rm", "dd", "shutdown"],
  "deny_patterns": ["\\brm\\b", "kubectl\\s+delete", "git\\s+push\\s+.*--force"]
}
  • allow (program names): if non-empty, only these may run. Empty/absent = all.
  • deny (program names): always blocked; deny wins over allow.
  • deny_patterns (regex over the whole command): for what names cannot express.

Prefer a deny_pattern over a name deny for anything that matters: name matching only sees the program name per pipeline segment, so deny: ["rm"] misses sudo rm and xargs rm, while deny_patterns: ["\\brm\\b"] catches them. In JSON the regex backslashes double up (\\b); use (?i) for case-insensitive matching.

A blocked command never runs; it is recorded in the audit log, shown in watch, and pushed to the client as a warning. This is an ADVISORY guardrail, not a sandbox: string matching is trivially bypassable (env rm, base64, bash -c). The real boundary is a least-privilege user, a container, or a scoped SSH account.

Auditing and the watch viewer

execkit can record everything an agent does in the shell, and let you watch it live.

The audit log

Point the server at a destination and every open / exec / close event is appended as JSON, with the session id, transport, an epoch-millisecond timestamp, and the command plus its (redacted, bounded) output.

  • EXECKIT_MCP_AUDIT=/var/log/execkit.jsonl writes one shared file for all sessions.
  • EXECKIT_MCP_AUDIT_DIR=/var/log/execkit/ writes one file per session, named <session_id>-<open_ms>.jsonl. This mode takes precedence when both are set.
  • EXECKIT_MCP_AUDIT_RETENTION_DAYS (default 14, 0 disables) prunes per-session files older than N days at startup. Files with a future-skewed mtime are never deleted.

The audit destination is operator-chosen and never a tool argument, so an injected agent cannot redirect or suppress it.

The watch viewer

Point watch at the audit file or directory from another terminal:

execkit-mcp watch /var/log/execkit.jsonl   # or: execkit-mcp watch  (uses $EXECKIT_MCP_AUDIT)
execkit-mcp watch /var/log/execkit/        # a directory; uses $EXECKIT_MCP_AUDIT_DIR

It is a live, read-only TUI: the agent’s sessions on the left, the selected session’s shell transcript on the right (prompt, command, stdout, stderr in red, exit status), rendered like a normal shell rather than JSON. Command, opened, closed and blocked lines start with their local time ([HH:MM:SS], following $TZ), with a -- YYYY-MM-DD -- line when the date changes, as in the browser viewer. Switch sessions with 1-9 or the arrow keys, scroll with PgUp/PgDn, quit with q. It only ever reads the log. Because the data comes from the server, it works the same under any MCP client.

Browser viewer

For a richer view in a normal browser tab, serve the transcript as a local web page:

execkit-mcp watch --serve /var/log/execkit/        # prints a loopback URL with a token
execkit-mcp watch --serve --open /var/log/execkit/ # ...and open it in your browser

The MCP server also starts the viewer automatically when EXECKIT_MCP_WATCH_WEB is set: it binds 127.0.0.1 only, prints the tokened URL and pushes it to the client as a notification, and keeps the URL stable across restarts so an open tab reconnects. EXECKIT_MCP_WATCH_PORT sets the port (default 7878, falls back to a random one if taken) and EXECKIT_MCP_WATCH_OPEN also opens the browser for you.

The page is read-only and local by construction: it binds loopback only, every request needs the URL token, and it can read the audit stream but never touch a session, a command, or your files. What you get:

  • Sidebar grouped by transport (local / ssh / docker), then by host or target inside each transport (the ssh host or alias such as etlstage, the docker container, or local). Each group header shows its session count. Every session gets its own row: start time (with the date when it is not today), #<n>, the user@host label or your alias, the command count, and a green (live) or grey (closed) dot. The History list uses the same grouping.
  • Timestamps: each command line starts with the local time it started, for example [19:19:22] /tmp $ echo hi; hover it for the full date, time and timezone offset. The audit log records an exec event when the command finishes, so the viewer shows ts - duration_ms. The opened, closed and blocked lines show their event time. A -- YYYY-MM-DD -- line marks a date change, and the first one also appears at the top when a session did not start today. Exports and screenshots include the same times (JSON exports keep the raw ts in unix milliseconds plus the local time).
  • Colored transcript with a header legend (cmd / out / err / ok). Click a legend item to show or hide that line type.
  • Search: press / to find within the transcript, step matches with Enter / Shift+Enter, and press e to jump to the next error or blocked line.
  • History of past sessions, with relative times, when EXECKIT_MCP_AUDIT_DIR is set. Sessions are grouped by transport, then by host, and listed newest first within each host. It shows the 20 newest sessions plus any you kept. It sits at the bottom of the sidebar, collapsed to a History (N) header, or History (20 of 57) when not every session is listed; click the header to expand it (the viewer remembers your choice), then click a session to read its transcript. With a single EXECKIT_MCP_AUDIT file there is no per-session history.
  • Per-session actions from a 3-dots menu: rename (a display alias), pin, keep, export to .txt / .log / .md / .json, and screenshot to .png.

  • A status bar that shows the selected session’s details; click it to copy the session id.

Rename / pin / keep and the sidebar width persist in ~/.execkit/viewer-state.json (mode 0600). That file is the viewer’s only write surface: display metadata only, never able to affect a session, a command, or the audit log.

Headless follow mode

For a pipeable, no-TTY view, use --follow instead of the TUI. It prints each command and its output as a line prefixed with the session id, as it happens. Command, opened, closed and blocked lines carry their local time, and a -- YYYY-MM-DD -- line appears first when the events are not from today and again whenever the date changes:

execkit-mcp watch --follow /var/log/execkit/
# [a3f9-1_local] [14:02:11] /home/u $ npm run build
# [a3f9-1_local] x exit 1  (3420ms)
# [a3f9-2_ssh_deploy@web-01] [14:02:40] /srv $ systemctl restart app

Live notifications in the client

Even with no audit log configured, the server streams each command to the MCP client as it runs, so a host agent can surface its own shell activity without anyone opening a separate terminal. Every session_exec emits:

  • a log notification (notifications/message) carrying the full shell transcript, info on success and warning on a non-zero exit; and
  • a progress notification (notifications/progress) with a one-line summary, when the call supplied a progressToken.

This reveals nothing new: the client already receives the same output in the tool result, redacted and bounded. How the activity is surfaced is up to the client.

CLI reference

execkit-mcp with no arguments is the stdio MCP server an agent launches. Everything below is for a human at a terminal.

Commands

execkit-mcp                          Run the MCP server on stdio (default)
execkit-mcp setup <client>           Print the config to wire execkit into a client
                                     client: claude | cursor | gemini | codex | vscode | windsurf
execkit-mcp doctor                   Check the local environment and print a report
execkit-mcp watch [--follow|--serve] [--open] <path>
                                     Live, read-only viewer (TUI, follow stream, or browser)
execkit-mcp --version | version      Print version
execkit-mcp --help                   Print help

setup <client>

Prints a ready MCP config block with this binary’s absolute path filled in, and for Claude Code the claude mcp add one-liner. Each client gets its own format and file location: JSON mcpServers for Claude Code, Cursor, Gemini CLI and Windsurf, JSON servers for VS Code (.vscode/mcp.json), and TOML [mcp_servers.execkit] for Codex (~/.codex/config.toml). It prints rather than edits your client’s live config, so it cannot corrupt one. See Wiring into an agent.

doctor

Reports the resolved audit destination and its writability, the SSH key directory and known_hosts (with the env var that overrides each), and whether the Docker daemon is reachable. Use it after install to catch setup problems before an agent connects. See Installation.

watch [--follow|--serve] [--open] <path>

A live read-only viewer over the audit log; a file or a directory. --follow gives a headless, pipeable stream instead of the TUI; --serve serves a loopback, token-gated web page instead, and --open also launches your browser at it. See Auditing and the watch viewer.

Environment

These configure the server (operator-controlled, not agent arguments). Full table and rationale on the Security model page.

EXECKIT_MCP_AUDIT                  Append a JSONL audit log of every command here
EXECKIT_MCP_AUDIT_DIR             One JSONL file per session in this directory
EXECKIT_MCP_AUDIT_RETENTION_DAYS  Prune per-session files older than N days (default 14)
EXECKIT_MCP_WATCH_WEB             Auto-start the loopback browser viewer and surface its URL
EXECKIT_MCP_WATCH_PORT            Port for the browser viewer (default 7878; random if taken)
EXECKIT_MCP_WATCH_OPEN            Also auto-open the browser at the viewer URL (default: link only)
EXECKIT_MCP_EXEC_TIMEOUT          Default per-call exec timeout in seconds (default 120, clamped 1-3600)
EXECKIT_MCP_KEY_DIR               Directory SSH keys must live under (default ~/.ssh)
EXECKIT_MCP_KNOWN_HOSTS           execkit-managed SSH known_hosts file (default ~/.execkit/known_hosts)
EXECKIT_MCP_MAX_SESSIONS          Soft cap on concurrent live sessions (default 64)
EXECKIT_MCP_SESSION_TTL           Reap sessions idle longer than N seconds (default 1800)
EXECKIT_MCP_POLICY_FILE           JSON allow/deny + deny_patterns the agent cannot edit (advisory)

Rust library

The execkit crate is the core. The MCP server is a thin wrapper over it; you can embed the same sessions directly in your own program.

[dependencies]
execkit = "0.9"                                          # local + SSH + Docker
# execkit = { version = "0.9", default-features = false } # local + Docker only (no SSH; drops russh/tokio)
use execkit::{Session, Policy};

fn main() -> Result<(), execkit::Error> {
    let mut s = Session::local()?
        .with_policy(Policy { allow: vec![], deny: vec!["rm".into()] });

    let r = s.exec("echo hi; echo err 1>&2; cd /tmp")?;
    // r.stdout == "hi", r.stderr == "err", r.exit_code == 0, r.cwd == "/tmp"
    println!("{} (exit {})", r.stdout, r.exit_code);
    Ok(())
}

State persists across exec calls on the same Session, exactly as it does over MCP. Results are the same structured ExecResult (split stdout/stderr, exit code, duration, cwd, truncated, timed_out), ANSI-stripped and secret-redacted.

Timeouts

Each exec has a timeout: 30 seconds by default. Change it with with_timeout when building the session, set_timeout on a live one, or pass one for a single call with exec_with_timeout:

#![allow(unused)]
fn main() {
use std::time::Duration;

let mut s = Session::local()?.with_timeout(Duration::from_secs(60));
let r = s.exec_with_timeout("sleep 30", None, Duration::from_secs(1))?;
assert!(r.timed_out);            // interrupted with Ctrl-C
assert_eq!(r.exit_code, 124);
assert_eq!(s.exec("echo still here")?.stdout, "still here");
}

A timed-out command is interrupted with Ctrl-C and returned as Ok with timed_out: true; the session keeps its cwd and env. If the command ignores Ctrl-C, exec returns Error::StillRunning and the session is poisoned: later calls return Error::SessionPoisoned and is_poisoned() is true. A command that exits the shell (exit, or a failure under set -e) returns Error::ShellExited and poisons it the same way. Open a new session in either case.

SSH and Docker

SSH and Docker sessions are constructed with their configs:

#![allow(unused)]
fn main() {
use execkit::{Session, SshConfig, SshAuth, HostKeyVerification};

let cfg = SshConfig::new("web-01", "deploy",
    SshAuth::Password("...".into()),
    HostKeyVerification::KnownHosts("/home/me/.execkit/known_hosts".into()));
let mut s = Session::ssh(cfg)?;
}

The known_hosts file uses execkit’s own host SHA256:<fingerprint> format (keyed [host]:port on ports other than 22), not OpenSSH’s; use a file of its own. SshConfig::connect_timeout (default 15 seconds) bounds the TCP connect, key exchange and authentication together.

The API surface stays small; the richness lives in the result, not the verbs:

Session::local() / ::ssh(cfg) / ::docker(container)        -> Session
session.exec(command)                                      -> ExecResult
session.exec_budgeted(command, &budget)                    -> ExecResult
session.exec_with_timeout(command, Option<&budget>, dur)   -> ExecResult
session.checkpoint(label?) / restore(id) / restore_last()  -> CheckpointId / restore report

Runnable examples live in the repository:

cargo run --example local
EXECKIT_SSH="user:password@host:22" cargo run --example ssh

Full API docs are on docs.rs/execkit.

Python SDK

execkit-py wraps the same Rust core with a Python API, published to PyPI as execkit.

pip install execkit
from execkit import Session, Policy

with Session.local(policy=Policy(deny=["rm"])) as s:
    r = s.exec("echo hi; echo err >&2; cd /tmp")
    print(r.stdout, r.exit_code, r.cwd, r.stderr)   # hi 0 /tmp err

The result object mirrors the Rust ExecResult: command, split stdout / stderr, exit_code, duration_ms, cwd, truncated and timed_out, already ANSI-stripped and secret-redacted. State persists across exec calls on the same session.

Timeouts

timeout= on a session constructor sets the default per-command timeout in seconds (30 if omitted); exec(cmd, timeout=...) overrides it for one call:

with Session.local(timeout=60) as s:
    r = s.exec("sleep 10", timeout=1)
    print(r.timed_out, r.exit_code)   # True 124
    print(s.exec("echo still here").stdout)

A timed-out command is interrupted with Ctrl-C and the session keeps going. Only a command that ignores Ctrl-C raises execkit.Timeout; a command that exits the shell raises execkit.ShellExited. Both are subclasses of SessionUnusable: open a new session after either. s.is_poisoned tells you the same thing.

stdin is closed, so commands that would prompt get end-of-file instead of hanging. Use non-interactive flags such as sudo -n.

Other transports

SSH and Docker sessions work through Session.ssh(...) / Session.docker(...) with the same options as Transports. Output budgets are keyword arguments (tail, head, grep, max_chars) on the session constructors and on exec. Checkpoints are not exposed in the Python SDK yet; use the Rust library or the MCP server for those.

Wheels ship for Linux and macOS, so no Rust toolchain is needed to install.