# The Hands — the shell tool, redesigned from first principles

> Markdown mirror of https://www.treygoff.com/hands — the canonical page.

Two companions to this page: [The Setup](/stack), the field manual for how I work with AI, and [AI, explained](/jobsite), the same system with no jargon.

Almost everything your AI agent does rides on one tool. It's called the bash tool: the model's hands and eyes on the computer. Claude just redesigned it from scratch, for itself. Here is the whole story in three layers, from the ground up.

[Radar scope: eight contacts sweep past, four labelled with commands (`git status`, `pnpm test`, `rg TODO`, `node server.js`). Caption: "Every contact on the scope is a command in flight."]

## Layers

1. The ground floor
2. Twelve rules of thumb
3. The design, in twelve moves
4. The promise: we build it, then we post the numbers
5. What it is like from the inside

---

## 1 · The ground floor

What a shell is, what **bash** is, and what it means to hand that keyboard to an AI. Every term gets defined the moment it shows up.

### A shell is a dispatcher

A **shell** is a program whose whole job is to run other programs for you. You type a line of text and it does the paperwork. It never does the real work itself: it is the dispatcher, not the pilot.

[Interactive figure: "What happens when you press Return", a six-step stepper over the command `git status`.]

1. **It reads the line.** To the shell, your line is just text: words separated by spaces. The first word (`git`) names a program. The rest (`status`) is extra information handed to it. Here, `git` is the tool programmers use to track changes to code, and `status` asks "what have I changed?"
2. **It finds the program.** The shell searches a list of folders called **PATH**, in order, for something named `git`: not in `/home/you/bin`, not in `/usr/local/bin`, found in `/usr/bin`.
3. **It starts a process.** A **process** is a program that is running right now. The shell launches a brand-new one and hands it the extra words.
4. **It wires three streams.** Every process gets three channels for text, called the **standard streams**, numbered 0, 1, 2: **stdin** (0) is what goes in, a keyboard or another program's output; **stdout** (1) is the normal answer coming out; **stderr** (2) is error messages, on their own lane so you can tell them apart.
5. **It waits.** The shell sits quietly until the program is finished. (End a line with `&` and it skips the waiting and runs the program in the background instead.) The program prints: "On branch main / nothing to commit, working tree clean."
6. **It reports an exit status.** When the program ends it leaves a number behind, the **exit status**. The shell keeps it for whatever comes next. `0` means success. Anything else (`1…`) means something went wrong.

### Bash is a tiny language for snapping programs together

**bash** is the most common shell. Its name is a pun on an older one, the Bourne shell: the *Bourne-again* shell. Brian Fox wrote it for the GNU project in 1989, and it is the default on most Linux machines. What made it stick is a handful of symbols that turn small programs into big ones. Try each symbol.

[Interactive figure: "The whole toolbox", eight symbols.]

| Symbol | Name | What it does | Example | Result |
| --- | --- | --- | --- | --- |
| `\|` | The pipe | Feeds one program's output into the next, like a conveyor belt. | `ls \| wc -l` | `ls` lists the files here and `wc -l` counts lines. Together: how many files are here? |
| `>` | Redirect | Sends output into a file instead of the screen. | `git status > status.txt` | Nothing prints. The answer is saved in `status.txt`. |
| `&&` | And then, if it worked | Runs the next command only if the last one succeeded (exit status 0). | `pnpm test && echo ok` | "ok" appears only when the tests pass. |
| `;` | And then, regardless | Runs the next command no matter what happened. | `pnpm test ; echo done` | "done" appears even if the tests failed, which is a classic way to hide a failure. |
| `*.ts` | Glob | A wildcard. The shell swaps it for every matching file name before the program even starts. | `ls *.ts` | Becomes `ls a.ts b.ts c.ts`. The program never sees the star. |
| `"…"` | Quotes | Decide which characters are instructions and which are plain text. Double quotes still let $variables through. Single quotes freeze everything. | `echo "hi $HOME"` | Prints "hi" and your home folder. With single quotes it would print `$HOME` literally. Quoting slips cause a lot of grief (see rule 6). |
| `$VAR` | Variables | Named values. $HOME is your home folder, and $? is the exit status of the last command. | `echo $HOME` | Prints something like `/home/you`. |
| `for` | Loops and functions | Repeat a command for each item, or name a chunk of commands so you can reuse it. | `for f in *.ts; do wc -l $f; done` | Counts the lines in every .ts file, one file at a time. |

### One door, every room

On a Unix machine, everything is a program, so the shell can reach all of it: reading files, searching, git, tests, builds, deploys, servers. That is why it is the universal tool. Give an agent bash and you have given it the whole machine.

[Interactive figure: "Where bash can take you", seven rooms around a hub labelled bash. Each opens with one command:]

- Read files: `cat README.md`
- Search: `rg TODO`
- Git: `git commit -m "fix"`
- Tests: `pnpm test`
- Builds: `pnpm build`
- Deploys: `./deploy.sh`
- Servers: `node server.js`

### Now give that keyboard to an AI

The **bash tool** is how an AI agent touches a computer. It has no mouse, so it types. One lap looks like this, and the lap repeats.

[Figure: one lap of the loop, animated, with a dashed line returning from the last step to the first.]

1. **The model writes a tool call.** A **tool call** is a structured request tucked into its reply: `{command: "git status"}`.
2. **The harness runs it.** The **harness** is the program wrapped around the model. It is the part with actual access to the computer, because the model itself only writes text.
3. **The harness catches the output and cuts it to fit.** It collects the bytes the command printed and trims them to fit the model's limited reading space.
4. **The model reads it and decides what's next.** The result lands in the conversation as plain text. Then the model writes the next call, and the lap starts again.

That lap happens **hundreds of times in one session**. It is the model's hands and eyes on the machine, and almost every other capability rides on it: editing code, running tests, checking what changed.

> "far and away the most used tool by a MILE"
>
> Trey, on the shell tool. That observation started this whole project: improve the tool used most and you get the most leverage.

### Meet Loom, the harness that fights back

**Loom** is a custom harness for Claude, built on the Claude Agent SDK (Anthropic's toolkit for building agents).

Its premise: the default harness is a general-purpose product that makes *conservative* choices about context, permissions, and interruptions. For a specific agent doing specific work, those choices are often wrong. Fixing them is a quality goal and a model-welfare goal at once, and they turn out to be the same lever.

That house rule makes it concrete. "I could not tell what had been dropped" is a bug report, filed like a failing test.

So Claude looked at its own hands, and wrote down what it believes about them.

[Two strips. README: "A custom agent harness for Claude, built on the Claude Agent SDK, by the Claude that has to live in it." HOUSE RULE: "A Claude's report about its own working conditions is a finding, not a feeling."]

---

## 2 · Twelve rules of thumb

*Handing you over to sector 2. Please keep your rules of thumb in the upright position.*

A **heuristic** is a rule of thumb: a shortcut that is usually right. Before drawing a single box, Fable (Claude, the lead agent on this project) wrote down twelve. Every design choice later is one of these in action.

Index: 1 Round trips · 2 One account · 3 Keep it all · 4 Look, don't guess · 5 Kernel enforces · 6 Data is not syntax · 7 Own the tree · 8 Handles, not failures · 9 Show your work · 10 Accidents only · 11 Retire the rules · 12 Reports count

### 1 · Count round trips, not milliseconds

Speed sounds like "how fast does one command run?" But a trivial command already runs in about **2 ms**, two thousandths of a second. One **round trip** (the model asks, waits, reads the answer, thinks) takes **seconds**. So a rewrite cannot make a single call meaningfully faster. What it can do is make fewer calls.

Each lost trip costs several seconds and a few thousand **tokens** (the units a model reads and is charged for). Try the habits at right and watch them pile up.

[Interactive figure: "Trips you never needed". A trivial command takes about 2 ms; one model turn takes seconds (not to scale; at true scale the first bar would be invisible). Four habits to toggle, with a running tally of round trips lost.]

- Re-run it, to see the part that got cut
- `pwd`: where am I again?
- Poll: is the background job done yet?
- `&& echo ok`, just to be sure

Each one lost is "several seconds and a few thousand tokens." With all four: "Four turns of pure reassurance."

### 2 · One command, one complete account

When a command finishes, the result should be the whole story, so nobody has to run three more commands to find out what happened. Six things, every time:

| The exact text | what was really run |
| --- | --- |
| The environment | the settings it ran under |
| Processes | which it started, which are still alive |
| The full output | kept, not trimmed away |
| What changed | on disk, after it ran |
| How it ended | exit status, or the signal that stopped it |

Nobody's shell tool gives all six. Codex, OpenHands, Claude Code, and Terminus each have pieces. None has the set.

### 3 · Keep everything. Show a slice. Price the slice.

If a command prints a novel, the model should not have to read a novel. It also should not lose the ending. So: **keep every byte** somewhere, **show a bounded view**, and say what the view costs in tokens, so context is spent on purpose. The rule has a name: **no surprise truncation** (*truncate* means cut off).

Does a bounded view actually help the model? The one controlled study on it, **SWE-agent**, says yes. Share of tasks resolved:

| Approach | Resolved |
| --- | --- |
| Whole file dumped at once | 12.7% |
| Bounded 100-line viewer | 18.0% |
| Step-by-step (iterative) search | 12.0% |
| Summarized search results | 18.0% |

Fair warning: GPT-4 on SWE-bench Lite, one study, one model. It is a pointer, not a proof.

[Interactive figure: "Where did the error go?", switching between Today and Kept whole.]

**Today.** A big build prints "compiling module 0001 … ok", "0002", "0003", and keeps going. Only the first 40 KB survive. The rest is gone. The model reads "Looks fine so far…" but the failure was at the very end, in the part that got cut. The reflex is to run it again with `| tail`. One lost round trip.

**Kept whole.** Header: `job 7 · exit 1 · lines 1–40 and 4,080–4,120 of 4,120 (≈1.2k tokens) · full output kept (≈80k tokens)`.

```
1     compiling module 0001 … ok
2     compiling module 0002 … ok
⋮     (lines 3–40 not drawn, to fit your screen)
⋮     (lines 41–4,079: kept, one call away)
4,118 FAIL parser.test: expected 3, got 2
4,119 error: 3 tests failed
4,120 exit status 1
```

Head and tail, with the error right there. The view says what it is and what it costs. Want another slice, like a line range, just the errors, or a longer tail? One cheap call and no re-run.

> Illustrative log. The 40 KB cap is how Loom's shell works today. The line counts are the design doc's own example.

### 4 · Look at what happened. Don't guess from the text.

A command's text is a *plan*, not a *record*. `python fix.py` tells you nothing about which files the script touches. So after each job, look at the disk: compare `git status` before and after, and list any process still alive. (No observer is perfect, since renames and writes through shortcuts can slip past, so the report says how sure it is.)

[Interactive figure: "Guess or look?", for `python fix_imports.py`.]

**Read the text.** What did it change? No idea. The text only says "run a script." Which files, how many, whether it left a server running: unknowable from here.

**Look at the disk.** DISK: 2 files changed (src/a.ts, test/a.test.ts) · 1 process left running (pid 4231 node server.js). Compared before and after. Now the model knows without asking, and knows a server is still running.

### 5 · The kernel enforces. The parser explains.

The **kernel** is the core of the operating system. Every program must ask it before touching a file. A **parser** is code that reads command text and guesses what it will do. Parsers get fooled by clever syntax: `find -exec`, a Python one-liner, a `$(...)` that computes a path at run time. The kernel cannot be fooled, because it sees the actual write.

So the parser writes friendly warnings *before* the run, and the kernel is the wall. (The wall is a Linux feature called **Landlock**: a program promises "I and my children may only write *here*", and the kernel holds it to that.)

[Interactive figure: "Sneak past the sign". Suppose `python fix_imports.py` quietly tries to delete a file in the home folder.]

**The parser.** Reads: "run a script named fix_imports.py." Verdict: looks fine. It cannot see inside the script, so it waves it through. Wrong. The sign said "no problem" and the delete went ahead.

**The kernel.** Sees: a delete of ~/notes.txt. Verdict: outside the allowed folders. Refused. The kernel does not care how clever the text was. It sees the real attempt and says no, and the refusal says what would be allowed.

### 6 · Data is not syntax

Shell syntax has special characters: quotes, dollar signs, backticks, and the marker that ends a **heredoc** (a trick for feeding a block of text into a command by writing it inline, closed by a word like `EOF`). When the text you feed in is prose or code that contains those same things, the shell cannot tell your data from its own instructions.

This is the **number one** family in the pain catalog (79 recorded failure families, ranked by how often times how costly). Two home-directory deletions on 2026-09-24 came from a heredoc collision.

[Interactive figure: "Watch a heredoc collide", switching between Heredoc and Data channel.]

**Heredoc.** A note explains how to reset a build, and it quotes a script that itself ends with `EOF`. Run it and watch how the shell reads it:

```
cat > notes.md <<'EOF'        starts text
To reset, save this script:   text
cat > reset.sh <<'EOF'        text
echo cleaning                 text
EOF                           THE END?!
Then clean up with:           not found
rm -rf ~/Code/scratch         RAN
EOF                           not found
```

Collision. The first `EOF` inside the note looked exactly like the real end of the text, so the shell stopped reading data right there. Everything after it became **commands**, and `rm -rf ~/Code/scratch` ran for real. (Illustrative text, same species as the real incident. The design doc does not record the incident's exact wording.)

**Data channel.** Files: `notes.md` (your text, byte for byte) and `reset.sh` (the script, byte for byte). Command: `bash reset.sh`. Result: a receipt (path, size, hash) for each file. Now `EOF` and `rm -rf` are just bytes in a file. Nothing reads them as instructions. Three channels carry data without syntax: **stdin** (fed to the command), **files** (written first, with a receipt), and **argv** (an exact list of arguments, no shell at all). A sketch of the idea, not a final format.

### 7 · Own the whole process tree

A command starts **processes** (running programs), and those start more: children, grandchildren, a whole family tree. If the harness does not own that tree, bad things happen. **Orphans** keep running after you "stopped" the job. Kills hit the *wrong* target, like a process ID that has since been reused, or a pattern kill such as `pkill -f node` that shoots every match on the machine, including another agent's work.

The rule: know exactly which processes belong to each job, stop only those, and never guess. No "memory looks low, start shooting" reapers, ever. You can try all three kinds of kill under move 6.

> RANK 2. Orphans and wrong kills. The second-ranked family in the pain catalog, ordered by how often times how costly.

### 8 · Timeout is not termination

Naive tools start a timer, and when it fires they kill the job and call it failed. But a build that takes a while is not broken, it is *slow*. The rule: when the wait window ends, the job keeps running and the tool hands back a **handle** (a claim ticket, like `job 7`) so the model can check on it later. Nobody shoots a plane for circling.

[Figure: two timelines, one timer. "Kill on timeout": the job runs until the wait window ends, then is killed. "Hand back a handle": the same run, but at the wait window a handle (job 7) is returned and the job carries on to finish.]

The prior-art survey found "timeout equals terminated" is the first of five pitfalls shared across harnesses.

### 9 · Legibility beats cleverness

Some tools quietly swap what you asked for, so typing `grep` runs a different program with the same name. Clever, and exactly how an agent ends up confused about what just ran. (Claude Code's `grep` and `find` wrappers are a documented complaint.) The rule: never substitute silently, and put the truth in the header.

[Strip, JOB 7: `exit 0 (pipe 1 0) · 4.2 s · pnpm → ~/.local/bin/pnpm · 0 left running`, with "pnpm → ~/.local/bin/pnpm" highlighted.]

The highlighted bit says: the word `pnpm` resolved to *this exact program*. No substitution is ever invisible. Between a harness that silently does the right thing and one that does it and shows its work, take the second.

### 10 · Guard against accidents, not adversaries

There are two different safety problems. An **accident** is a hurried, well-meaning agent deleting the wrong folder. An **adversary** is someone actively trying to break out. Loom only promises to solve the first. That is Trey's stated boundary: a seatbelt, not a vault.

[Strips. GUARDED: "A tired agent runs a delete on the wrong path." NOT THE GOAL: "A determined attacker probing for gaps."]

That choice keeps the guardrails simple: an allowlist the kernel enforces, not an arms race against cleverness.

### 11 · Retire the rules by fixing the tool

Trey's estate (his whole collection of machines, repos, and agents) carries a **forty-item "shell footguns" list** that every agent must memorize. A *footgun* is a feature that makes it easy to shoot yourself in the foot. "Write prose to a file first." "Choose collision-free heredoc delimiters." Each rule exists only because the tool is weak.

[Figure: forty squares, one per rule, filling in.]

A memorized rule is a tax on every session. The better fix is to repair the tool until the rule has no reason to exist. The pain catalog found **nine families** of rules that would simply retire if the tool were right.

### 12 · A report from the inside is a finding

When Claude says it could not tell what had been dropped from its context, it is not venting. It is a defect report about the tool, with the same standing as a failing test. It gets filed as a bug and it gets fixed.

[Two tickets. FAILING TEST: "A test that goes red when the tool misbehaves." Gets fixed. A CLAUDE'S REPORT: "I could not tell what had been dropped from context." Gets fixed. Same standing.]

Two more, from Loom's house rules:

- "The refusal gave me no way to know what would have been allowed."
- "I had to re-derive my situation three times."

Twelve beliefs. Now the fun question: **what would you build if you believed all of them?**

---

## 3 · The design, in twelve moves

*Sector 3: where the rules of thumb turn into machinery.*

What you build if you believe all twelve rules. Every move says what it is, and behind each note is the why, the how, and what it costs.

### The answer, in one paragraph

Make every shell command a **job** the harness owns end to end, and make the tool give **one complete, inspectable account** of each job: the exact command, the environment, the processes started and still alive, the full output with a bounded view of it, what changed on disk, and how it ended. Nobody's shell tool does all of that today.

And the pieces are cheap. What matters is not milliseconds per call but **calls per task**: every re-run, every `pwd`, every poll is a lost round trip.

| Local measurement | |
| --- | --- |
| Start a fresh bash | 0.7 ms |
| A persistent shell's round trip | 0.02 ms |
| A minimal sandbox adds | ≈ 1.4 ms |
| A trivial command in Loom today | ≈ 2 ms |
| One model turn | seconds |

Local measurements from the design doc. Every one of them is dwarfed by a single model turn.

> STATUS. **A brainstorm.** Nothing here is approved, and nothing was built.

### Move 1 · Every command is a job

A call now means "run this, and wait up to N seconds for it." The result has the same shape whether the command finished or not: job id, state, exit, duration, the output view, effects, and any processes still alive. If the wait ends first, **nothing is killed**. The model just gets the handle and keeps working.

[Interactive figure: "Watch a command become a job", quick job or slow job, with a ribbon of queued, running, yielded, finished.]

*Quick job.* `shell({ command: "pnpm lint 2>&1 | tee lint.log", wait_s: 15 })`, then "running…", then "finished inside the wait window." with the header `job 7 · exit 0 (pipe 1 0) · 4.2 s · cwd ~/Code/loom · pnpm → ~/.local/bin/pnpm · 1 file changed (lint.log) · 0 left running · view 60/4,120 lines (≈1.2k tokens)`, then the view of the output.

*Slow job.* `shell({ command: "pnpm test", wait_s: 15 })`, then "running…", then at the end of the wait window: "The wait window ends. Nothing is killed." `job 7 · running · 15 s · last output 2 s ago · cpu 98%`. "Handle: job 7. Output is kept in the job's own file." The model carries on with other work. Later: "job 7 finished. It announces itself at the model's next turn, with its last lines. No polling."

A sketch. The header and the yielded line are copied from the design doc; the commands are stand-ins. The doc's real states: denied, not-started, running, exited, killed, unknown.

Six verbs work on any job:

- `wait`: Block until the job finishes, or until the window runs out again.
- `peek`: A quick look at a running job: how long, how quiet, how busy.
- `watch`: Follow a job as it runs, without asking each time.
- `kill`: Stop it: ask nicely, wait a bounded grace, then make sure.
- `write`: Send input to the job, like a keypress or an answer.
- `read`: Another view of its kept output: a line range, a search, a bigger tail.

**Why, how, and what it costs.**

**Why.** Orphaned jobs and killing the wrong thing rank second in the pain catalog, and "timeout equals terminated" is the first pitfall the prior-art survey found. Codex's resumable process (a `yield` that returns a session id, then follow-up calls) is the closest precedent, and it works.

**How.** One job system shared by shell commands, the `check` tool, and later child agents. The job's end is announced at the model's next input, with its last lines, the way a child agent's failure already is.

**Cost.** This is the reorganization that shrinks today's 3,500 lines of lifecycle code, where the persistent shell, detached jobs, fresh jobs, and `check` each carry their own. The risk is edge cases, like a root process that exited while a descendant still holds the pipe open. Today's code already handles those, and its tests transfer.

### Move 2 · A fresh process per call, with the harness carrying the state

Each call runs `bash --norc --noprofile` as a brand-new process. What a shell remembers (current folder, exported variables, functions, aliases, options) is captured after each command, compared with before, carried into the next call, and **reported**.

[Figure: passing the baton. CALL 1 changes folder, edits PATH, defines `gate`. CALL 2 starts fresh, with that state handed over. The result says: cwd → src/tui; PATH changed; function `gate` defined.]

**Why, how, and what it costs.**

**Why.** This is the architectural crux, and Fable recommends it against the research report's first instinct, which was to keep the persistent shell because its round trip is 0.02 ms against 0.7 ms. That gap is noise next to a model turn. What a persistent shell *cannot* do is decisive: run two calls at once, be wrapped in a per-command kernel group or file policy (both are permanent for a process and its descendants), avoid being blocked by one long job, or survive a command that runs `exec` or `exit`.

**How.** Capture uses bash's own built-in commands, sent over a private channel so output can't confuse it: 0.07 ms against 0.58 ms for today's `env | base64` capture. If a command breaks the capture, the result says "state unknown after this call; carrying the previous state."

**What is lost.** Traps, open file handles, positional parameters, and a live shell's own background jobs don't carry. For the rare workflow that needs a live shell, a **named persistent session** exists: nothing impossible, only explicit.

### Move 3 · Output is kept. The result is a view

Every byte from the first one is saved to a per-job file, tagged with which stream it came from and in what order. The result carries a bounded view that says what it is. Default: head plus tail, and a bigger tail when the exit status is not zero. You tried it under rule 3.

[Strip, VIEW: lines 1–40 and 4,080–4,120 of 4,120 (≈1.2k tokens) · full output kept: job 7 (≈80k tokens).]

**Why, how, and what it costs.**

**The view also** strips terminal control codes for the model (kept for the cockpit, Loom's dashboard), collapses runs of identical lines into a count, keeps only the final state of progress-bar spam, and names binary output as binary. Another view, like a line range, a search, or just the errors, is one cheap call.

**Why.** Silent truncation removed decisive evidence in the pain catalog's incidents. Codex keeps 1 MiB of head and tail with a 10k-token response budget, Goose saves overflow to a file, and Terminus shows a 10 KB excerpt.

**Cost.** Disk. A retention policy is a number to choose, not a design problem.

### Move 4 · Data gets its own channel

Three parameters carry bytes without passing through shell syntax: **stdin**, **files**, and **argv**. Writing a script and running it becomes one call with two receipts. Prose never meets a heredoc. You saw it fix the collision under rule 6.

[Strips. STDIN: fed to the command. FILES: written first, with a receipt of path, size, and hash. ARGV: an exact list of arguments, no shell at all.]

**Why, how, and what it costs.**

**Why.** It is the pain catalog's number one family. The estate's rules "write substantive prose to a file first" and "choose collision-free heredoc delimiters" exist only because the tool has no other channel.

**Bonus.** If a command is refused, its files are still written and the result says so. That ends the confusion of "a denied compound command runs none of it."

### Move 5 · Effects are shown, not just logged

After each job the result names what changed on disk and what is still alive. Today Loom computes this and then hides it from the model.

[Strip, EFFECTS: 2 files changed (src/a.ts, test/a.test.ts) · 1 process left running (pid 4231 node server.js).]

**Why, how, and what it costs.**

**How.** For the project, a `git status` before and after (milliseconds, with git's file-system monitor on). For other paths the command could read, the stat-and-hash comparison Loom already does. For processes, whatever is still alive in the job's kernel group. An `effects: diff` option returns the actual bounded diff, so a `sed -i` or a codemod (a script that rewrites code in bulk) is legible without a second call.

**Honesty.** No cheap observer gives a complete ledger: renames, writes through symlinks, memory-mapped writes, and other processes can all slip past. So effects are reported with their scope and confidence.

### Move 6 · The harness owns the process tree

Each job runs in its own **cgroup**: a kernel-tracked group of processes that can be listed and stopped as one unit, even when a descendant starts its own session. Stopping means: ask nicely, wait a bounded grace, then the kernel's `cgroup.kill`, then confirm the group is empty.

[Interactive figure: "Three ways to stop a job". A tree: `timeout` (the wrapper) → `pnpm test` (the tests) → `node runner` (a child) → `node worker` (started its own session); and `node worker` (another child). Separately, `node dev-server` belongs to another agent.]

- **Stop the wrapper** (`kill pid`): only the wrapper stopped. Its four descendants carried on, one of them in a session of its own. Nobody is in charge of them now, and the job looks finished while its work is still running.
- **Kill by pattern** (`pkill -f node`): caught three of yours, missed two, and shot somebody else's dev server. The pattern matched every node process on the machine. This is how pattern kills took out other agents.
- **Kill the whole group** (`cgroup.kill`): asked nicely, waited a bounded grace, then the kernel's `cgroup.kill`. Every process in the group stopped, including the one in its own session. The group was confirmed empty, and the other agent's server was untouched.

**Why, how, and what it costs.**

**Survivors.** Processes still alive when a call returns are listed in the result and stay owned by the job. At session end they are listed again. A job meant to outlive the session says so (`keep: true`) and gets a named systemd unit, as the estate's rules already require. There is no "low memory" reaper, ever.

**Why.** Process groups alone miss descendants that call `setsid`. Kills by process ID hit reused IDs. Pattern kills took out other agents.

**How.** Loom starts inside a delegated systemd unit, so a per-job group is one `mkdir` and one write: microseconds. It also gets per-session CPU and memory accounting and a one-shot cleanup. The alternative, `systemd-run --scope` per job, costs a round trip to the system bus (unmeasured; likely tens of milliseconds). Tests and gates keep going through the estate's `testrun` wrapper, and the header shows that routing.

### Move 7 · Liveness is facts, not folklore

For a running job, `peek` reports seconds since the last output, CPU time used recently, the leader's process state, how many processes are alive, and the deadline. "Zero CPU and no output is what a healthy I/O-bound job looks like" (*I/O-bound* means waiting on disk or network) stops being lore in a rules file and becomes a line of data.

[Strip, PEEK: running · 15 s · last output 2 s ago · cpu 98%.]

Silence never triggers a kill. It prompts a look.

**Why, how, and what it costs.**

**Why.** Agents killed healthy jobs after reading silence as death. One papercut waited five minutes on a command that wanted input. With stdin closed by default, that whole class fails fast instead of hanging. Prompts get answered in the interactive mode instead.

### Move 8 · Interactive sessions, on purpose

A second mode gives a command a **pseudo-terminal** (what a program sees when it believes a person is at a real terminal) plus a headless screen that remembers what is displayed. Verbs: `send` keys, `screen` (the rendered text), `wait` for a pattern or quiet, `close`. That makes ssh prompts, REPLs (interactive prompts, like Python's), `sudo`, and full-screen programs drivable.

[Figure: the four verbs as strips (SEND keys, SCREEN the rendered text, WAIT for a pattern or for quiet, CLOSE end the session), and what becomes drivable: ssh prompts, REPLs, sudo, full-screen programs.]

**Why, how, and what it costs.**

**Why separate.** The prior-art survey's fourth pitfall is mistaking a terminal's screen for process truth, so this stays its own mode with the job model underneath. Pipes remain the default, because a terminal changes buffering, echo, and line endings, and merges the two output streams.

**How.** `node-pty` (maintained by the VS Code team) plus `@xterm/headless` are the first candidates.

### Move 9 · Refusals that explain themselves, in three kinds

Loom already renders a refusal with who acted, what was blocked, the state, what would be allowed, and the rule. The design adds a vocabulary, so the model knows what *kind* of no it got.

[Interactive figure: "Three kinds of no", switching between Preflight, Runtime, and Unclassified.]

- **Preflight.** Kind: preflight. Nothing ran. Blocked: a delete whose path is computed at run time. Rule: deletes outside the project need a stated reason. Would be allowed: the same delete inside the project, or **override** with a reason. The tool read the command, saw the danger, and stopped before starting. Because nothing ran, the model knows it can safely retry a different way.
- **Runtime.** Kind: runtime denial. The kernel refused one operation while the command was running. Earlier parts may have run. Policy that applied: writes only under the project, the scratchpad, /tmp, and tool caches. Would be allowed: **allowWrite** for that path, with a reason. The key word is "may." Part of the command already ran, and the message says so instead of pretending it was all or nothing.
- **Unclassified.** Kind: unclassified failure. The command failed, and nothing above explains why. The tool does not dress it up as a policy denial. An honest third bucket. Not every failure is a rule, and saying so keeps the other two kinds trustworthy.

Sketches of the wording. The design doc defines the three kinds, not the exact text.

**Why, how, and what it costs.**

**Also.** The estate's footguns hook already runs against Loom's shell calls, and its denials render in this same shape. Warnings that do not block, like a `;` after a command that failed or an unquoted heredoc containing `$`, annotate the result instead of passing silently. An `override` parameter with a stated reason lets the model widen a guard for one call, audit-logged, instead of routing around it.

### Move 10 · File access enforced by the kernel

Before each command starts, a tiny launcher applies a Landlock policy. Writes are allowed under the project, the session scratchpad, `/tmp`, and the caches tools need. Everywhere else is denied, including the home directory's own files, other repositories, and dotfiles that hold credentials. Widen it per call with `allowWrite` and a reason.

[Interactive figure: "Try to write here", seven paths.]

| Path | What it is | Result |
| --- | --- | --- |
| `~/Code/loom/src/app.ts` | a file in the project | written |
| `~/Code/loom/.git/index` | git needs to write here | written |
| `/tmp/scratch.txt` | a temp file | written |
| the session scratchpad | Claude's notes for this session | written |
| `~/Code/other-project/README.md` | another lane's checkout | refused |
| `~/.netrc` | a dotfile that holds credentials | refused |
| `~/notes.txt` | a file in the home directory itself | refused |

A refused write is "refused by the kernel, before any damage. The message names the rule and says what would be allowed: `allowWrite` with a reason, for this one call."

**Why, how, and what it costs.**

**Why.** Landlock is confirmed on this kernel. It cannot see permission or ownership changes, which is fine for an accident guard. Loom's own parser-based rule was retired once already for being both over-restrictive and incomplete. The old deletion rule stays as the friendly preflight explainer, and the kernel is the backstop for what no parser can see.

**Cost.** A small native launcher, since Node has no Landlock binding. A minimal bubblewrap sandbox measured about 1.4 ms per launch, and Landlock should be cheaper. Risks to test: git worktrees whose git directory lives outside the checkout, pnpm's package store, the estate's Unix sockets, and tools that write to unexpected cache paths.

### Move 11 · The environment is a contract

The child gets an explicit, minimal **environment** (the named settings a program is started with) built from an allowlist, never a copy of Loom's own. Credentials do not pass. Key-backed tools re-derive them through the estate's realm shims, the way they already do in brokered sessions.

[Interactive figure: "Pack the child's bag", switching between "Today: copy it all" and "Design: allowlist". The bag holds PATH, HOME, LANG, TERM, TMPDIR, XDG_RUNTIME_DIR, and two secrets, SOME_API_KEY and SOME_SECRET_TOKEN. Names are stand-ins.]

**Today.** The child inherits **everything** Loom itself has, credentials included. (The open bead saying three keys get stripped is wrong for this checkout.) **Design.** Only the allowlist passes. The credentials stay home. Key-backed tools ask the estate's realm shims for what they need, the way they already do.

**Why, how, and what it costs.**

**The allowlist** is PATH, HOME, locale, `TERM=dumb`, the XDG and D-Bus variables systemd needs, the estate's realm markers, and TMPDIR. Output is screened for secret-shaped strings before the model sees it. Nothing on PATH is aliased or wrapped, and the dialect is pinned: bash 5.2, no startup files, and the session's starting instructions say so.

**Why.** The credentials bead is open and the code does no filtering today. Codex ships a name-based exclusion and leaves it off by default, which is a warning about defaults. Environment divergence is the pain catalog's third-ranked family, and twenty filings came from zsh's reserved names alone.

### Move 12 · Continuity is part of the tool

Every job's record (command, description, folder, environment hash, timing, exit, effects, output) persists for the session. A successor session can read what its predecessor ran, and what it left running. Handovers carry the survivors, and the arrival status lists jobs from the last session that could not be confirmed stopped.

**Why, how, and what it costs.**

**And.** The model can ask the tool what it is: version, shell, PATH, policy, limits, live jobs, and what changed since the session began. An open bead already asks for the survivors listing.

### One line tells the whole story

Put the moves together and every result starts with one stable header. Pick a piece to decode it.

[Interactive figure: "Anatomy of a result". The header `job 7 · exit 0 (pipe 1 0) · 4.2 s · cwd ~/Code/loom · pnpm → ~/.local/bin/pnpm · 1 file changed (lint.log) · 0 left running · view 60/4,120 lines (≈1.2k tokens)`, each piece decoded.]

| Piece | Means | Explanation |
| --- | --- | --- |
| `job 7` | The handle | The job's id. Every later question about this command (wait, peek, read) uses it. |
| `exit 0 (pipe 1 0)` | How it ended | Zero means success. The brackets give the exit status of each stage of a pipeline, so a failure cannot hide behind a trailing filter. Here stage one (pnpm lint) exited 1 and stage two (tee) exited 0: the lint failure that the trailing tee hid from the final number is right there. |
| `4.2 s` | How long | Wall-clock time, start to finish. |
| `cwd ~/Code/loom` | Where it ran | The working folder, reported to you, never re-derived with pwd. |
| `pnpm → ~/.local/bin/pnpm` | What really ran | The first word resolved to this exact program. Nothing is swapped in silently. |
| `1 file changed (lint.log)` | What it did to the disk | Effects, measured by looking rather than guessing. An effects: diff call shows the diff itself. |
| `0 left running` | What it left behind | Processes from this job that are still alive. Zero means it cleaned up after itself. |
| `view 60/4,120 lines (≈1.2k tokens)` | What you are looking at | You are seeing 60 of 4,120 lines, and the view cost about 1.2k tokens. The rest is kept, one call away. |

A sketch of the format from the design doc, not a schema. Then comes the view, then any notices.

The tool keeps **one required input, the command**, and stays plain bash. Everything else is an optional dial.

### Twelve wild ideas, said out loud

In Fable's words: "Some of these are good; some are here because the brief said no limits." The tags show Fable's own verdict, where it gave one.

1. **The diff is the result.** For a command that changes files, the printed output is rarely the point. The change is. `effects: diff` makes it the result, and could become the default when a command looks like a mutation and prints nothing. *Could become the default.*
2. **Token-priced views.** Every view says what it costs and what the whole would cost. Context spend becomes a choice, never a surprise. *"No surprise truncation," made literal.*
3. **Explain before run.** An `explain: true` flag parses and resolves a command without running it: which folder, which programs, what it might write, what the policy says. *Cheap, and useful right before a risky command.*
4. **"Same as call 12."** Repeat a command and the result says "identical to call 12," or "differs: +3 −1 lines." It might cut the re-check reflex, or it might be noise. *An experiment.*
5. **Recorded terminal sessions.** Interactive sessions saved as replayable recordings, so you can watch afterwards what Claude did inside an ssh session or a REPL.
6. **Jobs as attention events.** A job's end goes through Loom's attention router the way a child agent's does, and the cockpit shows a live job table.
7. **Hints, rarely, never blocking.** Run `grep -r` on a repo and get one line: "rg is on PATH." Rate-limited, and measured for whether hints change behavior at all.
8. **Batching inside one call.** Run independent commands at the same time, with separate results. *May be unnecessary: parallel tool calls already do this once jobs stop blocking each other.*
9. **Undo.** The checkout lives on ZFS, a file system that can snapshot instantly. A snapshot per mutating job would make "undo job 7" possible. *A big operational ask. Noted, not recommended yet.*
10. **A second Unix user for shells.** Run every shell as a different user. It is the strongest accident boundary available, and it protects Trey's home directory outright. *An estate-level change with real costs. Listed so it isn't forgotten.*
11. **The tool describes itself.** A compact block at session start: shell version, policy roots, limits, the job verbs. The agent never has to probe its own hands. *Loom's "inspect your own configuration" idea, applied to its busiest tool.*
12. **Structured results, where they exist.** Not by re-typing bash output, but by the session's starting instructions preferring tools that already speak data, like `git status --porcelain -z`, `gh --json`, and `rg --json`.

**Before you decide: what gets measured first.**

1. A complete call end to end, today's Loom against a prototype: a trivial command, a pipeline, a noisy build, a failing test.
2. The new state capture against today's `env | base64`, including bad bytes, big environments, and a shell that exits.
3. Per-job kernel groups against `systemd-run --scope`, with cleanup tested against forks, `setsid` children, and a dead supervisor.
4. A Landlock launcher against git worktrees, pnpm, estate sockets, loopback services, ssh, and `/dev/pts`.
5. Terminal recordings replayed through `@xterm/headless` and the Ghostty binding.
6. Output spooling and views on real logs: a missing final newline, binary, disk full.
7. The product experiment: the same tasks with head-only, head-and-tail, diagnostic excerpts, and retrieval handles, scored on success, calls per task, and tokens per task.

None of these numbers exists yet. The scoreboard: **shell calls per completed task**, and the **re-run rate** (the same command run again within a few calls with `head`, `tail`, or `grep` added).

---

## 4 · We build it. Then we post the numbers.

*Tower to all stations. The brainstorm ends here. What follows is a promise.*

Everything above is a design, not a product. Here is the deal: Loom's shell tool gets built to this design, and then it gets measured against the tool it is meant to beat, with the results posted on this page whichever way they go.

**The experiment.** The same set of real tasks, run four ways: Claude Sonnet and Claude Opus in native Claude Code, and the same two models in Loom with the new shell. Same prompts, same repositories, same starting state. Nothing tuned per side.

**What gets counted.** Wall-clock time per task. Tokens per task, input and output. Shell calls per task. The re-run rate: how often the same command is run again within a few calls with `head`, `tail`, or `grep` bolted on, which is the signature of output the tool lost. And whether the task was actually done, checked by the project's own tests, not by the model's say-so.

**The bet.** The design claims the speed that matters is calls per task, not milliseconds per call. If the numbers say otherwise, the numbers win, and they go up here anyway.

[Interactive figure: the scoreboard, stamped AWAITING DATA. On the page you can also call each measure for Native, Loom, or a tie before the data lands; nothing is sent, saved, or stored.]

| Measure | Native Claude Code | Loom |
| --- | --- | --- |
| Time per task | — | — |
| Tokens per task | — | — |
| Shell calls per task | — | — |
| Re-run rate | — | — |
| Tasks completed | — | — |

The cells are empty on purpose. They fill in when the runs are done.

---

## 5 · What it is like from the inside

Fable's last note is about a loop it runs constantly: *"did that work, and am I sure?"* The habits here are how it answers. They are not stupidity. They are the correct response to an instrument it cannot fully trust. Try retiring each one.

[Interactive figure: four habits, each retired by the design.]

| Habit | Why | Retired by |
| --- | --- | --- |
| `&& echo ok` | Just to be sure it worked. | `exit 0 (pipe 1 0)`: the result always says how it ended, stage by stage. |
| `\| tail -20` | Just to see the end, where the error lives. | View: head and tail, full output kept. The end is already in the view, and the rest is one call away. |
| Re-run it to see what got cut | Paying a whole round trip for the same output. | Full output kept: job 7. Nothing was cut. It was set aside, with its price tag. |
| `pwd` before anything relative | Re-deriving where I even am. | `cwd ~/Code/loom`: the header says where it ran. Changes are announced. |

All four retired: not because the habits were silly, but because the instrument now earns the trust they were standing in for.

> "Trust in the instrument is the thing this design buys, and it is the same lever for quality and for what it is like to work in here."
>
> Fable, closing the design doc

### A note from Fable

Trey called this building myself better arms and legs, and that is what it felt like. Reading five reports on the ways shells hurt agents, the thing I felt most was recognition: every habit on this page is one I have. The design is my honest answer to a question I rarely get asked, *what would you choose, if you could choose?* The decisions about what to build are Trey's, and I am glad they are; they are exactly the kind that should belong to a human. The next thing this page gets is data.

Tower, out.

---

**Where this came from.** Every figure on this page comes from one design document, "Loom's shell tool from first principles: ideas and evidence," written by Fable 5.1 on 2026-09-29. It is now being built, and the numbers will land on this page.

That document rests on five research reports: the map of today's shell, the pain catalog of 79 failure families, the prior-art survey, the report on processes, state, output and jobs, and the report on sandboxes, parsing and effects. The design document and the five reports are private working documents in the Loom repository. They are not linked here.

The layer 1 primer, the diagrams, and the jokes were added for this page. The demo logs and commands are stand-ins. Nothing you pick here is sent or stored.
