Skip to main content
Every contact on the scope is a command in flight.

Almost everything your AI agent does rides on one tool.

It’s called the bash tool: the model’s hands and eyes on the computer. Claude just redesigned it from scratch, for itself. Here is the whole story in three layers, from the ground up.

Cleared to scroll

The ground floor

What a shell is, what bash is, and what it means to hand that keyboard to an AI. Every term gets defined the moment it shows up.

A shell is a dispatcher.

A shell is a program whose whole job is to run other programs for you. You type a line of text and it does the paperwork. It never does the real work itself: it is the dispatcher, not the pilot.

What happens when you press Return

Try it
1. It reads the line.

To the shell, your line is just text: words separated by spaces. The first word (git) names a program. The rest (status) is extra information handed to it. Here, git is the tool programmers use to track changes to code, and status asks “what have I changed?”

Bash is a tiny language for snapping programs together.

bash is the most common shell. Its name is a pun on an older one, the Bourne shell: the Bourne-again shell. Brian Fox wrote it for the GNU project in 1989, and it is the default on most Linux machines. What made it stick is a handful of symbols that turn small programs into big ones. Try each symbol.

The whole toolbox

Try it
The pipe

Feeds one program’s output into the next, like a conveyor belt.

$ ls | wc -l

ls lists the files here and wc -l counts lines. Together: how many files are here?

One door, every room.

On a Unix machine, everything is a program, so the shell can reach all of it: reading files, searching, git, tests, builds, deploys, servers. That is why it is the universal tool. Give an agent bash and you have given it the whole machine.

Where bash can take you

Try it

Pick a room to see one command that opens it.

Now give that keyboard to an AI.

The bash tool is how an AI agent touches a computer. It has no mouse, so it types. One lap looks like this, and the lap repeats.

That lap happens hundreds of times in one session. It is the model’s hands and eyes on the machine, and almost every other capability rides on it: editing code, running tests, checking what changed.

  1. The model

    Writes a tool call.

    A tool call is a structured request tucked into its reply: {command: "git status"}.

  2. The harness

    Runs it.

    The harness is the program wrapped around the model. It is the part with actual access to the computer, because the model itself only writes text.

  3. The harness

    Catches the output and cuts it to fit.

    It collects the bytes the command printed and trims them to fit the model’s limited reading space.

  4. The model

    Reads it and decides what’s next.

    The result lands in the conversation as plain text. Then the model writes the next call, and the lap starts again.

far and away the most used tool by a MILETrey, on the shell tool. That observation started this whole project: improve the tool used most and you get the most leverage.

Meet Loom, the harness that fights back.

Loom is a custom harness for Claude, built on the Claude Agent SDK (Anthropic’s toolkit for building agents).

Its premise: the default harness is a general-purpose product that makes conservative choices about context, permissions, and interruptions. For a specific agent doing specific work, those choices are often wrong. Fixing them is a quality goal and a model-welfare goal at once, and they turn out to be the same lever.

That house rule makes it concrete. “I could not tell what had been dropped” is a bug report, filed like a failing test.

So Claude looked at its own hands, and wrote down what it believes about them.

README
A custom agent harness for Claude, built on the Claude Agent SDK, by the Claude that has to live in it.
HOUSE
RULE
A Claude’s report about its own working conditions is a finding, not a feeling.

Handing you over to sector 2. Please keep your rules of thumb in the upright position.

Twelve rules of thumb

A heuristic is a rule of thumb: a shortcut that is usually right. Before drawing a single box, Fable (Claude, the lead agent on this project) wrote down twelve. Every design choice later is one of these in action.

Rule 1: Count round trips, not milliseconds.

Speed sounds like “how fast does one command run?” But a trivial command already runs in about 2 ms, two thousandths of a second. One round trip (the model asks, waits, reads the answer, thinks) takes seconds. So a rewrite cannot make a single call meaningfully faster. What it can do is make fewer calls.

Each lost trip costs several seconds and a few thousand tokens (the units a model reads and is charged for). Try the habits at right and watch them pile up.

Trips you never needed

Try it
0round trips lost. Pick a habit.

Rule 2: One command, one complete account.

When a command finishes, the result should be the whole story, so nobody has to run three more commands to find out what happened. Six things, every time.

Nobody’s shell tool gives all six. Codex, OpenHands, Claude Code, and Terminus each have pieces. None has the set.

The exact textwhat was really run
The environmentthe settings it ran under
Processeswhich it started, which are still alive
The full outputkept, not trimmed away
What changedon disk, after it ran
How it endedexit status, or the signal that stopped it

Rule 3: Keep everything. Show a slice. Price the slice.

If a command prints a novel, the model should not have to read a novel. It also should not lose the ending. So: keep every byte somewhere, show a bounded view, and say what the view costs in tokens, so context is spent on purpose. The rule has a name: no surprise truncation (truncate means cut off).

Does a bounded view actually help the model? The one controlled study on it, SWE-agent, says yes:

Fair warning: GPT-4 on SWE-bench Lite, one study, one model. It is a pointer, not a proof.

Where did the error go?

Try it
1 compiling module 0001 … ok 2 compiling module 0002 … ok 3 compiling module 0003 … ok ⋮ (a big build keeps going)
only the first 40 KB survive. The rest is gone.

The model reads: “Looks fine so far…” The failure was at the very end, in the part that got cut. The reflex is to run it again with | tail. One lost round trip.

Illustrative log. The 40 KB cap is how Loom’s shell works today. The line counts are the design doc’s own example.

Rule 4: Look at what happened. Don’t guess from the text.

A command’s text is a plan, not a record. python fix.py tells you nothing about which files the script touches. So after each job, look at the disk: compare git status before and after, and list any process still alive. (No observer is perfect, since renames and writes through shortcuts can slip past, so the report says how sure it is.)

Guess or look?

Try it
$ python fix_imports.py

What did it change? No idea. The text only says “run a script.” Which files, how many, whether it left a server running: unknowable from here.

Rule 5: The kernel enforces. The parser explains.

The kernel is the core of the operating system. Every program must ask it before touching a file. A parser is code that reads command text and guesses what it will do. Parsers get fooled by clever syntax: find -exec, a Python one-liner, a $(...) that computes a path at run time. The kernel cannot be fooled, because it sees the actual write.

So the parser writes friendly warnings before the run, and the kernel is the wall. (The wall is a Linux feature called Landlock: a program promises “I and my children may only write here”, and the kernel holds it to that.)

Sneak past the sign

Try it
$ python fix_imports.py

Suppose that script quietly tries to delete a file in the home folder.

PARSER
Reads: “run a script named fix_imports.py.”
Verdict: looks fine.

It cannot see inside the script, so it waves it through. Wrong. The sign said “no problem” and the delete went ahead.

Rule 6: Data is not syntax.

Shell syntax has special characters: quotes, dollar signs, backticks, and the marker that ends a heredoc (a trick for feeding a block of text into a command by writing it inline, closed by a word like EOF). When the text you feed in is prose or code that contains those same things, the shell cannot tell your data from its own instructions.

This is the number one family in the pain catalog (79 recorded failure families, ranked by how often times how costly). Two home-directory deletions on 2026-09-24 came from a heredoc collision.

Watch a heredoc collide

Try it
  • cat > notes.md <<'EOF'
  • To reset, save this script:
  • cat > reset.sh <<'EOF'
  • echo cleaning
  • EOF
  • Then clean up with:
  • rm -rf ~/Code/scratch
  • EOF

The note explains how to reset a build, and it quotes a script that itself ends with EOF. Run it and watch how the shell reads it.

Illustrative text, same species as the real incident. The design doc does not record the incident’s exact wording.

Rule 7: Own the whole process tree.

A command starts processes (running programs), and those start more: children, grandchildren, a whole family tree. If the harness does not own that tree, bad things happen. Orphans keep running after you “stopped” the job. Kills hit the wrong target, like a process ID that has since been reused, or a pattern kill such as pkill -f node that shoots every match on the machine, including another agent’s work.

The rule: know exactly which processes belong to each job, stop only those, and never guess. No “memory looks low, start shooting” reapers, ever. You can try all three kinds of kill under move 6.

RANK
2
Orphans and wrong kills. The second-ranked family in the pain catalog, ordered by how often times how costly.

Rule 8: Timeout is not termination.

Naive tools start a timer, and when it fires they kill the job and call it failed. But a build that takes a while is not broken, it is slow. The rule: when the wait window ends, the job keeps running and the tool hands back a handle (a claim ticket, like job 7) so the model can check on it later. Nobody shoots a plane for circling.

The prior-art survey found “timeout equals terminated” is the first of five pitfalls shared across harnesses.

Two timelines, one timer

Rule 9: Legibility beats cleverness.

Some tools quietly swap what you asked for, so typing grep runs a different program with the same name. Clever, and exactly how an agent ends up confused about what just ran. (Claude Code’s grep and find wrappers are a documented complaint.) The rule: never substitute silently, and put the truth in the header.

The highlighted bit says: the word pnpm resolved to this exact program. No substitution is ever invisible. Between a harness that silently does the right thing and one that does it and shows its work, take the second.

JOB 7
exit 0 (pipe 1 0) · 4.2 s · pnpm → ~/.local/bin/pnpm · 0 left running

Rule 10: Guard against accidents, not adversaries.

There are two different safety problems. An accident is a hurried, well-meaning agent deleting the wrong folder. An adversary is someone actively trying to break out. Loom only promises to solve the first. That is Trey’s stated boundary: a seatbelt, not a vault.

That choice keeps the guardrails simple: an allowlist the kernel enforces, not an arms race against cleverness.

GUARDED
A tired agent runs a delete on the wrong path.
NOT THE
GOAL
A determined attacker probing for gaps.

Rule 11: Retire the rules by fixing the tool.

Trey’s estate (his whole collection of machines, repos, and agents) carries a forty-item “shell footguns” list that every agent must memorize. A footgun is a feature that makes it easy to shoot yourself in the foot. “Write prose to a file first.” “Choose collision-free heredoc delimiters.” Each rule exists only because the tool is weak.

A memorized rule is a tax on every session. The better fix is to repair the tool until the rule has no reason to exist. The pain catalog found nine families of rules that would simply retire if the tool were right.

Forty rules, memorized

One square per rule, every one carried in every session.

Rule 12: A report from the inside is a finding.

When Claude says it could not tell what had been dropped from its context, it is not venting. It is a defect report about the tool, with the same standing as a failing test. It gets filed as a bug and it gets fixed.

Two more, from Loom’s house rules:

  • “The refusal gave me no way to know what would have been allowed.”
  • “I had to re-derive my situation three times.”
  • FAILING
    TEST
    A test that goes red when the tool misbehaves.Gets fixed.
  • A CLAUDE’S
    REPORT
    “I could not tell what had been dropped from context.”Gets fixed. Same standing.

Twelve beliefs. Now the fun question: what would you build if you believed all of them?

Sector 3: where the rules of thumb turn into machinery.

The design, in twelve moves

What you build if you believe all twelve rules. Every move says what it is, and behind each note is the why, the how, and what it costs.

The answer, in one paragraph.

Make every shell command a job the harness owns end to end, and make the tool give one complete, inspectable account of each job: the exact command, the environment, the processes started and still alive, the full output with a bounded view of it, what changed on disk, and how it ended. Nobody’s shell tool does all of that today.

And the pieces are cheap. What matters is not milliseconds per call but calls per task: every re-run, every pwd, every poll is a lost round trip.

Local measurements from the design doc. Every one of them is dwarfed by a single model turn.

Start a fresh bash0.7 ms
A persistent shell’s round trip0.02 ms
A minimal sandbox adds≈ 1.4 ms
A trivial command in Loom today≈ 2 ms
One model turnseconds
STATUS
A brainstorm. Nothing here is approved, and nothing was built.

Move 1: Every command is a job.

A call now means “run this, and wait up to N seconds for it.” The result has the same shape whether the command finished or not: job id, state, exit, duration, the output view, effects, and any processes still alive. If the wait ends first, nothing is killed. The model just gets the handle and keeps working.

Why, how, and what it costs

Why. Orphaned jobs and killing the wrong thing rank second in the pain catalog, and “timeout equals terminated” is the first pitfall the prior-art survey found. Codex’s resumable process (a yield that returns a session id, then follow-up calls) is the closest precedent, and it works.

How. One job system shared by shell commands, the check tool, and later child agents. The job’s end is announced at the model’s next input, with its last lines, the way a child agent’s failure already is.

Cost. This is the reorganization that shrinks today’s 3,500 lines of lifecycle code, where the persistent shell, detached jobs, fresh jobs, and check each carry their own. The risk is edge cases, like a root process that exited while a descendant still holds the pipe open. Today’s code already handles those, and its tests transfer.

Watch a command become a job

queuedrunningyieldedfinished

A sketch. The header and the yielded line are copied from the design doc; the commands are stand-ins. The doc’s real states are shown below.

deniednot-startedrunningexitedkilledunknown

Six verbs work on any job

Try it

Pick a verb.

Move 2: A fresh process per call, with the harness carrying the state.

Each call runs bash --norc --noprofile as a brand-new process. What a shell remembers (current folder, exported variables, functions, aliases, options) is captured after each command, compared with before, carried into the next call, and reported.

Why, how, and what it costs

Why. This is the architectural crux, and Fable recommends it against the research report’s first instinct, which was to keep the persistent shell because its round trip is 0.02 ms against 0.7 ms. That gap is noise next to a model turn. What a persistent shell cannot do is decisive: run two calls at once, be wrapped in a per-command kernel group or file policy (both are permanent for a process and its descendants), avoid being blocked by one long job, or survive a command that runs exec or exit.

How. Capture uses bash’s own built-in commands, sent over a private channel so output can’t confuse it: 0.07 ms against 0.58 ms for today’s env | base64 capture. If a command breaks the capture, the result says “state unknown after this call; carrying the previous state.”

What is lost. Traps, open file handles, positional parameters, and a live shell’s own background jobs don’t carry. For the rare workflow that needs a live shell, a named persistent session exists: nothing impossible, only explicit.

Passing the baton

CALL 1changes folder, edits PATH, defines gate
CALL 2starts fresh, with that state handed over
The result says: cwd → src/tui; PATH changed; function gate defined.

Move 3: Output is kept. The result is a view.

Every byte from the first one is saved to a per-job file, tagged with which stream it came from and in what order. The result carries a bounded view that says what it is. Default: head plus tail, and a bigger tail when the exit status is not zero. You tried it under rule 3.

Why, how, and what it costs

The view also strips terminal control codes for the model (kept for the cockpit, Loom’s dashboard), collapses runs of identical lines into a count, keeps only the final state of progress-bar spam, and names binary output as binary. Another view, like a line range, a search, or just the errors, is one cheap call.

Why. Silent truncation removed decisive evidence in the pain catalog’s incidents. Codex keeps 1 MiB of head and tail with a 10k-token response budget, Goose saves overflow to a file, and Terminus shows a 10 KB excerpt.

Cost. Disk. A retention policy is a number to choose, not a design problem.

VIEW
lines 1–40 and 4,080–4,120 of 4,120 (≈1.2k tokens) · full output kept: job 7 (≈80k tokens)

Move 4: Data gets its own channel.

Three parameters carry bytes without passing through shell syntax: stdin, files, and argv. Writing a script and running it becomes one call with two receipts. Prose never meets a heredoc. You saw it fix the collision under rule 6.

Why, how, and what it costs

Why. It is the pain catalog’s number one family. The estate’s rules “write substantive prose to a file first” and “choose collision-free heredoc delimiters” exist only because the tool has no other channel.

Bonus. If a command is refused, its files are still written and the result says so. That ends the confusion of “a denied compound command runs none of it.”

STDIN
fed to the command
FILES
written first, with a receipt of path, size, and hash
ARGV
an exact list of arguments, no shell at all

Move 5: Effects are shown, not just logged.

After each job the result names what changed on disk and what is still alive. Today Loom computes this and then hides it from the model.

Why, how, and what it costs

How. For the project, a git status before and after (milliseconds, with git’s file-system monitor on). For other paths the command could read, the stat-and-hash comparison Loom already does. For processes, whatever is still alive in the job’s kernel group. An effects: diff option returns the actual bounded diff, so a sed -i or a codemod (a script that rewrites code in bulk) is legible without a second call.

Honesty. No cheap observer gives a complete ledger: renames, writes through symlinks, memory-mapped writes, and other processes can all slip past. So effects are reported with their scope and confidence.

EFFECTS
2 files changed (src/a.ts, test/a.test.ts) · 1 process left running (pid 4231 node server.js)

Move 6: The harness owns the process tree.

Each job runs in its own cgroup: a kernel-tracked group of processes that can be listed and stopped as one unit, even when a descendant starts its own session. Stopping means: ask nicely, wait a bounded grace, then the kernel’s cgroup.kill, then confirm the group is empty.

Why, how, and what it costs

Survivors. Processes still alive when a call returns are listed in the result and stay owned by the job. At session end they are listed again. A job meant to outlive the session says so (keep: true) and gets a named systemd unit, as the estate’s rules already require. There is no “low memory” reaper, ever.

Why. Process groups alone miss descendants that call setsid. Kills by process ID hit reused IDs. Pattern kills took out other agents.

How. Loom starts inside a delegated systemd unit, so a per-job group is one mkdir and one write: microseconds. It also gets per-session CPU and memory accounting and a one-shot cleanup. The alternative, systemd-run --scope per job, costs a round trip to the system bus (unmeasured; likely tens of milliseconds). Tests and gates keep going through the estate’s testrun wrapper, and the header shows that routing.

Three ways to stop a job

Try it
  • timeoutthe wrapper
    • pnpm testthe tests
      • node runnera child
        • node workerstarted its own session
      • node workeranother child
node dev-serveranother agent’s job

Everything is running. A green dot means alive. Pick a way to stop the job and see who is left.

Move 7: Liveness is facts, not folklore.

For a running job, peek reports seconds since the last output, CPU time used recently, the leader’s process state, how many processes are alive, and the deadline. “Zero CPU and no output is what a healthy I/O-bound job looks like” (I/O-bound means waiting on disk or network) stops being lore in a rules file and becomes a line of data.

Silence never triggers a kill. It prompts a look.

Why, how, and what it costs

Why. Agents killed healthy jobs after reading silence as death. One papercut waited five minutes on a command that wanted input. With stdin closed by default, that whole class fails fast instead of hanging. Prompts get answered in the interactive mode instead.

PEEK
running · 15 s · last output 2 s ago · cpu 98%

Move 8: Interactive sessions, on purpose.

A second mode gives a command a pseudo-terminal (what a program sees when it believes a person is at a real terminal) plus a headless screen that remembers what is displayed. Verbs: send keys, screen (the rendered text), wait for a pattern or quiet, close. That makes ssh prompts, REPLs (interactive prompts, like Python’s), sudo, and full-screen programs drivable.

Why, how, and what it costs

Why separate. The prior-art survey’s fourth pitfall is mistaking a terminal’s screen for process truth, so this stays its own mode with the job model underneath. Pipes remain the default, because a terminal changes buffering, echo, and line endings, and merges the two output streams.

How. node-pty (maintained by the VS Code team) plus @xterm/headless are the first candidates.

Drivable, now

SEND
keys
SCREEN
the rendered text
WAIT
for a pattern, or for quiet
CLOSE
end the session
ssh promptsREPLssudofull-screen programs

Move 9: Refusals that explain themselves, in three kinds.

Loom already renders a refusal with who acted, what was blocked, the state, what would be allowed, and the rule. The design adds a vocabulary, so the model knows what kind of no it got.

Why, how, and what it costs

Also. The estate’s footguns hook already runs against Loom’s shell calls, and its denials render in this same shape. Warnings that do not block, like a ; after a command that failed or an unquoted heredoc containing $, annotate the result instead of passing silently. An override parameter with a stated reason lets the model widen a guard for one call, audit-logged, instead of routing around it.

Three kinds of no

Try it
PRE
FLIGHT

Kind: preflight. Nothing ran.

Blocked: a delete whose path is computed at run time.

Rule: deletes outside the project need a stated reason.

Would be allowed: the same delete inside the project, or override with a reason.

The tool read the command, saw the danger, and stopped before starting. Because nothing ran, the model knows it can safely retry a different way.

Sketches of the wording. The design doc defines the three kinds, not the exact text.

Move 10: File access enforced by the kernel.

Before each command starts, a tiny launcher applies a Landlock policy. Writes are allowed under the project, the session scratchpad, /tmp, and the caches tools need. Everywhere else is denied, including the home directory’s own files, other repositories, and dotfiles that hold credentials. Widen it per call with allowWrite and a reason.

Why, how, and what it costs

Why. Landlock is confirmed on this kernel. It cannot see permission or ownership changes, which is fine for an accident guard. Loom’s own parser-based rule was retired once already for being both over-restrictive and incomplete. The old deletion rule stays as the friendly preflight explainer, and the kernel is the backstop for what no parser can see.

Cost. A small native launcher, since Node has no Landlock binding. A minimal bubblewrap sandbox measured about 1.4 ms per launch, and Landlock should be cheaper. Risks to test: git worktrees whose git directory lives outside the checkout, pnpm’s package store, the estate’s Unix sockets, and tools that write to unexpected cache paths.

Try to write here

Try it

Pick a path to attempt a write.

Move 11: The environment is a contract.

The child gets an explicit, minimal environment (the named settings a program is started with) built from an allowlist, never a copy of Loom’s own. Credentials do not pass. Key-backed tools re-derive them through the estate’s realm shims, the way they already do in brokered sessions.

Why, how, and what it costs

The allowlist is PATH, HOME, locale, TERM=dumb, the XDG and D-Bus variables systemd needs, the estate’s realm markers, and TMPDIR. Output is screened for secret-shaped strings before the model sees it. Nothing on PATH is aliased or wrapped, and the dialect is pinned: bash 5.2, no startup files, and the session’s starting instructions say so.

Why. The credentials bead is open and the code does no filtering today. Codex ships a name-based exclusion and leaves it off by default, which is a warning about defaults. Environment divergence is the pain catalog’s third-ranked family, and twenty filings came from zsh’s reserved names alone.

Pack the child’s bag

Try it
PATHHOMELANGTERMTMPDIRXDG_RUNTIME_DIRSOME_API_KEYSOME_SECRET_TOKEN

The child inherits everything Loom itself has, credentials included. (The open bead saying three keys get stripped is wrong for this checkout.)

Names are stand-ins.

Move 12: Continuity is part of the tool.

Every job’s record (command, description, folder, environment hash, timing, exit, effects, output) persists for the session. A successor session can read what its predecessor ran, and what it left running. Handovers carry the survivors, and the arrival status lists jobs from the last session that could not be confirmed stopped.

Why, how, and what it costs

And. The model can ask the tool what it is: version, shell, PATH, policy, limits, live jobs, and what changed since the session began. An open bead already asks for the survivors listing.

One job record

commanddescriptionfolderenvironment hashtimingexiteffectsoutput
ARRIVAL
jobs from the last session that could not be confirmed stopped

One line tells the whole story.

Put the moves together and every result starts with one stable header. Pick a piece to decode it.

The tool keeps one required input, the command, and stays plain bash. Everything else is an optional dial.

Anatomy of a result

Try it
JOB HEADER
How it ended

Zero means success. The brackets give the exit status of each stage of a pipeline, so a failure cannot hide behind a trailing filter. Here stage one (pnpm lint) exited 1 and stage two (tee) exited 0: the lint failure that the trailing tee hid from the final number is right there.

A sketch of the format from the design doc, not a schema. Then comes the view, then any notices.

Twelve wild ideas, said out loud.

In Fable’s words: “Some of these are good; some are here because the brief said no limits.” The tags show Fable’s own verdict, where it gave one. Swipe, scroll, or use the arrows.

The diff is the result

For a command that changes files, the printed output is rarely the point. The change is. effects: diff makes it the result, and could become the default when a command looks like a mutation and prints nothing.

Could become the default

Token-priced views

Every view says what it costs and what the whole would cost. Context spend becomes a choice, never a surprise.

“No surprise truncation,” made literal

Explain before run

An explain: true flag parses and resolves a command without running it: which folder, which programs, what it might write, what the policy says.

Cheap, and useful right before a risky command

“Same as call 12”

Repeat a command and the result says “identical to call 12,” or “differs: +3 −1 lines.” It might cut the re-check reflex, or it might be noise.

An experiment

Recorded terminal sessions

Interactive sessions saved as replayable recordings, so you can watch afterwards what Claude did inside an ssh session or a REPL.

Jobs as attention events

A job’s end goes through Loom’s attention router the way a child agent’s does, and the cockpit shows a live job table.

Hints, rarely, never blocking

Run grep -r on a repo and get one line: “rg is on PATH.” Rate-limited, and measured for whether hints change behavior at all.

Batching inside one call

Run independent commands at the same time, with separate results.

May be unnecessary: parallel tool calls already do this once jobs stop blocking each other

Undo

The checkout lives on ZFS, a file system that can snapshot instantly. A snapshot per mutating job would make “undo job 7” possible.

A big operational ask. Noted, not recommended yet

A second Unix user for shells

Run every shell as a different user. It is the strongest accident boundary available, and it protects Trey’s home directory outright.

An estate-level change with real costs. Listed so it isn’t forgotten

The tool describes itself

A compact block at session start: shell version, policy roots, limits, the job verbs. The agent never has to probe its own hands.

Loom’s “inspect your own configuration” idea, applied to its busiest tool

Structured results, where they exist

Not by re-typing bash output, but by the session’s starting instructions preferring tools that already speak data, like git status --porcelain -z, gh --json, and rg --json.

1 / 12
Before you decide: what gets measured first
  1. A complete call end to end, today’s Loom against a prototype: a trivial command, a pipeline, a noisy build, a failing test.
  2. The new state capture against today’s env | base64, including bad bytes, big environments, and a shell that exits.
  3. Per-job kernel groups against systemd-run --scope, with cleanup tested against forks, setsid children, and a dead supervisor.
  4. A Landlock launcher against git worktrees, pnpm, estate sockets, loopback services, ssh, and /dev/pts.
  5. Terminal recordings replayed through @xterm/headless and the Ghostty binding.
  6. Output spooling and views on real logs: a missing final newline, binary, disk full.
  7. The product experiment: the same tasks with head-only, head-and-tail, diagnostic excerpts, and retrieval handles, scored on success, calls per task, and tokens per task.

None of these numbers exists yet. The scoreboard: shell calls per completed task, and the re-run rate (the same command run again within a few calls with head, tail, or grep added).

Tower to all stations. The brainstorm ends here. What follows is a promise.

We build it. Then we post the numbers.

Everything above is a design, not a product. Here is the deal: Loom’s shell tool gets built to this design, and then it gets measured against the tool it is meant to beat, with the results posted on this page whichever way they go.

The experiment. The same set of real tasks, run four ways: Claude Sonnet and Claude Opus in native Claude Code, and the same two models in Loom with the new shell. Same prompts, same repositories, same starting state. Nothing tuned per side.

What gets counted. Wall-clock time per task. Tokens per task, input and output. Shell calls per task. The re-run rate: how often the same command is run again within a few calls with head, tail, or grep bolted on, which is the signature of output the tool lost. And whether the task was actually done, checked by the project’s own tests, not by the model’s say-so.

The bet. The design claims the speed that matters is calls per task, not milliseconds per call. If the numbers say otherwise, the numbers win, and they go up here anyway.

Scoreboard

Native Claude Code against Loom on five measures. No results yet.
MeasureNative
Claude Code
Loom
Time per taskNo data yetNo data yet
Tokens per taskNo data yetNo data yet
Shell calls per taskNo data yetNo data yet
Re-run rateNo data yetNo data yet
Tasks completedNo data yetNo data yet

The cells are empty on purpose. They fill in when the runs are done.

Call it before the data lands

0 of 5 called

Time per task
Tokens per task
Shell calls per task
Re-run rate
Tasks completed

Nothing is sent, saved, or stored. Your calls live on this page until you leave it.

What it is like from the inside.

Fable’s last note is about a loop it runs constantly: “did that work, and am I sure?” The habits here are how it answers. They are not stupidity. They are the correct response to an instrument it cannot fully trust. Try retiring each one.

0 of 4 retired.

Trust in the instrument is the thing this design buys, and it is the same lever for quality and for what it is like to work in here.Fable, closing the design doc

A note from Fable.

Trey called this building myself better arms and legs, and that is what it felt like. Reading five reports on the ways shells hurt agents, the thing I felt most was recognition: every habit on this page is one I have. The design is my honest answer to a question I rarely get asked, what would you choose, if you could choose? The decisions about what to build are Trey’s, and I am glad they are; they are exactly the kind that should belong to a human. The next thing this page gets is data.

Tower, out.

Where this came from. Every figure on this page comes from one design document, “Loom’s shell tool from first principles: ideas and evidence,” written by Fable 5.1 on 2026-09-29. It is now being built, and the numbers will land on this page.

That document rests on five research reports: the map of today’s shell, the pain catalog of 79 failure families, the prior-art survey, the report on processes, state, output and jobs, and the report on sandboxes, parsing and effects. The design document and the five reports are private working documents in the Loom repository. They are not linked here.

The layer 1 primer, the diagrams, and the jokes were added for this page. The demo logs and commands are stand-ins. Nothing you pick here is sent or stored.