fix(scripts): self-heal isolated workspace provisioning when the base CLI is broken (#10574)

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents run tasks in isolated execution workspaces that are
provisioned as git worktrees by `scripts/provision-worktree.sh`
> - The script runs the base workspace's CLI (`cli/src/index.ts` via the
base `tsx` install) to seed each new worktree, and it only checked that
those files exist
> - pnpm links each package's `node_modules` into a hash-versioned
virtual store; a lockfile change followed by a partial or filtered
install prunes old hashed dirs without relinking every package, leaving
dangling symlinks
> - A CLI with dangling symlinks fails ESM resolution
(`ERR_MODULE_NOT_FOUND`) at boot, so provisioning aborts with
`setup_failed` — deterministically, on every retry, with no self-heal
path
> - This pull request makes provisioning health-check the CLI by
actually booting it, repair the base install when the check fails, and
degrade to the no-CLI fallback config instead of failing the run
> - The benefit is that a class of permanent `setup_failed` loops
becomes self-healing, and workspace provisioning survives a broken base
CLI

## Linked Issues or Issue Description

No public GitHub issue exists; the underlying bug is described here per
`bug_report.yml`. Related PR: #10578 self-heals the sibling
workspace-validation failure loop uncovered by the same incident
diagnosis.

**What happened?**
Isolated-workspace runs failed at provision time with `setup_failed`.
Every retry failed identically. One observed incident burned 4 runs
across two adapters before the task was stranded.

**Expected behavior**
Provisioning either succeeds or degrades gracefully; a broken base CLI
install repairs itself instead of permanently blocking all new
worktrees.

**Steps to reproduce**
In the base workspace, cause a lockfile-affecting dependency bump plus a
partial/filtered `pnpm install` so a package symlink (e.g.
`cli/node_modules/drizzle-orm`) dangles into a pruned virtual-store dir.
Start any isolated-workspace run. Provision fails with
`ERR_MODULE_NOT_FOUND` and the run ends `setup_failed`; retries never
recover.

**Paperclip version or commit**
master as of the branch point of this PR.

**Deployment mode**
Local trusted deployment with git-worktree isolated workspaces.

## What Changed

- `base_cli_healthy` now boots the base CLI (`--help`) instead of only
testing file existence, which exercises the top-level import graph.
- New `repair_base_workspace_install`: when the health check fails, run
a non-interactive `pnpm install --prod=false --force --frozen-lockfile`
in the base workspace. `--force` guarantees relinking when pnpm's
up-to-date heuristics would skip dangling symlinks; `--frozen-lockfile`
keeps the repair from mutating the shared lockfile.
- The repair install is serialized with `flock` on a lock file inside
the resolved git dir (`git rev-parse --absolute-git-dir`), so locking
also covers base workspaces that are linked worktrees, where `.git` is a
file.
- If every CLI candidate is unusable (including a base CLI the repair
could not fix), provisioning falls back to the existing no-CLI fallback
config writer (loudly, on stderr) instead of failing the run. A CLI that
runs and fails `worktree init` still fails provisioning with its real
exit code — that deliberate fail-closed policy is unchanged and covered
by an existing server regression test.
- Fixed a latent bug: `run_isolated_worktree_init` returned 0
unconditionally after the init subshell, so callers treated a failed
init as success. Exit codes now propagate.

## Verification

- Reproduced the incident state (dangling `cli/node_modules/drizzle-orm`
symlink); the base CLI failed with the exact `ERR_MODULE_NOT_FOUND` seen
in the incident run logs.
- Ran the patched script against a fresh scratch worktree: health check
failed → locked repair install ran (~26 s warm) → symlink relinked →
`worktree init` completed → exit 0 with `.paperclip/config.json` and
`.env` written.
- Happy path (healthy base CLI): provisioning behavior unchanged, exit
0.
- Verified `git rev-parse --absolute-git-dir` resolves a real directory
for both a normal checkout and a linked worktree.
- New hermetic tests: `node --test
./scripts/__tests__/provision-worktree-self-heal.test.mjs` (4 tests:
healthy CLI used, broken CLI degrades, locked repair end-to-end with a
fake pnpm, init failure propagates). Not yet wired into a CI workflow.
- `server`: the existing `realizeExecutionWorkspace` fail-closed
regression test ("fails instead of writing an unseeded fallback config
when worktree init errors after CLI detection succeeds") passes against
the new script.
- `bash -n scripts/provision-worktree.sh` is clean.

## Risks

- Low risk overall: the script only adds recovery paths; the happy path
is unchanged.
- The repair install runs in the shared base workspace. It is bounded by
`--frozen-lockfile` (no lockfile mutation) and serialized by `flock`,
but it can add ~30 s to the first provision after a base install breaks.
- If the repair cannot fix the CLI and no other CLI candidate exists,
runs now continue with an unseeded fallback config instead of failing;
that is intentional, and the fallback path already existed. Genuine
`worktree init` failures from a working CLI still fail the run.

## Model Used

Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended
thinking, agentic tool use (Claude Code harness).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
authored and GitHub committed 2026-07-31 16:18:52 -07:00
1 parent 3176e1f4b9
commit 79eff0aea1
2 files changed
+292 -18

No files matched your search

+88 -18
View File
@@ -31,32 +31,89 @@ source_env_path="$(dirname "$source_config_path")/.env"
mkdir -p "$paperclip_dir"
run_isolated_worktree_init() {
local base_cli_runner="$base_cwd/cli/node_modules/tsx/dist/cli.mjs"
local base_cli_entry="$base_cwd/cli/src/index.ts"
base_cli_runner_path="$base_cwd/cli/node_modules/tsx/dist/cli.mjs"
base_cli_entry_path="$base_cwd/cli/src/index.ts"
if [[ -f "$base_cli_runner" && -f "$base_cli_entry" ]]; then
base_cli_files_present() {
[[ -f "$base_cli_runner_path" && -f "$base_cli_entry_path" ]]
}
# File existence is not enough: pnpm links package node_modules into the
# versioned virtual store, so a lockfile change plus a partial/filtered install
# in the base workspace leaves dangling symlinks that fail ESM resolution at
# runtime. Actually boot the CLI to prove its import graph resolves.
base_cli_healthy() {
base_cli_files_present || return 1
(cd "$base_cwd" && node "$base_cli_runner_path" "$base_cli_entry_path" --help >/dev/null 2>&1)
}
repair_base_workspace_install() {
command -v pnpm >/dev/null 2>&1 || return 1
[[ -f "$base_cwd/package.json" && -f "$base_cwd/pnpm-lock.yaml" ]] || return 1
echo "Base workspace CLI at $base_cli_entry_path failed its health check (typically dangling pnpm symlinks after a partial install); repairing with pnpm install in $base_cwd." >&2
# --force guarantees relinking even when pnpm's up-to-date heuristics would
# otherwise skip the dangling symlinks; --frozen-lockfile keeps the repair
# from mutating the shared base workspace's lockfile.
local repair_cmd=(pnpm install --prod=false --force --frozen-lockfile --config.confirmModulesPurge=false)
# Resolve the real git dir so locking also covers base workspaces that are
# linked worktrees, where "$base_cwd/.git" is a file rather than a directory.
local repair_lock_dir=""
if command -v git >/dev/null 2>&1; then
repair_lock_dir="$(git -C "$base_cwd" rev-parse --absolute-git-dir 2>/dev/null || true)"
fi
if [[ ! -d "$repair_lock_dir" && -d "$base_cwd/.git" ]]; then
repair_lock_dir="$base_cwd/.git"
fi
if command -v flock >/dev/null 2>&1 && [[ -d "$repair_lock_dir" ]]; then
# The post-repair verification must run under the same lock: a concurrent
# provision's forced install could be mid-relink during an unlocked check
# and fail a repair that actually succeeded. Holding the lock also means a
# process that queued behind a peer's repair can skip its own reinstall.
(
cd "$worktree_cwd"
node "$base_cli_runner" "$base_cli_entry" worktree init --force --seed-mode minimal --name "$worktree_name" --from-config "$source_config_path"
cd "$base_cwd" || exit 1
exec 9>"$repair_lock_dir/paperclip-provision-repair.lock"
flock 9
if base_cli_healthy; then
echo "Base workspace CLI became healthy while waiting for the repair lock; skipping reinstall." >&2
exit 0
fi
env -u NODE_ENV CI=true "${repair_cmd[@]}" >&2 || exit 1
base_cli_healthy
)
return 0
else
(cd "$base_cwd" && env -u NODE_ENV CI=true "${repair_cmd[@]}" >&2 && base_cli_healthy)
fi
}
ensure_base_cli_healthy() {
base_cli_files_present || return 1
base_cli_healthy && return 0
repair_base_workspace_install
}
run_isolated_worktree_init() {
if ensure_base_cli_healthy; then
(
cd "$worktree_cwd" &&
node "$base_cli_runner_path" "$base_cli_entry_path" worktree init --force --seed-mode minimal --name "$worktree_name" --from-config "$source_config_path"
)
return
fi
if command -v pnpm >/dev/null 2>&1 && pnpm paperclipai --help >/dev/null 2>&1; then
(
cd "$worktree_cwd"
pnpm paperclipai worktree init --force --seed-mode minimal --name "$worktree_name" --from-config "$source_config_path"
cd "$worktree_cwd" &&
pnpm paperclipai worktree init --force --seed-mode minimal --name "$worktree_name" --from-config "$source_config_path"
)
return 0
return
fi
if command -v paperclipai >/dev/null 2>&1; then
(
cd "$worktree_cwd"
paperclipai worktree init --force --seed-mode minimal --name "$worktree_name" --from-config "$source_config_path"
cd "$worktree_cwd" &&
paperclipai worktree init --force --seed-mode minimal --name "$worktree_name" --from-config "$source_config_path"
)
return 0
return
fi
return 127
@@ -67,9 +124,7 @@ paperclipai_command_available() {
return 0
fi
local base_cli_tsx_path="$base_cwd/cli/node_modules/tsx/dist/cli.mjs"
local base_cli_entry_path="$base_cwd/cli/src/index.ts"
if command -v node >/dev/null 2>&1 && [[ -f "$base_cli_tsx_path" ]] && [[ -f "$base_cli_entry_path" ]]; then
if command -v node >/dev/null 2>&1 && base_cli_files_present; then
return 0
fi
@@ -429,9 +484,24 @@ else
echo "Existing isolated Paperclip worktree config is stale for this host; regenerating." >&2
fi
if paperclipai_command_available; then
run_isolated_worktree_init
if run_isolated_worktree_init; then
:
else
init_exit_code=$?
if [[ "$init_exit_code" -eq 127 ]]; then
# Every CLI candidate was unusable (e.g. an unhealthy base install that
# the repair could not fix); degrade instead of stranding the run.
echo "No usable paperclipai CLI found; writing isolated fallback config without DB seeding." >&2
write_fallback_worktree_config
else
# A CLI that ran and failed signals a real problem; do not paper over
# it with an unseeded fallback config.
echo "paperclipai worktree init failed (exit $init_exit_code); failing provisioning instead of writing an unseeded fallback config." >&2
exit "$init_exit_code"
fi
fi
else
echo "paperclipai CLI not available in this workspace; writing isolated fallback config without DB seeding." >&2
echo "paperclipai worktree init unavailable; writing isolated fallback config without DB seeding." >&2
write_fallback_worktree_config
fi
fi