mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
ae9711da4818357ec72b669ce5befa8104d52f6c
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ae9711da48 |
docs(release): re-date the 2026.818.0-beta.1 stable notes to v2026.824.0 (#12113)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Stable versions date the promotion, and the promotion reads its notes from master > - The merged notes for beta 2026.818.0-beta.1 assumed an Aug 21 promotion; the beta soaked longer > - This pull request re-dates the header to today's resolved version, v2026.824.0 > - The benefit is a GitHub Release whose title, date, and body agree ## Linked Issues or Issue Description **What existing behavior does this improve?** The stable notes header for the promotion happening today. **Current behavior** `releases/beta/v2026.818.0-beta.1.md` is titled `# Paperclip v2026.821.0`, `> Released: 2026-08-21`. **Proposed behavior** `# Paperclip v2026.824.0`, `> Released: 2026-08-24` — matching `./scripts/release.sh stable --date 2026-08-24 --print-version`. **Reason and benefit** The file publishes verbatim as the GitHub Release body; the header should match the version actually minted. ## What Changed - Three header/intro lines re-dated. Nothing else. ## Verification - `./scripts/release.sh stable --date 2026-08-24 --print-version` → `2026.824.0`. ## Risks - None; docs-only. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
83fefaadd1 |
fix(grok_local): do not warn when the default model sentinel is unavailable (#12062)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Each agent runs under an adapter. The `grok_local` adapter runs the Grok Build CLI. > - The adapter has an environment test. It probes the CLI and reports checks to the operator. > - `DEFAULT_GROK_LOCAL_MODEL` is `"grok-build"`. This value is a sentinel. It means "use the Grok CLI's own default model". > - `execute.ts` only passes `--model` when the configured model differs from the sentinel. So the sentinel is never sent to grok. > - The environment test still compared the sentinel to the models that `grok models` lists. Real grok never lists `grok-build`. > - So every probe emitted a false "Configured model not found" warning, even on a correctly configured agent. > - This pull request stops the false warning and keeps the real check for user-set models. > - The benefit is an accurate environment test: operators see a warning only when it is real. ## Linked Issues or Issue Description No public issue exists. The problem, in bug-report form: **What happened?** The `grok_local` environment test always warns `Configured model "grok-build" not found in available models`, even when the agent works. `grok-build` is the default sentinel, not a real model id, and it is never sent to the CLI. **Expected behavior** When the model is left at the default, the test reports the CLI's own default model as info and does not warn. It warns only when a user sets a real model that `grok models` does not list. **Steps to reproduce** 1. Create a `grok_local` agent and leave the model at its default. 2. Run the adapter environment test. 3. See the `grok_model_not_found` warning, although `grok models` and the hello probe succeed. **Agent adapter(s) involved** grok_local (Grok Build CLI). ## What Changed - `packages/adapters/grok-local/src/server/test.ts`: the model check now treats the default sentinel as valid and reports it as info (`Using the Grok CLI's default model (<default>)`). It still warns when an explicitly configured, non-sentinel model is absent from the discovered list. This matches `execute.ts`, which never sends the sentinel to grok. - `packages/adapters/grok-local/src/server/test.test.ts`: adds a test that the default sentinel does not warn when it is absent from the real model list, and a test that a real, unavailable model still warns. ## Verification - `pnpm exec vitest run packages/adapters/grok-local/src/server/test.test.ts` — 5 passed. ## Risks Low risk. The change only affects one adapter's environment-test reporting. It does not change how runs pass `--model`. No schema, no runtime behavior change. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (n/a — no doc change) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0dfa0fb988 |
ci: refresh general-server shard duration manifest (#12075)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its slowest check sets the feedback time for all contributors > - The general-server test lane splits its vitest suites across five runners with a duration-weighted partition (`scripts/general-server-shard.mjs`) > - The partition reads a duration manifest that was sampled on 2026-08-04, when the lane had 279 suites and 946s of serial time > - The lane has since grown to 405 suites and 1274s; 126 suites had no recorded duration and one suite grew from 37s to 123s > - The stale weights made the partition uneven: in the fully green actions run 32708351172, "General tests (server (2/5))" ran 364s and was the slowest check in the whole run, while sibling shards ran 292-330s > - This pull request refreshes the manifest with per-suite durations measured from that same run > - The benefit is a level five-shard split (255s ±1s of predicted suite time per shard), which removes ~50s from the slowest PR check ## Linked Issues or Issue Description **Describe the current behavior** In the fully green PR actions run [32708351172](https://github.com/paperclipai/paperclip/actions/runs/32708351172) (2026-08-24), the check "General tests (server (2/5))" completed in 364s. Its test step ran 315s while sibling shards ran 241-276s. It was the slowest check in the run. **Describe the improvement** The duration manifest `scripts/general-server-shard-durations.json` is stale. It holds 279 suites sampled on 2026-08-04, but the lane now has 405 suites. The 126 unknown suites fall back to the median weight (~1.3s), and `server/src/__tests__/workspace-runtime.test.ts` grew from 37.4s to 123.3s. The partition therefore predicts a level split but produces an uneven one. Refreshing the manifest restores the level split without any code change. **Expected impact** All five server shards level at ~255s of predicted suite time (~310s job time). The slowest PR check drops from 364s to about 317s, so the PR critical path improves by roughly 50s. ## What Changed - Regenerated `scripts/general-server-shard-durations.json` from actions run 32708351172 (2026-08-24): 405 suites, 1274s total serial time (was 279 suites, 946s from 2026-08-04) - Updated the `$comment` field to name the new sample run and date - No code changes; the partition logic in `scripts/general-server-shard.mjs` is untouched ## Verification - Parsed all five "General tests (server (n/5))" job logs from run 32708351172 with the consecutive-completion-timestamp method described in the manifest `$comment`; asserted that the parsed suite set equals the exact file list that `run-vitest-stable.mjs` collects (405/405, no misses, no extras) - Ran `node scripts/run-vitest-stable.mjs --mode general --group general-server --shard-index N --shard-count 5 --dry-run` for N=0..4 with the new manifest: each shard predicts 255s (±1s) of suite time, and the five shards form a complete, non-overlapping cover of all 405 suites - Ran `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs`: 13/13 pass ## Risks - Low risk. The change is data-only. Wrong weights cannot break correctness: the partition always covers every suite exactly once, so the worst case of a bad weight is an uneven shard, which is the current state. ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable (no code change; existing partition tests pass) - [x] I have updated relevant documentation to reflect my changes (manifest `$comment` updated) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #11528 (balanced the serialized server shards by recorded duration), #10923 (split serialized tests into five shards), #11156 (split workspaces-a into two shards). Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
87d68f476b |
fix: harden the sandbox bridge gateway against crashes and queue wedge (#12060)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents on remote sandbox targets reach the Paperclip API through the sandbox callback bridge: a loopback HTTP gateway inside the sandbox queues request files for a host-side worker > - The gateway process has no supervisor: nothing inside the sandbox respawns it, so a crash leaves a dead loopback port for the rest of the run > - The gateway also never cleaned up request files whose responses never arrived, so a stalled host wedged the queue at its depth cap and every later request got an immediate 503 > - #12052 made the host-side worker survive transient faults; this pull request hardens the other half of the relay > - The benefit is that a gateway fault degrades one request instead of severing the agent from the control plane until run end ## Linked Issues or Issue Description Refs #12052 (host-side worker half of the same relay). Refs #9904 and #8977 (adjacent bridge behavior). No public issue exists for this defect. The description below follows the bug report template. **What happened?** During a staging run, an agent's API calls to the bridge's loopback port began failing at the connection level (curl reported HTTP 000) partway through the run. A dead gateway process is the only mechanism that produces connection-level failures on that port, and nothing restarts it. Separately, request files for timed-out requests stayed in the queue; after 64 accumulated, the gateway answered every request with `503 Bridge request queue is full.` until the run ended. **Expected behavior** An uncaught fault in the gateway must not kill the loopback listener. A request that times out must not leave its file counting toward the queue-depth cap. A queue full of orphaned files must recover instead of rejecting until run end. **Steps to reproduce** 1. Start a remote-sandbox run and stop the host-side bridge worker. 2. Send requests to the gateway until they time out; the request files stay in `requests/`. 3. After 64 such files, every request gets an immediate 503, even after the host recovers. 4. Independently, raise any uncaught exception in the gateway process; the loopback port dies for the rest of the run. ## What Changed - The generated gateway source installs global `uncaughtException` / `unhandledRejection` handlers that log to stderr (already redirected to `logs/bridge.log`) and keep serving. The relay holds no state a fault can corrupt beyond the one request it interrupted. - Survival is gated on readiness: before the gateway has written its readiness file (file mode) or sent its READY frame (duplex mode), the same handlers exit(1) instead. A startup fault (failed bind, failed readiness write) means the process can never serve, and surviving there would only leave an un-ready zombie while the host waits out its readiness poll. - The file gateway attaches an explicit `error` listener to its server and pins the event loop with a keepalive until the bind settles. Newer Node runtimes do not reliably surface a failed bind through `uncaughtException` in this shape: the process can drain and exit 0 before the error event is delivered (reproduced on Node 24/25; Node 22 delivered it). The duplex gateway already had an explicit listener. - A request that times out waiting for the host now deletes its own request file. The host's response write is guarded on that file, so the removal also signals that no caller waits anymore. - At the queue-depth cap, the gateway sweeps request files older than the response deadline (orphans from killed callers or a previous gateway process) before rejecting with 503. - Host-side, `processRequestFile` treats a request file that vanished before the read as the benign caller-gave-up race and skips it quietly instead of escalating into the recovery pass. ## Verification - `npx vitest run packages/adapter-utils/src/sandbox-callback-bridge.test.ts` — 43 passed, verified on both Node 22 and Node 25. - New end-to-end test: with no worker running, a request times out (502), its file is cleaned, and the same gateway then serves a 200 once a worker starts — no wedge, no dead port. - New end-to-end test: with `maxQueueDepth: 1` and a backdated orphan file at the cap, the gateway sweeps the orphan and admits the request instead of answering 503. - New worker test: a request file that vanishes before the read is skipped without a handler call, a response write, or a run-level error. - New generated-source test: spawned directly against an already-occupied port, the gateway exits 1 promptly with the `EADDRINUSE` fault on stderr instead of lingering un-ready (or exiting 0 silently, the pre-existing behavior on Node 24/25). - A pin keeps the crash handlers, the readiness gate, and the sweep in the generated source. - `pnpm --filter @paperclipai/adapter-utils typecheck`. ## Risks - Keeping a Node process alive after `uncaughtException` is normally suspect; here the alternative is a dead loopback port for the rest of the run, and the gateway is a stateless per-request relay. The fault is logged with its stack to `bridge.log`, and survival applies only after readiness — startup faults still fail fast. - Deleting a timed-out request file could race a host that is mid-processing. The host's response write is already guarded on request-file existence, and the new host-side skip treats the vanished file as a no-op, so no duplicate mutation path is introduced. - The stale sweep runs only at the depth cap and only removes files older than the response deadline plus a 2 s grace, so a live caller's file is never swept. - Orphaned response files (host responded after the caller gave up) still linger; that pre-existing minor leak is unchanged here. ## Model Used - Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended thinking enabled, agentic tool use via Claude Code (CLI harness), 200k context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fc9e9b704f |
fix: stop teaching agents to curl literal {id} route templates (#12061)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Adapters inject prompt text that teaches agents how to call the
Paperclip API, including copy-pasteable curl examples
> - Some of those URLs contained brace placeholders like
`/api/issues/{id}/checkout`
> - Agents paste such lines verbatim; the placeholder reaches the server
as `/api/issues/%7Bid%7D` and 404s, and request logs show agents doing
exactly that
> - The acpx engine's API note already avoids this by using
`$PAPERCLIP_TASK_ID`, and its test pins `/api/issues/{id}` out of the
prompt
> - This pull request applies the same standard to the gemini adapter,
the shared prompt template, and the openclaw gateway workflow
> - The benefit is that agents stop burning turns on placeholder 404s
and doc examples stay safe to execute as written
## Linked Issues or Issue Description
No public issue exists for this defect. The description below follows
the bug report template.
**What happened?**
Server request logs show agents issuing `GET /api/issues/%7Bid%7D` — the
literal, percent-encoded text `{id}` — which 404s. The source is adapter
prompt text: the gemini adapter's API note embeds a curl example with
`/api/issues/{id}/checkout` in the URL, the shared agent prompt template
mentions `/api/issues/{issueId}` endpoints, and the harness checkout
notice names `/api/issues/{id}/checkout`. Models copy these strings into
real requests.
**Expected behavior**
URL paths in prompt text must carry environment variables or real ids,
never brace placeholders, in every string an agent might execute
verbatim. Where a placeholder is unavoidable, the prompt must state
explicitly that the literal text must never be sent.
**Steps to reproduce**
1. Give an agent the gemini adapter's API access note.
2. Watch it call `curl ...
"$PAPERCLIP_API_URL/api/issues/{id}/checkout"` as written.
3. The server logs `POST /api/issues/%7Bid%7D/checkout 404`.
## What Changed
- gemini-local's API note curl example now uses `$PAPERCLIP_TASK_ID` and
tells the agent to substitute a real issue id when that variable is
absent — the same convention as the acpx engine's API note.
- The shared agent prompt template (`server-utils.ts`) uses
`$PAPERCLIP_TASK_ID` in its interaction-creation and resume-endpoint
mentions, and the harness checkout notice names `POST
/api/issues/$PAPERCLIP_TASK_ID/checkout`.
- openclaw-gateway's endpoint workflow keeps its `{issueId}`
placeholders — they are defined by its "determine issueId" step — but
now states explicitly that the literal text must never be sent in a URL.
- `server-utils.test.ts` pins the new form and adds negative pins that
keep `/api/issues/{id}` and `/api/issues/{issueId}` out of the shared
prompt template, mirroring the existing acpx-engine negative pin.
- The `confirmation:{issueId}:plan:{revisionId}` idempotency-key
template is untouched: it is a value-construction pattern, not a URL.
## Verification
- `npx vitest run packages/adapter-utils/src/server-utils.test.ts
packages/adapters/gemini-local packages/adapters/openclaw-gateway` — 152
passed. The single failure (`pre-selects gemini-api-key auth in the
managed HOME for sandbox execution`) is a pre-existing
environment-specific failure on the development machine, unrelated to
prompt text; CI is authoritative for it.
- `pnpm --filter @paperclipai/adapter-utils --filter
@paperclipai/adapter-gemini-local --filter
@paperclipai/adapter-openclaw-gateway typecheck`.
## Risks
- Low risk: prompt-text and test changes only; no runtime logic changes.
- Agents that memorized the old example strings keep working — the
routes are unchanged, only the placeholder text in prompts is.
## Model Used
- Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended
thinking enabled, agentic tool use via Claude Code (CLI harness), 200k
context window.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
633e102971 |
fix: verify issue-update writes instead of inferring success (#12051)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents report task state to the control plane with `PATCH
/api/issues/{id}` at the end of each heartbeat
> - On remote sandbox targets those writes cross a relay that can fail
at the connection level
> - An agent that pipes its status curl through `head` cannot see that
failure; the write is lost but the run reports success
> - The issue then stays `in_progress` with no disposition, and the
missing-disposition recovery must repair it
> - This pull request makes the issue-update helper verify every write,
and it teaches the shared skill to require verified writes
> - The benefit is that a lost status write becomes a visible, retried
failure instead of a silent success
## Linked Issues or Issue Description
No public issue exists for this defect. The description below follows
the bug report template.
**What happened?**
A sandboxed heartbeat run answered its issue in a comment. It then sent
`PATCH /api/issues/{id}` with `status: done` through `curl -sf ... |
head -c 400`. The relay dropped the connection. The `-f` flag suppressed
the error output, and the pipe replaced curl's exit code with the exit
code of `head`. The agent saw empty output and exit 0. It reported the
write as an "empty 2xx" success and exited. The issue stayed
`in_progress`, and the successful-run recovery had to close it in a
corrective run.
**Expected behavior**
A status write that does not reach the server must surface as a failure.
The helper script must retry transient failures. It must exit non-zero
when the write is unconfirmed. Skill guidance must forbid write patterns
that hide failures.
**Steps to reproduce**
1. Point `PAPERCLIP_API_URL` at an endpoint that drops connections
intermittently.
2. Finalize an issue with `curl -sf -X PATCH
"$PAPERCLIP_API_URL/api/issues/$ID" -d '{"status":"done"}' | head -c
400`.
3. Observe exit code 0 with empty output while the server never received
the PATCH.
## What Changed
- `scripts/paperclip-issue-update.sh` now captures `%{http_code}`,
retries a retryable failure (connection-level, 429, 5xx) once — two
attempts total, which matches the shared bounded-write-retry rule —
rejects an empty 2xx body, and confirms the response echoes the
requested status before it exits 0.
- Failure output states plainly that the write was NOT saved, so the
calling agent reports it accurately.
- `skills/paperclip/SKILL.md` Step 8 adds a required "Verify writes —
never infer them" rule: a successful PATCH always returns the updated
issue JSON, disposition writes must never run through `head`/`tail`
pipelines, and an unconfirmed write must be reported as FAILED.
- `server/src/__tests__/paperclip-skill-utils.test.ts` pins the new
skill rule; a new `paperclip-issue-update-helper.test.ts` exercises the
helper's behavior end-to-end.
## Verification
- `bash -n scripts/paperclip-issue-update.sh`
- `server/src/__tests__/paperclip-issue-update-helper.test.ts` runs the
helper end-to-end against a local HTTP server: confirmed-echo success
(exit 0), empty 2xx (exit 1), wrong echoed status (exit 1), 422 reject
(exit 1, exactly one request), 503 then success (two requests),
connection refused (two attempts, then exit 1 with a "NOT saved"
report).
- `npx vitest run
server/src/__tests__/paperclip-issue-update-helper.test.ts
server/src/__tests__/paperclip-skill-utils.test.ts
server/src/__tests__/cli-invocation-safety.test.ts` — 50 passed.
## Risks
- Low risk. The success-path output is unchanged (the updated issue
JSON).
- The helper now exits non-zero on unconfirmed writes. Callers that
previously missed silent failures now see explicit errors. That is the
intended behavior change.
- The single retry re-sends the PATCH after a retryable failure. If the
first request committed and only its response was lost, an attached
comment can post twice. The duplicate is visible and benign; the prior
behavior lost the write silently.
## Model Used
- Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended
thinking enabled, agentic tool use via Claude Code (CLI harness), 200k
context window.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
c7f4bc1300 |
fix: survive transient sandbox exec failures in the callback bridge worker (#12052)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents on remote sandbox targets reach the Paperclip API through the sandbox callback bridge: a loopback gateway inside the sandbox writes request files, and a host-side worker polls them over the provider's exec channel and forwards them to the server > - The worker's poll loop had one terminal catch: a single reset or slow exec ended the relay for the rest of the run > - The in-sandbox gateway kept queueing requests against the dead worker, so every later API call from the agent stranded, including its final status write > - A relay that dies on one transient fault turns a routine provider hiccup into a lost issue disposition > - This pull request restructures the loop so transient faults back off and retry, while the watchdog remains the escalation path for sustained outages > - The benefit is that one flaky exec no longer severs an agent from the control plane mid-run ## Linked Issues or Issue Description Refs #9904. Refs #8977. Both touch adjacent bridge behavior (curl shim, header forwarding); neither addresses worker-loop lifetime. No public issue exists for this defect. The description below follows the bug report template. **What happened?** During a staging run, the host-side bridge worker hit one failed sandbox exec while relaying requests. The poll loop's only catch is terminal: it failed the pending requests and set the worker to settled, with no restart. The agent's later API calls saw connection-level failures or bridge errors until the run ended. Its final `PATCH status: done` was lost, and the missing-disposition recovery had to repair the issue in a corrective run. **Expected behavior** One transient exec failure must not end the relay for the rest of the run. The worker must back off and retry. A sustained outage must still fail queued requests fast through the watchdog. In-flight request semantics (abort plus 504 backstop, retry-safe 503) must not change. **Steps to reproduce** 1. Start a remote-sandbox run and let the provider exec channel reject or stall one call while the bridge worker polls. 2. The worker hits one `listJsonFiles` failure or one request-attempt timeout, and the loop exits through its terminal catch. 3. Every later bridge request strands. The loopback gateway keeps accepting requests that never complete, and after 64 queued files it answers every request with 503 until run end. ## What Changed - `startSandboxCallbackBridgeWorker`'s poll loop now separates three failure domains: - A failed poll backs off exponentially (capped at 5 s) and retries instead of dying. The first failure of a streak still lands on the run trace through the workerFailed span; later repeats only warn. - A failed or hung request attempt runs the same recovery pass the loop previously died on — abort the in-flight handler (its 504 backstop keeps the caller from stranding) and 503 the unclaimed queued requests — and the loop then continues and serves the caller's retry. - A listing where every file already has an in-flight attempt sleeps one poll interval, like an empty listing. The previous immediate re-list was a hot spin: an exec storm against a real channel, and a pure-microtask loop that starved every timer in the process when the queue client resolves synchronously. - The watchdog, the claim/finalize fences, and the stop/drain semantics are unchanged. ## Verification - `npx vitest run packages/adapter-utils/src/sandbox-callback-bridge.test.ts` — 38 passed. - New regression test: a request queued behind three consecutive poll failures is still delivered. - Updated tests: the stalled-poll test now expects recovery (the handler's real 200) instead of a terminal 503; the sustained-outage 503 path remains proven by the dedicated watchdog test; the recovery-503-write-retry test triggers the recovery pass through a hung request read, because a hung poll no longer runs that pass. - `pnpm --filter @paperclipai/adapter-utils typecheck`. ## Risks - Behavior change: a transiently failing poll no longer mass-fails queued requests on the first error. Callers wait through the backoff window, bounded by the existing in-sandbox 30 s response deadline, or the watchdog fails them after 20 s of no successful iteration. This trades fast-but-terminal degradation for recovery. - A hard-down channel now retries every ≤5 s for the rest of the run instead of stopping. Each retry is one exec attempt against a channel that already fails. - In-flight mutation safety is unchanged: the guard map and the claim protocol still prevent a double-applied host mutation, and a guarded file is always finalized by its own attempt or by its 504 backstop. ## Model Used - Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended thinking enabled, agentic tool use via Claude Code (CLI harness), 200k context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c62bb4b16b |
feat: environment delete with agent reassignment and consented sandbox destroy (#12053)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments define where agent runs execute: local, SSH, or provider sandboxes > - Operators can create and edit environments, but the UI has no way to delete one > - The server already exposes `DELETE /environments/:id` and a delete-blast-radius preflight, but no UI consumes them, and a delete blocked by reusable sandbox leases gives the operator no path forward > - This pull request adds the delete flow to the environment configuration page: a preflight-driven modal that reassigns dependent agents, names the workspaces that hold blocking sandbox leases, and can destroy those sandboxes with explicit consent > - The benefit is that operators can retire stale environments from the UI without database surgery, and dependent agents move to a chosen replacement instead of silently falling back ## Linked Issues or Issue Description Refs #8554 Refs #11124 **Subsystem affected** Environments (server routes, environment runtime service, and the environment settings UI). **Problem or motivation** The environment configuration page has no delete control. The server delete endpoint exists, but nothing in the UI calls it. When reusable sandbox leases block a delete, the 409 error names no owner, so the operator cannot find the blocking workspace. Agents that use the environment as their default lose it silently through the FK `on delete set null`. **Proposed solution** Add a delete button with a confirmation modal on the environment edit page. The modal reads the delete-blast-radius preflight. It offers a dropdown to reassign dependent agents to another environment before the delete. It lists each workspace that holds a blocking reusable sandbox lease, with a link. When those leases are the only blocker, the confirm button destroys the sandboxes inline (`?destroyReusableSandboxLeases=true`) and then deletes. A failed teardown falls back to `pending_cleanup` for the sweep, so no sandbox is orphaned. ## What Changed - `ui/src/pages/CompanyEnvironments.tsx`: delete button on the edit page header, confirmation modal with agent reassignment select, lease-holder list, impact notes, and a consent-labeled destroy-and-delete action - `ui/src/api/environments.ts`: `deleteBlastRadius` and `remove` client methods; `remove` takes an optional `destroyReusableSandboxLeases` flag - `server/src/routes/environments.ts`: `DELETE /environments/:id` accepts `?destroyReusableSandboxLeases=true`; it destroys the environment's reusable sandbox leases first, but only when those leases are the sole delete blocker, then re-checks the blast radius before it deletes - `server/src/services/environment-runtime.ts`: new `destroyReusableSandboxLeasesForEnvironment` — destroys every reusable sandbox lease an environment still owns while the environment config (provider credentials) is still available - `server/src/services/environments.ts`: the delete blast radius now returns `reusableSandboxLeaseHolders` (lease id, workspace, issue) so clients can name what blocks a delete - `packages/shared/src/types/environment.ts`: `EnvironmentDeleteReusableLeaseHolder` type on the blast radius - Tests: route gating for the consent flag (destroy runs, mixed-blocker rejection, surviving-lease rejection), runtime destroy scoped to an environment, blast-radius holder join, and UI tests for the reassignment flow, holder links, and the consent button ## Verification - `npx vitest run server/src/__tests__/environment-routes.test.ts server/src/__tests__/environment-service.test.ts server/src/__tests__/environment-runtime.test.ts ui/src/pages/CompanyEnvironments.test.tsx` - Manual: open Settings → Environments → edit an environment. The trash icon opens the modal. With agents on the environment, pick a reassignment target and confirm; agents move and the environment deletes. With reusable sandbox leases, the modal names the holding workspaces and the confirm button reads "Destroy N sandboxes and delete". ## Risks - The consented path destroys provider sandboxes. It runs only when reusable leases are the sole blocker, so a delete that would still be rejected never destroys anything. A failed teardown routes to `pending_cleanup` and the delete stays blocked until the sweep resolves it. - Agent reassignment issues one PATCH per agent from the client. A mid-sequence failure leaves some agents reassigned; the reassignments are valid on their own and the UI refreshes to the actual state. - Hard blockers (managed local, instance default, pending cleanup) keep the existing 409 behavior and disable the confirm button. ## Model Used - Claude (Anthropic) — Fable 5, model id `claude-fable-5`, extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
627eef7cbd |
fix(plugins): retry errored plugins at boot instead of leaving them dead (#12054)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Plugins extend the server with sandbox providers, tools, and jobs; a loader activates them at boot > - When activation fails, the loader marks the plugin `error` and skips it on every later boot > - Activation failures are often environmental — missing package dependencies, a stale build output, a module that moved under a pull — and the fix lands on disk without any write to the plugin row > - The plugin therefore stays dead forever, and every feature behind it (sandbox destroys, cleanup sweeps, probes) silently stops working until an operator flips the row by hand > - This pull request makes `loadAll` retry errored plugins once per boot: flip to `ready`, attempt activation, and re-record the error if the attempt fails > - The benefit is that a plugin recovers on the next boot after its environment is fixed, with no manual database or lifecycle intervention ## Linked Issues or Issue Description **What happened?** Several sandbox-provider plugins sat in `error` status for weeks after a transient activation failure (a module resolution error from an older checkout state). The boot loader only loads plugins in `ready` status, so it never retried them. Environments backed by those providers lost sandbox destroys, cleanup sweeps, and probes with no visible signal other than the stale `last_error`. **Expected behavior** A plugin whose activation failure has been fixed on disk recovers on the next server boot. A plugin that still fails stays in `error` with a fresh error message. **Steps to reproduce** 1. Install a plugin whose worker cannot start (for example, delete one of its dependencies), then boot the server. The plugin lands in `error` status. 2. Restore the dependency. 3. Restart the server. Before this change, the plugin stays in `error` forever. After this change, the boot retries it and the plugin activates. ## What Changed - `server/src/services/plugin-loader.ts`: `loadAll` also fetches plugins in `error` status, flips each to `ready`, and activates it with the normal batch. The flip runs before activation because the `error` status only legally transitions to `ready` or `uninstalled`; a retry that failed while still in `error` could not re-mark itself. A failed flip logs a warning and never aborts the boot load. The stale comment at the `markError` site now describes the retry. - `server/src/__tests__/plugin-loader-error-retry.test.ts`: covers the flip-then-retry flow, the failed-flip isolation, and the empty case. ## Verification - `npx vitest run server/src/__tests__/plugin-loader-error-retry.test.ts server/src/__tests__/bundled-plugins.test.ts server/src/__tests__/plugin-lifecycle-restart.test.ts server/src/__tests__/cloud-image-bundled-plugins.test.ts` - Manual: mark an installed plugin's status to `error`, restart the server, and observe the loader log line `retrying plugins that failed activation on a previous boot` followed by a successful activation (or a fresh `last_error` if the plugin is genuinely broken). ## Risks - A genuinely broken plugin now costs one bounded activation attempt per boot (the attempts run in parallel with the ready batch under `Promise.allSettled`). It cannot crash-loop within a running process, and it returns to `error` with a fresh message. - The flip clears `last_error` before the attempt. If the process dies between the flip and the activation, the row is `ready` with no error text; the next boot simply loads it as a ready plugin. - Operators who relied on `error` as a manual "keep this off" latch should use the `disabled` status, which this change does not touch. ## Model Used - Claude (Anthropic) — Fable 5, model id `claude-fable-5`, extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fbd20b28d3 |
fix(grok-local): stop defaulting --permission-mode to dontAsk (#11898)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The `grok_local` adapter runs the native Grok Build CLI in headless
mode for unattended agent heartbeats
> - Grok CLI 1.0 started to enforce the `dontAsk` permission mode as
deny-by-default, and it takes precedence over `--always-approve`
> - The adapter passes both flags on every run, so each run dies on its
first tool call and is still recorded as a success
> - This pull request removes the `dontAsk` default so unattended runs
rely on `--always-approve` alone
> - The benefit is that `grok_local` agents can execute tools again on
current Grok CLI releases
## Linked Issues or Issue Description
No public issue exists. Description per the bug template:
**What happened?**
Every `grok_local` run on Grok CLI 1.0.x stops on its first tool call.
The stream shows the tool call move from `pending` to `failed` with
"User cancelled the execution for tool `run_terminal_command`", and the
session ends with `stopReason: "cancelled"` after one turn. The CLI
exits 0, so Paperclip records the run as succeeded with no work done,
and the issue lands in missing-disposition recovery.
**Expected behavior**
Unattended runs must auto-approve tool executions. The adapter already
passes `--always-approve` for this.
**Steps to reproduce**
In a clean Linux environment with Grok CLI 1.0.3 and `XAI_API_KEY` set,
run the adapter's exact invocation shape:
`grok --output-format streaming-json --permission-mode dontAsk
--always-approve --disable-web-search --single "Run the shell command:
echo ok"`
The tool call is denied. Drop `--permission-mode dontAsk` (or use
`--permission-mode bypassPermissions`) and the same command executes the
tool. On Grok 0.2.x the original combination worked because the CLI
accepted `dontAsk` without enforcing it; the 0.2.39 embedded docs state
the flag takes effect only for `bypassPermissions` / always-approve.
**Paperclip version or commit**
master (
|
||
|
|
adfbe2d4b9 |
feat(environments): refer to the managed default environment by name, not the sandbox driver key (#11838)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed deployments provision a platform-managed default environment
for agent runs; the UI shows this environment in selectors, the agent
form, run details, and the environments page
> - Those surfaces append the raw driver key to the environment name, so
users see labels like "Paperclip Computer (sandbox)", "Paperclip
Computer · sandbox", and fallback copy such as "Managed sandbox" and
"The sandbox has no ready authentication"
> - "sandbox" is infrastructure vocabulary, not the product name of the
environment; showing it next to the managed environment's name is
confusing and off-brand
> - This pull request renders platform-managed environments by name
alone and rewords the sandbox-phrased copy, while user-created
environments keep the driver suffix so mixed lists stay distinguishable
> - The benefit is that the default environment reads as one clear
product name everywhere, and self-hosted users lose nothing: their own
environments still show the driver
## Linked Issues or Issue Description
**What existing behavior does this improve?**
Display of the platform-managed default environment across the UI.
**Subsystem affected**
UI (environment selectors, agent config form, environments page, agents
page, run details) and the claude-local/codex-local adapter auth checks.
**Current behavior**
The agent form labels the inherited default environment as "Name
(sandbox)". Environment selectors and the environments list render "Name
· sandbox". The agents page describes the environment as "<provider>
sandbox provider". The agent form's fallback label is "Managed sandbox".
Adapter auth checks say "The sandbox has no ready authentication for
this adapter."
**Proposed behavior**
Platform-managed environment rows (`metadata.managedByPaperclip`) render
their name alone. The fallback label is "Paperclip Computer". The agents
page describes managed environments as "Managed by Paperclip". Run
details omit the driver suffix for sandbox-driver environments (the
adjacent Provider entry already identifies the mechanism). Adapter auth
checks say "This environment has no ready authentication for this
adapter."
**Reason and benefit**
The managed environment carries a product name. Appending the raw driver
key ("sandbox") to it is noise and contradicts the product naming.
User-created environments keep the driver suffix, so mixed lists stay
distinguishable.
**Breaking changes**
None. Message text of the auth check is not read programmatically; the
UI keys off `ADAPTER_AUTH_MISSING_CHECK_CODE`. Rows without the managed
marker render exactly as before.
## What Changed
- New `environmentDisplayLabel` helper in
`ui/src/lib/managed-sandbox-environment.ts`: managed rows → name alone;
other rows → "Name · driver".
- `AgentConfigForm`: inherited-default label uses the helper; fallback
copy "Managed sandbox" → "Paperclip Computer"; environment options use
the helper.
- `ProjectProperties`, `CompanyEnvironments`: environment selector
options use the helper; the environments-list row hides the driver
suffix on managed rows; the managed detail page's fallback description
no longer says "sandbox".
- `Agents` page: managed environments are described as "Managed by
Paperclip" instead of "<provider> sandbox provider".
- `CommentThread` run details: the driver suffix is omitted for
sandbox-driver environments.
- claude-local and codex-local adapters: auth-missing check message/hint
reworded from "sandbox" to "environment" (ACP and environment-test
paths); claude-local probe/effort/login hints reworded the same way.
- Run status lines: "Syncing workspace to sandbox", "Exporting git
changes from sandbox", "Starting adapter in sandbox", and friends now
say "environment"; "Finalizing sandbox workspace" → "Finalizing
workspace". Templated transfer-progress lines map the `sandbox`
transport key to "environment" for display (`runtime-progress.ts`).
- Agent form sign-in panel: "Sign in to the sandbox" → "Sign in to the
environment"; "Authenticated. The sandbox has credentials now." → "…The
environment has credentials now."
- Feature catalog + instance settings card: "Managed Sandbox Only" →
"Managed Environment Only" (setting key unchanged; the card keeps its
alphabetical slot).
- Server agents routes: execution-target failure and test-identity copy
no longer say "sandbox"; workspace-mode label "Cloud sandbox" → "Cloud
environment".
- Tests: new `environmentDisplayLabel` unit cases; new `AgentConfigForm`
render case asserting the managed default renders without "(sandbox)" or
"· sandbox"; status-line assertions updated across adapter-utils, server
heartbeat/live-run, and UI chat suites.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `pnpm --filter @paperclipai/adapter-claude-local typecheck` and
`--filter @paperclipai/adapter-codex-local typecheck` — clean.
- `vitest run` for `managed-sandbox-environment.test.ts`,
`AgentConfigForm.render.test.tsx`, `CompanyEnvironments.test.tsx`,
`Agents.test.tsx`, `CommentThread.test.tsx`, `NewAgent.test.tsx` — all
green (118 tests across the two runs).
## Risks
Low risk. Cosmetic label changes only; no data or API changes. Rows
without `metadata.managedByPaperclip` render exactly as before, so
self-hosted deployments with their own environments see no change. The
only self-hosted-visible wording changes are the adapter auth-check
message and the driver suffix omission on sandbox-driver rows in run
details.
## Model Used
- Claude (Anthropic) — claude-fable-5 (Claude Fable 5), Claude Code CLI,
extended thinking, tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
docs reference these labels)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
599ad7016c |
ci(release): raise npm publish visibility budget to 10 minutes per package (#11835)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release workflow publishes every public package to npm and polls each version's registry visibility before continuing > - #11834 raised the poll budget from 60 seconds to 5 minutes after npm CDN propagation lag failed four canary runs > - The very next canary run missed by ten seconds: `adapter-opencode-local@2026.821.0-canary.2` was accepted at 07:02:27 UTC and became visible at 07:07:40, just past the final poll > - This pull request doubles the per-package budget to 10 minutes > - The benefit is a release train that survives the one consistently slow package; the 90-minute publish job timeout from #11834 already absorbs it ## Linked Issues or Issue Description Follow-up to #11834. Evidence in the `Release` run for `16149a75f`: every package's publish became visible within seconds except `adapter-opencode-local`, which has lagged 3-5+ minutes on all of today's runs and exceeded the 5-minute budget by ten seconds on the latest. ## What Changed - `NPM_PUBLISH_VERIFY_ATTEMPTS` 30 → 60 (with `NPM_PUBLISH_VERIFY_DELAY_SECONDS: "10"`, a 10-minute per-package budget), plus the comment documenting the observed near-miss. ## Verification - Same env-override plumbing verified in #11834; only the numeric budget changes. The next master push (this merge) exercises the canary path. ## Risks - Low risk: a genuinely failed publish reports in up to 10 minutes; healthy publishes exit the poll on first visibility. ## Model Used Claude Fable 5 (Anthropic, `claude-fable-5`) with extended thinking and agentic tool use via the Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable (not applicable: numeric workflow env tuning) - [x] I have updated relevant documentation to reflect my changes (workflow comment updated) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
16149a75fd |
ci(release): give npm publish visibility polling a 10-minute budget (#11834)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The release workflow publishes every public package to npm per
master push (canary) and per promotion channel
> - `release.sh` polls the registry after each publish and aborts the
whole release when a version is not visible within 60 seconds
> - npm accepts publishes immediately, but its CDN can lag packument
propagation by several minutes; on 2026-08-21 this failed four
consecutive canary runs mid-loop even though every publish succeeded
> - This pull request sets the script's existing visibility-budget env
overrides at the workflow level to 10 minutes
> - The benefit is a release train that tolerates registry propagation
lag; a healthy publish still exits the poll on its first visible check
## Linked Issues or Issue Description
Not applicable for a `ci:` workflow tuning change. Evidence: four
consecutive `Release` runs on master failed in `publish_canary` with
"npm did not publish and expose <package>@<version>", while the raw
registry packument shows each of those versions present minutes later
(`2026.821.0-canary.0` accepted 01:36 UTC, visible 01:40;
`2026.821.0-canary.1` accepted 05:54, visible 05:57).
## What Changed
- Set `NPM_PUBLISH_VERIFY_ATTEMPTS: "30"` and
`NPM_PUBLISH_VERIFY_DELAY_SECONDS: "10"` in the `Release` workflow's
top-level `env`, raising `release.sh`'s post-publish visibility poll
from 60 seconds to 5 minutes per package for every channel. Both
variables are existing overrides read by the script
(`scripts/release.sh` lines 311-312); no script change.
- Raised the four publish jobs' `timeout-minutes` from 45 to 90 so
several laggard packages fit inside the job without exhausting it before
the tag push / Docker / release steps.
## Verification
- `release.sh` reads the two env overrides with defaults
(`${NPM_PUBLISH_VERIFY_ATTEMPTS:-12}` /
`${NPM_PUBLISH_VERIFY_DELAY_SECONDS:-5}`), so workflow-level env reaches
`publish_package_to_npm_and_wait` unchanged.
- Not run: a live release (needs the npm-canary environment). The next
master push exercises the canary path with the new budget.
## Risks
- Low risk: a genuinely failed publish now takes up to 5 minutes to
report instead of 1, and a pathological batch where most packages lag
the full budget still fails inside the 90-minute job — that pattern
means a real registry incident. The poll exits early on success, so
healthy releases are unaffected.
## Model Used
Claude Fable 5 (Anthropic, `claude-fable-5`) with extended thinking and
agentic tool use via the Claude Code CLI (release log forensics against
raw registry packument timestamps).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable (not applicable:
workflow env tuning)
- [x] I have updated relevant documentation to reflect my changes
(comment in the workflow documents the budget rationale)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
db4defdfbf |
feat: operator-configurable settings visibility via PAPERCLIP_HIDDEN_SETTINGS (#11823)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The instance settings surface (Access, Plugins, Adapters, General,
Experimental) assumes the person at the keyboard operates the whole
instance
> - Operators who host Paperclip for others — a managed cloud or an
internal shared server — expose settings pages and toggles that do not
apply to their deployment, and the related mutation APIs stay open
> - A hosted tenant can open Plugins or Adapters, try an action, and hit
a confusing failure, because only a few hardcoded platform floors exist
> - This pull request adds a generic, operator-configured visibility
mechanism: one env var hides declared settings surfaces in the UI and
floors their mutation routes with a stable 403 code
> - The benefit is a clean hosted-tenant settings surface for any
operator, with zero behavior change for normal self-hosted instances
## Linked Issues or Issue Description
**Subsystem affected**
Instance settings (server routes and UI), the shared settings registry
in `packages/shared`, and the `/api/health` bootstrap payload.
**Problem or motivation**
An operator who hosts Paperclip for other people cannot hide settings
surfaces that the platform manages. Tenants see Access, Plugins, and
Adapters pages, backup retention, and host-level experimental toggles
that do nothing useful for them. The mutation APIs behind these surfaces
also stay open, so a tenant admin can attempt actions the platform must
control. ROADMAP.md names a cleaner shared deployment story as a goal
("Teams should be able to run the same product in hosted or semi-hosted
environments without changing the mental model").
**Proposed solution**
Add a declarative registry of hideable settings surfaces and one env
var, `PAPERCLIP_HIDDEN_SETTINGS`. The server parses the list at boot,
reports it on `/api/health`, and rejects value-changing writes to hidden
surfaces with a stable `settings_operator_managed` 403 code. The UI
reads the list from the health payload and removes the hidden pages,
sections, and toggles from navigation, routes, and page content. Unknown
keys warn and are ignored, so one list can roll across a fleet with
mixed app versions. With the variable unset, behavior is byte-identical
to today.
**Alternatives considered**
- Hardcode the hidden set for cloud instances in this repo: rejected,
because each hosting operator needs a different policy, and policy does
not belong in shared code.
- Deliver the hidden set through the managed-config document: rejected,
because that channel is cloud-specific and fail-closed on unknown
fields; a plain env var works for any operator, including self-hosted
shared servers.
- Lock the controls with a badge instead of hiding them: rejected for
these surfaces, because they are meaningless to tenants, not merely
platform-controlled; the existing managed-overlay lock stays the right
tool for controlled flags.
**Roadmap alignment**
Supports the "shared deployment story" item in ROADMAP.md: hosted and
semi-hosted deployments keep the same product with a settings surface
that matches what the tenant can actually do.
## What Changed
- New `packages/shared/src/settings-visibility.ts`: registry of hideable
surfaces (every instance settings page — profile, environments, access,
heartbeats, experimental, plugins, adapters; every Instance → General
section; every experimental flag as `instance.experimental.<key>`), the
`PAPERCLIP_HIDDEN_SETTINGS` parser, and the `settings_operator_managed`
error code. The General page stays visible as the settings root and
redirect target.
- New `server/src/services/settings-visibility.ts`: parse-once accessor;
unknown keys log one warning and are ignored.
- `/api/health` reports `hiddenSettings` on every response shape; the
field is omitted when nothing is hidden.
- Server floors on hidden surfaces, with same-value echo tolerance (the
`executionMode` precedent): field-backed general sections and
experimental keys reject value-changing PATCHes, and hiding the whole
Experimental page floors every toggle; plugin lifecycle and config
writes, adapter management writes, and the Access admin routes (reads
included) return 403 `settings_operator_managed`. Reads the app itself
needs (plugin `ui-contributions`, adapter metadata, plugin job trigger)
stay open. Pages without instance-scoped mutation routes are hidden in
the UI only.
- UI: new `useHiddenSettings` hook and `HiddenSettingsPageGate` route
gate (hidden pages redirect to the settings root); the settings sidebar
and tab bar drop hidden entries; remembered settings paths remap to the
default page; `InstanceGeneralSettings` skips hidden sections; every
`ExperimentalToggleCard` now carries its flag key and renders nothing
when hidden.
- Removed the dead `InstanceSidebar` component (referenced only by its
own test).
- Docs: `docs/deploy/environment-variables.md` documents the variable
and the key registry.
## Verification
- `pnpm vitest run` over the new and extended suites: shared registry
and parser, representative floor tests per route class (changed-value
403, same-value echo 200, unset env 200, page-level Experimental
hiding), the health field, the route gate, nav filtering, and
section/card hiding with one hidden example per surface kind — 168 tests
pass.
- Full root `pnpm typecheck` passes.
- Manual: booted a server with the variable set. `/api/health` lists the
keys; an unknown key logs one warning and the server boots; hidden pages
redirect; hidden sections and cards do not render; hidden-field PATCH
returns 403 with `details.code = "settings_operator_managed"`; a
same-value echo returns 200. Unset the variable: the full settings
surface returns and responses are byte-identical to master.
## Risks
- Low risk for self-hosted instances: with the variable unset, the
hidden set is empty, the health field is omitted, and no floor
activates.
- Flooring plugin config writes assumes hosted deployments configure
plugins through the platform. If a future bundled plugin needs
tenant-entered config, the floor needs a narrow carve-out.
- Hidden-key floors tolerate same-value echoes, so API clients that
round-trip full GET responses keep working.
- Hiding a toggle does not change its value; operators pair hiding with
the desired default where the value matters.
## Model Used
Claude Fable 5 (Anthropic, `claude-fable-5`) with extended thinking and
agentic tool use, driven through the Claude Code CLI (file edits, test
execution, and live-server verification loops).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
de9645ab73 |
fix(ui): surface live runtime status in the task-chat tail before the first transcript token (#11802)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs on sandbox execution targets spend their first minutes in preparation phases — config seed, workspace and skills sync into the sandbox — before the agent CLI produces its first transcript token > - The engine already reports these phases through the runtime-progress mechanism (`onRuntimeProgress` → `recordCurrentHeartbeatRunRuntimeProgress` → `currentStatusMessage` on the live run), and the pre-task-chat issue view surfaced them in its live status line > - The chat-style task view's live tail dropped that affordance: with zero renderable transcript entries it shows an opaque "Waiting for transcript..." for minutes, which reads as a hang (and prompted a real is-this-broken investigation on a healthy run) > - This pull request surfaces the live run's `currentStatusMessage` as the tail's empty-state message, with the generic wait text as fallback > - The benefit is that operators watching a sandbox run see "Syncing workspace to sandbox" instead of wondering whether the run is stuck — for every adapter, with no adapter identities involved ## Linked Issues or Issue Description **What existing behavior does this improve?** The chat-style task view's live transcript tail (`TaskChatThread` → `TaskChatLiveTail`), introduced with the experimental chat-style task view. **Current behavior** While a live run has no renderable transcript entries yet — the normal state for the multi-minute sandbox preparation window — the tail shows a static "Waiting for transcript...". The run's live `currentStatusMessage` (e.g. "Syncing workspace to sandbox", emitted by the sandbox-managed runtime's progress reporting) is available on the same live-run object but unused by this surface, although the earlier issue-chat view did display it. **Proposed behavior** When the tail is streaming a live run and no transcript rows exist yet, the empty-state message prefers the run's `currentStatusMessage`; the generic wait text remains the fallback when no runtime status has been reported (e.g. local runs that produce output immediately, or the brief pre-status window). **Reason and benefit** The preparation phases are real, reportable progress that the engine already emits. Showing them turns a minutes-long apparent hang into a legible status, for every adapter and execution target, using data the view already receives. ## What Changed - `ui/src/components/TaskChatThread.tsx`: the live tail's `emptyMessage` prefers `liveRun.currentStatusMessage` (guarded to the run the tail is actually streaming) over the static "Waiting for transcript..." fallback. Queued runs keep "Waiting to start...". - `ui/src/components/TaskChatThread.test.tsx`: a test covering both branches — a live run with a runtime status shows it (and not the wait text), and a run without one keeps the generic message. ## Verification - `vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/task-chat/TaskChatLiveTail.test.tsx`: 22/22 pass (including the new test) - `pnpm --filter @paperclipai/ui typecheck`: clean - Reproduced live: a `claude_local` run on a Daytona sandbox environment showed "Waiting for transcript..." for the full sync window; with this change the same window shows the streamed preparation statuses ## Risks - Low: a one-expression change to an empty-state string, active only while a live run has produced no renderable transcript rows. The fallback path is byte-identical to today. - `currentStatusMessage` is truncated/humanized upstream by the runtime-progress reporter; this surface renders it verbatim in the same muted style as the wait text. ## Model Used - Anthropic, **Claude Fable 5** (`claude-fable-5`) via Claude Code, with repository, shell, and Git tooling. It traced the runtime-progress mechanism end-to-end (engine emitter → heartbeat recorder → live-run API → both thread views), identified the dropped affordance in the chat-style view, and wrote the fix and test. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
933749e01f |
test(server): deflake postgres teardown and pinned exposure port (#11667)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server test suite gates every merge and every release cut. > - Three server tests each failed exactly once on markdown-only or unrelated diffs, then passed on rerun. > - One of the three (the git-operation-scheduler owner/joiner race) was fixed on master independently by [#11671](https://github.com/paperclipai/paperclip/pull/11671) while this PR was open, so after rebasing this pull request carries the remaining two. > - A flaky gate makes release operators rerun CI and stop trusting red results. > - Each remaining flake has a real nondeterminism: a teardown race and a hard-coded host port. > - This pull request removes the nondeterminism from the two tests without weakening what they prove. > - The benefit is a test gate that fails only when the product is broken. ## Linked Issues or Issue Description - [x] I searched open and closed issues and pull requests for these test files and for these failures. I found no duplicate report or fix. **What happened?** Three one-off CI failures occurred during release operations, each on a diff that could not have caused it, and each passed on rerun: 1. Run [32086355930](https://github.com/paperclipai/paperclip/actions/runs/32086355930): `server/src/__tests__/interaction-resolution-cross-issue-cap-postgres.test.ts` — all 7 tests passed, but vitest recorded an Unhandled Error and failed the run: `TypeError: Cannot read properties of null (reading 'write')` at `postgres@3.4.9/src/connection.js:255 Immediate.nextWrite`. 2. Run [32096743814](https://github.com/paperclipai/paperclip/actions/runs/32096743814): `server/src/services/workspace-git-operation-scheduler.test.ts` — the test "coalesces the same canonical key and cleans single-flight state after success and failure" failed with an AssertionError: the two concurrent calls came back with the `singleFlightJoined` values swapped. *(Fixed on master by [#11671](https://github.com/paperclipai/paperclip/pull/11671) with an equivalent single-flight barrier while this PR was open; the fix was dropped from this PR on rebase and the file is no longer touched here.)* 3. Run [32196201529](https://github.com/paperclipai/paperclip/actions/runs/32196201529): `server/src/services/workspace-runtime-exposure.test.ts` — the test "keeps an existing runtime port that is already inside the dedicated range" failed once out of 610 recorded runs because the runtime came back on a relocated port instead of the pinned 42500. **Expected behavior** The tests pass on every run when the code under test is correct. A red result means a product defect, not scheduling luck on the CI host. **Steps to reproduce** Each flake is a low-probability race, but both remaining mechanisms reproduce deterministically: 1. Postgres teardown: the suite never ends the postgres.js pool behind `createDb`; `afterAll` only stops the embedded server. postgres.js batches small writes and flushes them with `setImmediate` (`connection.js` `nextWrite`), and `close()` nulls the socket. Stop the server while the pool is open and a pending flush can run after the socket is gone. 2. Exposure pinned port: hold any loopback socket on 42500 or 52500 (both are inside the default Linux ephemeral port range, 32768–60999) and run the test. The allocator correctly relocates, and the assertion fails with `expected 42000 to be 42500`. The client side of any loopback connection on the CI host can land on those ports. **Paperclip version or commit** Branched from `master` at `4b968d8c0`; rebased onto `5a1ce7aed`. **Privacy checklist** I reviewed this description and removed private instance URLs, internal task identifiers, credentials, and user paths. ## What Changed Both fixes are test-side. I found no product race. - `interaction-resolution-cross-issue-cap-postgres.test.ts`: `afterAll` now ends the drizzle/postgres.js pool (`db.$client.end()`) before it stops the embedded Postgres server. `end()` waits for in-flight queries, including a fire-and-forget wake that lands just after a response, and closes the sockets from the client side first. Sibling suites (for example `heartbeat-plugin-environment.test.ts`) already use this order; this suite had skipped the pool shutdown. - `workspace-runtime-exposure.test.ts`: the pinned-port test no longer hard-codes 42500. It scans the dedicated range with the suite's real loopback probe, finds the lowest free app/HMR pair, then pins the next free pair strictly above it. If the keep-preferred-port path broke, the ascending fallback scan would return the lower pair, so the assertion keeps its discriminating power while no longer betting on one fixed host port staying free. - *(Dropped on rebase: the `workspace-git-operation-scheduler.test.ts` coalescing fix, superseded by the equivalent barrier merged in [#11671](https://github.com/paperclipai/paperclip/pull/11671).)* ## Verification - Reproduced the exposure flake exactly: with a listener held on `127.0.0.1:52500`, the pre-fix test fails with `expected 42000 to be 42500`; the fixed test passes with the port still held. - The postgres flake is a probabilistic teardown race and I could not trigger it on demand. The mechanism is established from `postgres@3.4.9` source (`setImmediate`-batched `nextWrite` versus `close()` nulling the socket) and the fix removes the whole class by closing the pool before the server. - Repeat runs after the fix: the pinned-port exposure test 20/20 green while the Postgres suite looped concurrently for loopback churn; `interaction-resolution-cross-issue-cap-postgres.test.ts` 15/15 green with no unhandled errors. - Re-verified after rebasing onto `5a1ce7aed`: both changed test files pass and `tsc --noEmit` passes in `server/`. - Environment note: three unrelated tests in `workspace-runtime-exposure.test.ts` (the wildcard-bind diagnosis tests) fail on macOS before and after this change because they read `/proc`; they are untouched and pass on Linux CI. ## Risks - Low risk: both changes are test-only; no product code changed. - The pinned-port test keeps a tiny time-of-check/time-of-use window between its own probe and the runtime's bind. The window shrinks from "one fixed port must stay free across the whole CI fleet" to milliseconds on a pair just verified free. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Fable 5 (Claude Code) — model ID `claude-fable-5`, with repository tools and local code execution for reproduction and repeat-run verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1d0e826767 |
feat(acpx/ui): adapter-declared capabilities for verbose streaming backends (no behavior change by default) (#11761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local agent adapters stream their work through the shared acpx engine, which every ACP adapter (`claude_local`, `codex_local`, `gemini_local`, custom ACP) runs on > - Verbose streaming backends break two shared-engine behaviors: the auto-posted run summary concatenates every text delta including the thought stream (a long multi-tool run once auto-posted a ~50k character monologue as an issue comment), and token-by-token tool-argument streaming floods the run log with tens of thousands of placeholder-titled in-progress events per run > - These fixes were developed inside the `kimi_local` adapter PR, where Kimi Code's streaming volume (~16,000 text deltas per run vs ~290 for a comparable Claude run) surfaced both problems > - Changing this behavior for all adapters at once is a fleet-wide risk, and hardcoding adapter identities in shared code does not scale to many adapters (or work at all for externally-shipped plugin adapters) — so the behaviors become invocation-config parameters that an adapter's own acpx config builder sets, with defaults preserving today's behavior byte-for-byte > - The benefit is that the machinery lands fully tested with zero behavior change for existing adapters — pin tests prove it — the engine carries no adapter identities, and any adapter (including `custom_acp` configs for external backends) opts in declaratively ## Linked Issues or Issue Description - Refs #9967 — extracted from the `kimi_local` adapter PR and restructured to be inert by default; the commit preserves the original author's (@hawikk) authorship. **Current behavior** When an acpx-engine run ends without the agent leaving a comment, the auto-posted summary is every streamed text delta concatenated, thought stream included. Backends that stream tool arguments emit tens of thousands of placeholder-titled `in_progress` tool events into the stored run log, pinning the live activity indicator to a generic "tool call". There is no mechanism for an adapter to vary either behavior, and shared code must never branch on adapter identities. **Proposed behavior** Two engine invocation-config parameters, read with behavior-preserving defaults: `summaryStrategy` (`"full"` = existing concatenation, the default; `"lastOutputSegment"` = segment output at tool starts, exclude thought stream, post the last non-empty segment) and `coalescePlaceholderToolUpdates` (`false` = never drop an event, the default; `true` = coalesce placeholder-titled in-progress updates). An adapter opts in from its own acpx config builder — the engine has no per-adapter knowledge, no adapter identity appears anywhere in shared code, and `custom_acp` agent configs can set the same knobs for external verbose backends. **Reason and benefit** Existing adapters are provably unaffected — new pin tests assert the default path's summary and tool-event output byte-for-byte, so any future change that alters behavior for claude/codex/gemini/custom fails the suite. The verbose-backend handling still lands fully tested, activated declaratively by the adapter that needs it (the `kimi_local` adapter PR sets both knobs in its config builder). ## What Changed - `packages/adapter-utils/src/acpx-engine/execute.ts`: the run preparation parses `summaryStrategy` and `coalescePlaceholderToolUpdates` from the invocation config (validated, defaulted); summary accumulation and `emitRuntimeEvent` branch on the prepared values. The default path is the pre-existing code (`textParts.join("")`, no event filtering). `buildAcpxRunSummary` is the exported last-segment strategy. - `packages/adapter-utils/src/acpx-engine/execute.test.ts`: a pin test asserting the default path's exact summary (thought stream included) and full tool-event stream (placeholder-titled in-progress updates present, names restored); opt-in tests for each knob; a `buildAcpxRunSummary` unit test. - `ui/src/adapters/types.ts`: `UIAdapterModule` gains an optional `transcriptPresentation` capability — `maxVisibleEntries` (issue-chat transcript window, default 30) and `liveReasoningView` (`"ticker"` default; `"scrollLog"` renders live reasoning in a scrollable auto-following box with one entry per tool call). - `ui/src/lib/issue-chat-messages.ts` and `ui/src/components/IssueChatThread.tsx`: shared code resolves the hints via `findUIAdapter(adapterType)` with today's defaults as fallback — no adapter identities anywhere. The `scrollLog` rendering component ships here but is unreachable until an adapter declares it. - `ui/src/lib/issue-chat-messages.test.ts`: a capability test registers a synthetic verbose adapter and asserts the wider window; the pre-existing test keeps pinning the default 30-entry window. No adapter declares any of this in this PR — every adapter renders and summarizes exactly as before, and there is no per-adapter data anywhere. The `kimi_local` adapter PR (#9967, stacked on this branch) is the first consumer: it declares `transcriptPresentation` in its own UI module and sets the engine knobs in its own acpx config builder. ## Verification - Engine suite: 130/130 pass (126 existing + 4 new); chat suites (`issue-chat-messages`, `IssueChatThread`, `RunChatSurface`): all pass including the new synthetic-adapter capability test — 237 tests across the touched surfaces - The pin tests are the regression guard: the engine test encodes today's summary text and tool-event sequence for a default-config run, and the existing 30-entry-window test pins the default transcript window, so "nothing changed for Claude/Codex users" is an executable assertion, not a review judgment - `pnpm --filter @paperclipai/adapter-utils --filter @paperclipai/ui typecheck`: clean ## Risks - Low: with no adapter setting the knobs, every code path taken in production is the existing one. The only behavioral surface is additive (an unused strategy and an unused filter), exercised by tests. - The knobs are ordinary invocation-config keys, so a `custom_acp` agent config can also set them — intended: an external verbose backend gets the same handling without code changes. Both knobs only affect that agent's own run summaries and run-log verbosity. ## Model Used - Original implementation authored in #9967 by @hawikk (models documented there: Moonshot AI Kimi K3 Coding via Kimi Code CLI 0.27.0, OpenAI GPT-5 Codex, Anthropic Claude Opus 4.8). The commit preserves that authorship. - Extraction, restructuring into config-declared parameters, pin tests, and verification: Anthropic, **Claude Fable 5** (`claude-fable-5`) via Claude Code, with repository, shell, and Git tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Hawik <davapa@gmail.com> |
||
|
|
b5a3a863c3 |
feat(release): bootstrap new npm packages with a placeholder publish (#11757)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its release pipeline publishes a set of npm packages from CI with npm trusted publishing (GitHub OIDC), gated by `scripts/release-package-manifest.json` > - A brand-new package name cannot be published by CI directly: the PR bootstrap gate requires the name to resolve on npm, and a trusted-publisher rule can only be configured after the package page exists > - The current bootstrap helper closes that gap by building the package locally and publishing its real output from a maintainer machine — before the PR that adds the package has passed CI or review > - This pull request replaces that flow: the helper now publishes a minimal deprecated placeholder at version `0.0.0` that only reserves the name, so every real version ships from CI > - The benefit is that unreviewed build output never reaches npm, and the bootstrap runs from any checkout (including `master`, before the new package's PR merges) with no local build ## Linked Issues or Issue Description **What existing behavior does this improve?** The one-time npm bootstrap for a brand-new release package (`pnpm run release:bootstrap-package`). **Current behavior** The helper builds the target package locally and publishes the real build output from a maintainer machine. That content has not passed repository CI or review at publish time. The helper also requires the new package to exist in the local workspace, so it must run from the (unmerged) PR branch that adds the package. **Proposed behavior** The helper publishes a three-file placeholder at version `0.0.0` (manifest, README, and an `index.js` that throws a descriptive error), waits for the registry to show the package, then deprecates it. The PR bootstrap gate (`scripts/check-release-package-bootstrap.mjs`) only requires the name to resolve on the registry, so the placeholder satisfies it. The first real calver release from CI supersedes the placeholder, and a stable release moves `latest` off it — the same `latest` window that existed under the old flow, but containing an explicit inert stub instead of unreviewed code. **Reason and benefit** Real package content only ever reaches npm from CI, after review and merge. The bootstrap becomes safer (scope guard refuses names outside `@paperclipai/`, already-published names are rejected) and simpler (no local build, no workspace state, runs from any checkout). **Breaking changes** None at runtime. The helper's CLI surface changes: it now takes a package name only (no directory selector) and drops `--skip-build`. `doc/PUBLISHING.md` is updated to match. ## What Changed - `scripts/bootstrap-npm-package.mjs`: replaced the build-and-publish flow with a placeholder publish — stages `package.json` + `README.md` + throwing `index.js` at version `0.0.0` in a temp directory, previews with `npm publish --dry-run`, and publishes only with `--publish`. One-time passwords are prompted interactively (never passed as arguments, since they are single-use and would land in shell history), with re-prompt on a rejected or expired code. After publishing, the helper polls the registry until the package is visible (a first publish can lag by minutes; verified live at ~5 minutes), requiring two consecutive sightings before prompting for a second code and deprecating the placeholder so accidental installs warn loudly; on timeout or failure it prints the exact manual `npm deprecate` command. Added an `@paperclipai/`-scope guard and a fail-fast error when `--publish` runs without an interactive terminal. Removed the workspace-plan dependency so it runs from any checkout. - `scripts/bootstrap-npm-package.test.mjs`: rewrote for the new interface — argument parsing, scope validation, the generated placeholder files (manifest shape, throwing entry point, README), the OTP re-prompt loop, and the registry poll (consecutive-sighting requirement, timeout, transient-error tolerance) via injected fakes. - `doc/PUBLISHING.md`: rewrote the "One-time bootstrap sequence for a new package" section for the placeholder flow, including the `latest` dist-tag window and the trusted-publishing setup ordering (placeholder publish → trusted publisher rule → `"publishFromCi": true`). - `.github/scripts/check-pr-release-bootstrap.mjs` (+ test, + wiring in `run-quality-gates.mjs`): new informational commitperclip notice on PRs that need this bootstrap. It fires when the PR newly release-enables a package that is missing from npm, or adds an unpublished `publishFromCi: false` package that published packages declare a `workspace:*` dependency on, and names the exact maintainer command — so contributors know the red `policy` check is not theirs to fix. It never fails the gate (the `policy` job remains the enforcer), only looks up scope-validated names on the registry, and stays quiet on registry errors. ## Verification - `node --test scripts/bootstrap-npm-package.test.mjs`: 13/13 pass - `node --test .github/scripts/tests/*.test.mjs`: 147/147 pass (10 new for the PR notice) - `pnpm run test:release-registry`: 82/82 pass - Replayed the new PR notice against a real historical PR's live API data (files, manifest at base and head refs): with the registry in its pre-bootstrap state it produces the exact maintainer instruction; with the package bootstrapped it stays silent - Full live end-to-end run: the flow bootstrapped `@paperclipai/adapter-kimi-local` for real — dry-run preview (634-byte, 3-file tarball), publish, registry visibility after ~5 minutes of propagation lag, deprecation confirmed via `npm view ... deprecated` - Guards verified live: an already-published name is rejected, an out-of-scope name (`left-pad`) is rejected, unknown options (including the removed `--otp`) are rejected, and `--publish` in a non-interactive shell fails fast before any network call ## Risks - The `latest` dist-tag points at the deprecated `0.0.0` placeholder until the first stable release supersedes it. This window also existed under the old flow (which parked `latest` at a locally built version); internal consumers are unaffected because release version rewrites pin exact calver versions. - The registry poll caps at ~10 minutes. If propagation is slower than that, the helper prints the exact `npm deprecate ... --otp <code>` command to run manually once `npm view` resolves. - The helper no longer validates the name against the workspace release plan, so a typo within the `@paperclipai/` scope would reserve a wrong name. The dry-run preview shows the exact name before any publish. ## Model Used - Anthropic, **Claude Fable 5** (`claude-fable-5`) via Claude Code, with repository, shell, and Git tooling. It analyzed the existing bootstrap flow and the release scripts (`release-package-map.mjs`, `check-release-package-bootstrap.mjs`, `release.sh` dist-tag handling), wrote the replacement script and tests, updated the documentation, and ran the verification above. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
927ef58d04 |
docs(skills): changelog entries describe deltas, not repeats (#11663)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Stable release notes are drafted per release from the commit range
since the previous stable
> - Features evolve across consecutive releases, so a correct range
still produces entries that re-describe what the previous notes already
introduced
> - The first channel-train draft did exactly that for four themes, and
review caught it against the published previous changelog
> - This pull request adds the missing framing rule to the changelog
skill and mirrors it in the Discord skill
> - The benefit is that consecutive releases read as a progression
instead of repeating themselves
## Linked Issues or Issue Description
**What existing behavior does this improve?**
The authoring guidance in `.agents/skills/release-changelog/SKILL.md`
and `.agents/skills/release-changelog-discord-message/SKILL.md`.
**Current behavior**
The skills define the range and the sections but say nothing about
features the previous stable's notes already introduced. A draft for a
follow-up release naturally re-describes them as if they debuted
(observed in the v2026.821.0 draft: the chat conversation view, sandbox
output streaming, the raised import cap, and the channel system — all
already announced in v2026.817.0's notes).
**Proposed behavior**
The changelog skill instructs the author to read the previous stable's
notes first and phrase already-introduced features as deltas ("last
release introduced X; this release makes it the default"), demoting
follow-through themes out of Highlights. The Discord skill's highlight
guidance mirrors the rule.
**Reason and benefit**
Consecutive changelogs read as a progression; readers of both releases
never see the same debut twice.
## What Changed
- `.agents/skills/release-changelog/SKILL.md`: a "describe deltas, not
repeats" guideline in Step 4, with the read-the-previous-notes
instruction and the headline-demotion rule.
- `.agents/skills/release-changelog-discord-message/SKILL.md`: one
mirrored bullet in the template notes.
## Verification
- Docs-only; proofread. The rule matches the fix applied to the live
v2026.821.0 draft (#11661), which is the worked example of what it
prevents.
## Risks
- None; docs-only.
## Model Used
Claude Fable 5 (Claude Code)
## Pre-submission checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
|
||
|
|
ffaac1d7f9 |
docs(release): stable notes for the 2026.818.0-beta.1 soak (v2026.821.0) (#11661)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channels promote a soaked beta to stable, and the promotion reads the stable's release notes from `releases/beta/v<beta-version>.md` on master > - Beta `2026.818.0-beta.1` is soaking now, and the release workflow auto-drafted its notes skeleton on this branch at publish time > - This pull request rewrites that skeleton into the finished changelog for the upcoming stable > - The stable preflight fails without this file on master, so this PR must merge before promotion day > - The benefit is release notes reviewed during the soak instead of written at the gate ## Linked Issues or Issue Description **What existing behavior does this improve?** The release-notes flow for the upcoming stable (planned `v2026.821.0`, promoting beta `2026.818.0-beta.1`, source `664052f8e`). **Current behavior** The branch holds the machine-generated skeleton: grouped commit subjects for the 172-commit range from v2026.817.0's content baseline (`8f7b8b3fd`). **Proposed behavior** The finished changelog in the established release voice: verified Breaking Changes, Highlights, Improvements, Fixes, an Upgrade Guide covering migrations 0212–0222 and the new environment variables, and a verified community contributor list. **Reason and benefit** Notes get real review during the soak; the promotion just reads the merged file. ## What Changed - Rewrite `releases/beta/v2026.818.0-beta.1.md` from skeleton to finished notes. Headline themes: chat-style tasks as the default experience, Tailscale HTTPS managed runtime exposure, sandbox capability contract with live output streaming, in-product Claude/Codex sign-in, chunked resumable company imports, and the completed release-channel automation. - Five verified breaking/behavior changes, including the `enableTaskChatRedesign` → `enableClassicTaskInterface` toggle inversion and migration `0218`'s conservative resolver-policy remap. - Contributors: 172 commits from 32 non-bot authors; 22 verified community handles listed (founders/core excluded per the canonical list; 4 authors omitted as unverifiable rather than guessed). ## Verification - Every cited PR number resolves to a commit subject in the range. - No internal ticket identifiers. - Handles verified via noreply emails, co-author trailers, or `gh api users/<name>`. - The version/date header says the planned promotion (`v2026.821.0`, 2026-08-21); if promotion slips, only that header needs a touch-up — the beta-keyed filename is date-independent. ## Risks - Docs-only. If a newer beta supersedes `2026.818.0-beta.1` before promotion, this file stays as history and the new beta gets its own draft. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
e434829889 |
fix(release): draft-notes baseline survives candidate-cut stables (#11647)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The release workflow drafts the upcoming stable's notes skeleton at
beta publish, ranging from the last stable to the beta's source commit
> - Stables cut from candidate branches leave their tag off master's
lineage, so the ancestor-only baseline rule skips them and falls back a
whole release too far
> - The first live draft did exactly that: it ranged from v2026.722.0
instead of the just-shipped v2026.817.0 and re-included everything
already released
> - This pull request replaces the baseline rule with a merge-base walk
over stable tags and adds a candidate-branch regression fixture
> - The benefit is a correct draft range for every stable lineage:
master-cut, candidate-cut, and promoted-older-source
## Linked Issues or Issue Description
**What happened?**
The first `draft_stable_notes` run (for beta `2026.818.0-beta.1`)
generated `releases/beta/v2026.818.0-beta.1.md` with the range
`v2026.722.0..664052f8e` — 311-commits-worth of already-shipped
v2026.817.0 content re-included.
**Expected behavior**
The draft ranges from the point the shipped stable's content diverges
from the beta source. For v2026.817.0 (cut from
`candidate/release-2026.817.0`) that is the promoted source commit
`8f7b8b3fd`, giving the 172 commits of genuinely-new work.
**Steps to reproduce**
Ship a stable from a candidate branch (tag lands off master's lineage),
then publish a beta from master and read the generated skeleton's range
line.
**Paperclip version or commit**
master at `664052f8e`.
Related (not duplicates): #11567 introduced the generator; its
ancestor-only rule was itself a review fix for the newest-by-version
rule, and this PR is the second iteration with the lineage case the
first fix missed.
## What Changed
- `scripts/draft-stable-notes.sh`: the range start comes from walking
stable tags newest-first and taking the first whose merge-base with the
beta source is a proper ancestor of the source. Candidate-cut stables
resolve to the promoted commit, master-lineage stables to the tag
itself, and tags containing the source are skipped (an older-source
promotion cannot produce an empty range). The skeleton header prints a
runnable short-sha range with the stable tag as a labeled baseline.
- `scripts/draft-stable-notes.test.mjs`: new regression fixture with the
stable tag on an unmerged candidate branch; the existing older-source
and fallback fixtures still pass unchanged.
## Verification
- `node --test scripts/draft-stable-notes.test.mjs` — 8 pass, including
the new fixture.
- Regenerated the live `2026.818.0-beta.1` draft against the real
repository: range start resolves to `v2026.817.0 (merge-base
|
||
|
|
393da0f67c |
fix(adapter-utils): graft unrelated imported histories instead of failing the run (#11638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - At run finalize, the host imports the sandbox git history and reconciles it with the local worktree in `integrateImportedGitHead` > - Transported workspaces are depth-1 shallow clones, so the boundary commit reads as parentless inside the sandbox > - A `git commit --amend` there rewrites the boundary commit into a root commit, and the re-imported history no longer connects to the host history > - `git merge-tree` has no common base to merge against, so the sync throws "Failed to merge concurrent remote git histories" and the run fails with its work stranded in the sandbox > - This pull request grafts the imported tree onto the current head as a single commit instead of failing > - The benefit is that a history rewrite inside the sandbox can no longer lose a run's work ## Linked Issues or Issue Description No existing issue found. I searched issues and PRs for "unrelated histories", "Failed to merge concurrent", and "shallow". Depends on #11637 (merged; the graft commit reuses its identity constant). This PR is now rebased onto `master`. **What happened?** An agent run amended a commit inside its sandbox workspace to address review feedback. The sandbox clone is depth-1 shallow, so git treated the boundary commit as parentless and the amend produced a root commit. At finalize, the host-side sync failed with `Failed to merge concurrent remote git histories for <sha>` and the run was marked failed. A follow-up run had to repair the branch by hand: fetch the true parent from origin and rebuild the commit with `git commit-tree`. **Expected behavior** The sync must never strand completed work. When the imported history shares no ancestor with the local one, the imported tree should still land on the current head, with the imported message preserved and the graft recorded. **Steps to reproduce** 1. Start a run whose workspace transport uses the shallow clone path (`withShallowGitWorkspaceClone`, depth 1). 2. Inside the sandbox workspace, run `git commit --amend` on the boundary commit. The result is a parentless root commit. 3. Finish the run. The host-side `integrateImportedGitHead` finds no merge base, `merge-tree` fails, and the run fails. ## What Changed - `git-workspace-sync.ts`: new exported `createUnrelatedHistoryGraftCommit` helper. It reads the imported head's tree and message, and creates one commit on top of the current head with the deterministic sync identity and a trailer that records the graft and both shas. - `integrateImportedGitHead` (both the remote-git-sync version and the SSH copy in `ssh.ts`): when `merge-base` reports no common ancestor, graft instead of throwing. The ref update keeps the same compare-and-swap and concurrent-retry semantics as the merge path. - The graft is gated on `git merge-base` exiting with status 1 — the no-ancestor signal. Operational failures (timeout, missing object, repository error) keep the loud merge failure instead of rewriting the tip. - New regression tests: one builds the exact shallow-amend shape (a root commit rebuilt from the base tree) and asserts the graft lands on the current head with the imported tree, subject, and graft trailer; one integrates a well-formed sha the repository does not hold and asserts the integration still throws with the branch tip unchanged. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts` — 19/19 pass (includes the new graft test and the merge-base failure-discrimination test). - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - Full `pnpm vitest run packages/adapter-utils`: every file passes except `local-process-sandbox.test.ts`, which fails identically on an untouched `master` checkout on macOS (bubblewrap-dependent, pre-existing, unrelated). ## Risks - Behavioral shift: unrelated imported histories previously failed the integration; now they land as a squash-graft. In this degenerate case there is no base to merge against, so the imported tree is taken wholesale and concurrent local-only tree changes are superseded at the tip. The local commits keep their place in the graft's ancestry, and the trailer records both shas, so nothing is unrecoverable. The old behavior lost the imported work instead, which is the worse failure for an autonomous run. - The graft reuses the imported head's commit message, so branch history still reads naturally after a sandbox rewrite. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, via Claude Code CLI (tool use for code exploration, test runs, and verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no doc surface describes this internal sync path) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9ea8143c87 |
fix(adapter-utils): give sync-created merge commits a deterministic git identity (#11637)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs execute in transported workspaces; at finalize, the host syncs the sandbox git history back into the local worktree > - When both sides advanced, `integrateImportedGitHead` reconciles them with `git merge-tree` plus `git commit-tree` on the host > - Execution hosts are often containers with no git config and no resolvable hostname, so `commit-tree` fails with "Author identity unknown" > - That one local command failure marks the whole run as failed, even though the run's work succeeded > - This pull request gives sync-created merge commits an explicit, deterministic identity at the call site > - The benefit is that workspace finalize no longer depends on ambient host git configuration ## Linked Issues or Issue Description No existing issue found. I searched issues and PRs for "Author identity unknown", "unrelated histories", and "commit-tree identity". **What happened?** A run finished its work, but workspace finalize failed. The host-side sync ran `git commit-tree <tree> -p <localHead> -p <importedHead> -m "Paperclip remote git sync merge <sha>"`. Git exited with `Author identity unknown ... fatal: unable to auto-detect email address (got 'node@<container-id>.(none)')`. The adapter recorded the whole run as failed, and the host worktree kept the stale head. Any container deployment without a global gitconfig reproduces this; I observed it on a Paperclip Cloud stack. **Expected behavior** Commits that the sync machinery itself creates must not depend on ambient host git configuration. The merge commit is machine-authored, so it should carry a deterministic Paperclip identity. **Steps to reproduce** 1. Run the Paperclip server in a container with no `user.name`/`user.email` git config and a hostname git cannot turn into an email. 2. Let a run's sandbox branch diverge from the host worktree, so both sides advance. 3. Workspace finalize calls `integrateImportedGitHead`. The `git commit-tree` step fails with "Author identity unknown" and the run fails. ## What Changed - `git-workspace-sync.ts`: new exported `GIT_SYNC_COMMIT_IDENTITY_ARGS` (`-c user.name=Paperclip -c user.email=noreply@paperclip.ing`), applied to the `commit-tree` call in `integrateImportedGitHead`. - `ssh.ts`: the SSH-sync copy of `integrateImportedGitHead` applies the same identity args to its `commit-tree` call. - New regression test: builds divergent histories in a repo with no configured identity and asserts the sync merge commit is created with the deterministic identity, correct parents, and merged tree. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts` — 18/18 pass. - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - Negative proof: with the source fix stashed, the new test fails on the identity assertion. - Full `pnpm vitest run packages/adapter-utils`: every file passes except `local-process-sandbox.test.ts`, which fails identically on an untouched `master` checkout on macOS (bubblewrap-dependent, pre-existing, unrelated). ## Risks Low risk. The change only adds `-c` identity flags to two machine-generated commit invocations. `GIT_AUTHOR_*` / `GIT_COMMITTER_*` environment variables still take precedence over `-c` when an operator sets them, so existing deployments that configure an identity keep their behavior. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, via Claude Code CLI (tool use for code exploration, test runs, and verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no doc surface describes this internal sync path) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
664052f8ea |
feat(release): draft stable notes at beta publish, read them from master at promotion (#11567)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system promotes builds canary → nightly → beta → stable, and stable releases publish a GitHub Release from `releases/vYYYY.MDD.P.md` > - The stable lane requires that notes file to exist inside the promoted source commit, but the file is named for the promotion date, which is unknown when the source commit is created > - A promoted beta can therefore never pass the notes check: every happy-path stable is forced through the candidate-branch fix path, with a soak-gate justification, for a notes-only change > - This pull request drafts the notes automatically when the beta is published and lets the stable promotion read them from `master` > - The benefit is a walkable stable happy path: the soak gate stays exact, notes get a real review window during the soak, and the justification path returns to its real purpose (cherry-picked fixes) ## Linked Issues or Issue Description **What existing behavior does this improve?** The stable promotion path in the release channel system (`release.yml`, `scripts/release.sh`). **Current behavior** `release.sh stable` requires `releases/vYYYY.MDD.P.md` in the checked-out source tree, and `publish_stable` checks out the exact promoted SHA. The soak gate requires a `beta/v*` tag to point at that same SHA. No commit can satisfy both for a promoted beta, so a stable promotion must cut a candidate branch with a notes-only commit and bypass the soak gate with a written justification. Release notes are also written at promotion time, under time pressure, with no review window. **Proposed behavior** When a beta publishes, a `draft_stable_notes` job generates a grouped notes skeleton at `releases/beta/v<beta-version>.md` and pushes it to a machine-owned branch; a human opens the PR and edits it during the 3-day soak. The stable preflight resolves notes before the `npm-stable` approval gate: source-tree notes first (the candidate fix path, unchanged), then the merged beta-keyed file on `master`; it fails early with the missing path named when neither exists. After the stable ships, a canonicalization job pushes a branch that moves the file to `releases/vYYYY.MDD.P.md`. Related (not duplicates): #11006 and #11008 introduced the nightly and beta lanes this builds on; older changelog PRs (for example #10669) authored notes manually at promotion time, which is the flow this replaces. **Reason and benefit** The happy path becomes: promote the exact soaked SHA, no justification, notes reviewed during the soak instead of written at the gate. The `releases/vYYYY.MDD.P.md` invariant still holds durably via the canonicalization PR. ## What Changed - `scripts/release.sh`: new `--notes-file PATH` (stable only) overrides where the pre-publish notes check looks, so notes can live outside the source checkout without dirtying the worktree. - `scripts/create-github-release.sh`: same `--notes-file` override for the GitHub Release body. - `scripts/draft-stable-notes.sh` (new): deterministic skeleton generator — commit subjects from the newest stable tag (falling back to the previous beta, then full history) to the beta's source commit, grouped into Features / Fixes / Other. - `.github/workflows/release.yml`: - `draft_stable_notes` job after `publish_beta`: runs the generator and force-pushes `release-notes/v<beta-version>`; the job summary links the compare page. It recreates the beta tag locally if the tag push was rejected (the known workflows-permission case), so drafting is not blocked on manual tag recovery. - `preflight_stable`: computes the target stable version (`release.sh stable --print-version`) and resolves the notes source (`source_tree` → `master_beta` → fail early / warn on dry run); new outputs. - `publish_stable`: materializes `master`-side notes into `RUNNER_TEMP` and passes `--notes-file` to both scripts; outputs the published stable version. - `canonicalize_stable_notes` job: pushes the `git mv` branch after a stable that used `master`-side notes. - `doc/RELEASING.md`, `doc/RELEASE-CHECKLIST.md`: document the drafted-notes flow, the preflight resolution order, and the canonicalization step; the LLM changelog flow now targets the draft branch during the soak. - `.agents/skills/release-changelog/SKILL.md`, `.agents/skills/release-changelog-discord-message/SKILL.md`: the notes-authoring skills now describe this flow — range ends at the beta source commit (not `HEAD`), the file is beta-keyed on the `release-notes/v<beta-version>` branch (seeded with `scripts/draft-stable-notes.sh` for betas that predate the automation), and the canonicalization link caveat is called out for announcements. ## Verification - `node --test scripts/draft-stable-notes.test.mjs` — 6 tests, temp git-repo fixtures: grouping, stable-tag range, previous-beta and full-history fallbacks, default output path, malformed version, missing tag. - `node --test scripts/release-lib.test.mjs` — unchanged suite still green. - `bash -n` on both changed shell scripts; `release.yml` re-parsed as YAML. - `./scripts/release.sh stable --print-version` unchanged (prints the next stable version); `--notes-file` on a non-stable channel fails with a clear error. - Not exercised end-to-end: the new workflow jobs need a real beta publish to run. The first beta after merge is the live test; the draft job is additive and cannot affect the publish result (it runs after `publish_beta` completes). ## Risks - Low risk to publishing itself: `--notes-file` defaults preserve today's behavior everywhere; the draft and canonicalization jobs are additive and run after the publishes succeed. - The preflight now fails a real stable run when no notes are found. That is the intended fail-early behavior (it previously failed later, inside `publish_stable`, after the `npm-stable` approval). - `draft_stable_notes` force-pushes only the machine-owned `release-notes/v<beta-version>` branch; a beta re-cut regenerates it cleanly. - The stable version computed at preflight could differ from the published one if a run crosses UTC midnight between the two jobs; the materialized notes are passed by path, so the publish still succeeds, and the canonicalization job uses the actually-published version. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
6691c57e54 |
docs(release): reconcile v2026.817.0 release notes to master (#11590)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Stable release notes live at `releases/vYYYY.MDD.P.md` on master as the durable record > - v2026.817.0 shipped from a candidate branch, so its notes file exists on that branch and in the GitHub Release, but not yet on master > - This pull request lands the exact shipped notes at the canonical path > - The benefit is a complete stable-notes record on master, and the FULL RELEASE NOTES link in announcements resolves ## Linked Issues or Issue Description **What existing behavior does this improve?** The `releases/` record on master. **Current behavior** `releases/` on master ends at `v2026.722.0.md`. The v2026.817.0 notes exist only on `candidate/release-2026.817.0` and in the published GitHub Release. **Proposed behavior** `releases/v2026.817.0.md` exists on master, byte-identical to what the release published. **Reason and benefit** Complete history at the canonical path; the standard notes link (`releases/v2026.817.0.md` on master) resolves for the announcement. ## What Changed - Add `releases/v2026.817.0.md`, taken verbatim from the candidate branch the release was published from (tag `v2026.817.0`, commit `213dabab4`). ## Verification - `git diff v2026.817.0 -- releases/v2026.817.0.md` against this branch is empty (byte-identical to the shipped file). - Docs-only; no code surface. ## Risks - None; docs-only. The candidate branch is deleted after this merges (the `v2026.817.0` tag keeps its commits reachable). ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
49217aadf0 |
refactor: balance serialized server shards by recorded suite duration (#11528)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its wall-clock time sets the feedback loop for all contributors > - In a recent successful PR run (actions run 32012408876), the slowest check was "Verify serialized server suites (1/5)" at 337s, while its four sibling shards finished in 212-238s > - The serialized lane assigns suites to shards round-robin over an alphabetical list, so the heavy heartbeat and issues suites cluster on one runner > - The general-server lane already solves this with a duration-aware LPT partition backed by a recorded manifest > - This pull request reuses that partitioner for the serialized lane with a fresh per-suite duration manifest > - The benefit is a balanced serialized matrix: the measured 968s suite total levels to about 194s per shard, which removes about 80-100s from the run's slowest check ## Linked Issues or Issue Description **What existing behavior does this improve?** The `Verify serialized server suites` shard matrix in `.github/workflows/pr.yml` distributes route/authz test suites across five runners. **Subsystem affected** CI / test infrastructure (`scripts/run-vitest-stable.mjs`). **Current behavior** `selectSerializedSuites` assigns suites round-robin (`index % shardCount`) over the alphabetically sorted file list. The heavy suites cluster on shard 1/5. In actions run 32012408876, shard 1/5 spent 291s in its test step while the other shards spent 170-201s, which made that job (337s total) the slowest check of the whole PR run. **Proposed behavior** Partition the serialized suites with the same duration-aware LPT algorithm the general-server lane already uses (`scripts/general-server-shard.mjs`), backed by a new per-suite duration manifest. All five shards then carry about 194s of measured test time. **Reason and benefit** The slowest check bounds PR feedback time. Balancing the serialized matrix removes about 80-100s from that bound without adding runners. **Breaking changes** None. The partition remains deterministic, complete, and non-overlapping; suites missing from the manifest get the median weight. ## What Changed - Added `scripts/serialized-shard-durations.json`: per-suite wall-clock durations (ms) for all 134 serialized suites, sampled from actions run 32012408876 by diffing consecutive per-suite label timestamps in the shard logs (captures vitest spawn overhead, not just reported test time) - `scripts/run-vitest-stable.mjs`: `selectSerializedSuites` now uses the existing LPT partitioner (`selectGeneralServerShard`) with the new manifest instead of round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: added a manifest-freshness test and a shard-balance test for the serialized lane, mirroring the general-server ones - `.github/workflows/pr.yml`: updated the serialized matrix comment with the new measurement and mechanism ## Verification - `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` passes (13 tests), including the existing test that the serialized shards form a complete, non-overlapping partition - Dry-run of all five shards shows estimated totals of 194/194/194/194/193s (round-robin was 276/175/160/172/187s): `node scripts/run-vitest-stable.mjs --mode serialized --shard-index N --shard-count 5 --dry-run` - The `Verify serialized server suites` jobs on this PR run the real partition end to end ## Risks - Low risk. Selection logic only; the vitest invocation per suite is unchanged - A stale manifest degrades gracefully: unknown suites get the median weight, and a dedicated test fails if fewer than half the current suites have recorded durations ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK); no extended-thinking mode ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #10923 (split serialized tests into five shards), #10925 (general-server duration manifest), #11156 (workspaces-a native shards). Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
48f4ae16ac |
fix(codex-local): keep a promoted device-login credential when re-seeding the managed home (#11578)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The codex_local adapter supports a device login that runs in a trusted sandbox and promotes the credential into the per-company managed Codex home > - The managed-home re-seeding step treats every regular-file `auth.json` as apikey-mode residue and removes it so the shared-home symlink can be restored > - The promotion writes the company credential as a regular file, so the first environment Test or run after a successful login deletes it > - On a server with no shared Codex login — any containerized deployment — nothing replaces the file, and the UI reports that the sandbox has no ready authentication right after it reported a successful login > - This pull request makes the cleanup identity-anchored: a subscription credential whose identity the shared source does not hold survives re-seeding > - The benefit is that a device login stays usable after Test and runs, on hosts with and without a shared Codex login ## Linked Issues or Issue Description No public GitHub issue covers this. The problem is described in-PR following the bug template. Related public PRs: [#11237](https://github.com/paperclipai/paperclip/pull/11237) added the sandbox device login and the credential promotion, [#11097](https://github.com/paperclipai/paperclip/pull/11097) added its building blocks, and the `ensureSymlink` heal for stale copies came from the fix for #5028. [#9621](https://github.com/paperclipai/paperclip/pull/9621) touches the adjacent sandbox auth sync-back lane but not this defect. **Subsystem affected** packages/adapters/codex-local — managed `CODEX_HOME` seeding (`codex-home.ts`). **Current behavior** A successful device login promotes the subscription `auth.json` into the company Codex home as a regular file, and the UI reports the login as authenticated. The next `seedManagedCodexHome` call — the environment Test probe and every execute both run it — removes any regular-file `auth.json` when no API key is configured, because the cleanup assumes such a file is apikey-mode residue left by a previous run. It then symlinks `auth.json` from the shared source home. On a server whose shared home has no Codex login (a container image, for example), there is no source to symlink, so the home ends with no credential at all. The Test probe then reports "The sandbox has no ready authentication for this adapter" immediately after a successful login, and a fresh login repeats the same cycle. On a server whose shared home does hold a login, the symlink silently replaces the promoted account with the host account. **Expected behavior** The credential a device login promoted stays in the company home across Test probes and runs. The #5028 heal (a stale regular-file copy of the shared credential becomes a symlink to the live source) and the apikey-residue cleanup keep working. **Steps to reproduce** 1. Run the server in an environment whose shared Codex home (`$CODEX_HOME` or `~/.codex`) has no `auth.json`. 2. Complete a Codex device login for a company; the promotion writes the company home `auth.json` and the UI reports authenticated. 3. Click Test on a codex_local agent (or start a run). The probe reports no ready authentication, and the promoted `auth.json` is gone from the company home. **Proposed solution** Make the cleanup identity-anchored, the same rule the promotion and the cache vend already use. A regular-file `auth.json` survives re-seeding when it holds a usable subscription identity that the shared source does not also hold, and the shared symlink does not replace it. A same-identity regular file is still the #5028 stale copy and is still healed into the symlink, because the symlink serves the same account with live, rotating tokens. An apikey-mode or unreadable file is still removed. ## What Changed - `seedManagedCodexHome` reads the target `auth.json` before the cleanup and keeps it when `readSubscriptionAccountId` yields an identity the shared source `auth.json` does not hold. The kept file is excluded from the shared symlink pass, and the function logs a fixed line when it keeps the file. - The function doc comment states the kept-promoted-credential rule. - Four new `seedManagedCodexHome` test cases: a promoted credential with no shared auth, a promoted credential with a different shared identity, the same-identity #5028 heal, and apikey-mode residue removal. ## Verification ```sh cd packages/adapters/codex-local npx tsc --noEmit # clean npx vitest run # 28 files, 321 passed, 1 skipped ``` The four new cases fail on the previous code: the first two observed the promoted file deleted (and, with a shared login present, replaced by the shared symlink). ## Risks Low risk. The change narrows one deletion path. Deployments that never use the device login see no difference: without a promoted subscription file, the cleanup and the symlink behave exactly as before, and the #5028 heal is pinned by an existing test plus a new same-identity test. The one deliberate behavioral shift: after a device login, the promoted company credential now stays authoritative over the shared host login for that company — which is the promotion's documented contract ("the company credential slot"). ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use and code execution — investigation, implementation, and tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
adfb357223 |
docs(skills): maintain the release-credits exclusion list (#11588)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Stable release notes end with a Contributors section that credits
community contributors
> - The release-changelog skill excluded founders from that list, but
had no rule for non-founder core contributors
> - A release-notes draft therefore credited two core contributors as
community contributors
> - This pull request makes the exclusion list canonical in the
changelog skill and adds the core contributors
> - The benefit is that release credits consistently mean "community",
with one list to maintain
## Linked Issues or Issue Description
**What existing behavior does this improve?**
The Contributors rules in `.agents/skills/release-changelog/SKILL.md`
and the Community-section rules in
`.agents/skills/release-changelog-discord-message/SKILL.md`.
**Current behavior**
The changelog skill says to exclude "Paperclip founders (e.g.
`cryppadotta`, `forgottendev`, `devinfoley`, `sockmonster`,
`scotttong`)". Core contributors who are not founders are not covered,
so drafts credit them in the community list. The Discord skill refers to
"founders" in two places.
**Proposed behavior**
The changelog skill holds the canonical exclusion list of specific folks
— now including `nguyenm7`, `nickyleach`, and `tonio-alucema` — and both
Discord-skill references defer to that one list.
**Reason and benefit**
Release credits consistently mean community contributions, and there is
exactly one place to update when the core team changes.
## What Changed
- `.agents/skills/release-changelog/SKILL.md`: the exclusion rule names
specific folks, is the canonical list, and adds `nguyenm7`,
`nickyleach`, and `tonio-alucema`.
- `.agents/skills/release-changelog-discord-message/SKILL.md`: both
references ("Community" template note and the final checklist) defer to
the changelog skill's canonical list.
## Verification
- Docs-only change; rendered and proofread. No workflow, script, or test
surface is touched.
## Risks
- None beyond docs accuracy. If the canonical-list wording lands after
#11567 (which touches other sections of the same files), Git merges them
cleanly — the edited regions do not overlap.
## Model Used
Claude Fable 5 (Claude Code)
## Pre-submission checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
|
||
|
|
962e98b1be |
feat(server): operator declaration for platform edge TLS termination on the Claude login guard (#11579)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Claude local adapter supports a setup-token subscription login, and its confidential routes pass a fail-closed transport guard > - The guard accepts direct socket TLS, a local_trusted loopback peer, or an allowlisted proxy peer that forwards https — and deliberately never reads the global `TRUST_PROXY` > - On a managed platform the edge terminates TLS, the app socket is always plain HTTP, and the edge-proxy peer addresses are not stable or documented, so none of the three cases can hold > - Every login on such a deployment shows the clear-text transport warning although the user's connection is HTTPS, and the agent-scoped confidential routes fail closed entirely > - This pull request adds a dedicated operator declaration that the platform edge terminates TLS, as a fourth guard case > - The benefit is a correct transport decision on managed platforms with the default posture unchanged everywhere else ## Linked Issues or Issue Description No public GitHub issue covers this. The problem is described in-PR following the enhancement template. Related public PRs: [#11347](https://github.com/paperclipai/paperclip/pull/11347) added the new-agent login flow and the non-blocking transport advisory, and [#11286](https://github.com/paperclipai/paperclip/pull/11286) added the setup-token login and the guard with its `CLAUDE_LOGIN_TRUSTED_PROXIES` allowlist. **Subsystem affected** server/ — the confidential transport guard for the Claude setup-token login (`services/setup-token-session.ts`, `routes/agents.ts`, `app.ts`). **Current behavior** The guard allows a confidential response on direct socket TLS, on a `local_trusted` loopback peer, or when the immediate peer is on the dedicated `CLAUDE_LOGIN_TRUSTED_PROXIES` allowlist and forwards `https`. Behind a managed platform's TLS-terminating edge (Railway, Render, Fly, and similar), the app socket is plain HTTP and the edge-proxy peer addresses are not operator-visible or stable, so the allowlist cannot express them — IPv6 entries match by exact string only. The result: the login panel shows "This connection is not encrypted" for a connection that is HTTPS to the user, and the agent-scoped confidential routes return the fixed no-secret error. **Proposed behavior** `CLAUDE_LOGIN_EDGE_TLS_TERMINATED=true` is an explicit, single-purpose operator declaration that every client request reaches the server through the platform's TLS-terminating edge. Under the declaration the guard treats a request as confidential unless the edge itself labels the client hop as plain `http` in `X-Forwarded-Proto`. The declaration is never derived from the global `TRUST_PROXY` setting, which the guard still never reads. Without the declaration, nothing changes. **Reason and benefit** The guard's spoofing concern does not apply to this deployment shape: a client cannot pick its transport, because the platform admits HTTPS only, and the header the guard consults is set by the platform edge, not the client. A blanket warning that is always wrong teaches users to ignore it. The declaration keeps the strict default for every deployment that does not opt in, and it keeps the allowlist as the precise tool for operators who do know their proxy addresses. ## What Changed - `ConfidentialTransportConfig` gains optional `edgeTlsTerminated` (default false), documented as the operator declaration for platform edge TLS termination. - `evaluateConfidentialTransport` adds the declaration as a guard case: allowed unless the forwarded protocol's first hop is explicitly `http` (reason `edge_labeled_plain_http` then; `operator_edge_tls_termination` when allowed). - `assessConfidentialStartup` reports `edge_tls_termination_declared`, so the startup log shows why forwarded requests pass. - `app.ts` parses `CLAUDE_LOGIN_EDGE_TLS_TERMINATED` (truthy: `1/true/yes/on`) and passes it to the agent routes; the routes build the guard config from it. - The SR-7 operator-requirement comment on the setup-token routes documents the new variable next to the allowlist. - Tests: five new guard unit cases and a route case asserting the prompt and code responses carry no `transportAdvisory` under the declaration. ## Verification ```sh cd server npx tsc --noEmit # clean npx vitest run src/services/setup-token-session.test.ts \ src/routes/setup-token-route.test.ts \ src/__tests__/openapi-routes.test.ts # 3 files, 89 passed ``` The new "keeps failing closed when the declaration is absent" case pins the unchanged default posture. ## Risks The declaration is an operator statement the server cannot verify; an operator who sets it on a deployment whose edge does not terminate TLS re-labels plain-HTTP requests as confidential. This is the same trust class as `CLAUDE_LOGIN_TRUSTED_PROXIES` (a wrong allowlist entry has the same effect) and is opt-in, off by default, and scoped to the login routes only. The guard still fails closed when the edge explicitly labels a request `http`. No schema change, no API shape change — `transportAdvisory` was already nullable. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use and code execution — investigation, implementation, and tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2ec984502a |
fix(release): stop smoke_beta silently skipping on promote-mode betas (#11582)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system re-smokes every published beta as post-publish verification (`smoke_beta`) > - The candidate-branch beta lane (#11209) added `verify_beta_candidate` to `publish_beta`'s needs; that job is skipped on every normal promote-mode beta > - `smoke_beta`'s condition has no status-check function, so GitHub attaches an implicit `success()` that evaluates the needs chain transitively — a skipped ancestor makes it false > - This pull request makes the condition explicit so promote-mode betas smoke again, and pins the shape in the workflow wiring test > - The benefit is that the post-publish beta gate actually runs instead of silently skipping ## Linked Issues or Issue Description **What happened?** Beta `2026.818.0-beta.0` (run 32082007439) published successfully, but its post-publish `smoke_beta` job was skipped. No configuration or input asked for that: the run was a plain `channel: beta` dispatch with `dry_run` at its default `false`, and the same expression `!inputs.dry_run` evaluated true inside `publish_beta`'s own steps (the Docker dispatch step ran). **Expected behavior** Every non-dry-run beta publish is followed by the release smoke suite against the exact published version, as documented in `doc/RELEASING.md` and `doc/RELEASE-CHECKLIST.md`. **Steps to reproduce** Dispatch `release.yml` with `channel: beta` promoting a nightly (promote mode). `verify_beta_candidate` is skipped by design; `publish_beta` runs through its explicit `!cancelled()` condition; `smoke_beta` then skips because its implicit `success()` sees the skipped ancestor in the transitive needs chain (actions/runner#2205 semantics). The beta published on 2026-08-11 predated #11209, so this never surfaced before. **Paperclip version or commit** master at `43ab441f0` (workflow file, current head). Related (not duplicates): #11209 introduced the candidate lane whose skipped job triggers this; #11208 covers the adjacent tag-push failure playbooks. ## What Changed - `smoke_beta`'s condition becomes `!cancelled() && needs.publish_beta.result == 'success' && !inputs.dry_run` — an explicit status-check function suppresses the implicit `success()`, and the result check keeps the dependency on a successful publish. - A comment above the job records why the explicit form is load-bearing. - `scripts/__tests__/release-verify-workflow.test.mjs` pins the new shape so the implicit form cannot silently return. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 8 pass, including the new assertion. - `release.yml` re-parsed as YAML. - The exact skip is visible on run 32082007439 (`smoke_beta: skipped` after `publish_beta: success`); the coverage gap for that beta was closed manually by dispatching `release-smoke.yml` with `paperclip_version: beta` (run 32084880767). - Not exercised end-to-end: the corrected condition needs the next real promote-mode beta to demonstrate; the expression change is minimal and the semantics are the documented actions/runner behavior. ## Risks - Low risk: condition-only change on one job plus a test. Dry runs still skip the smoke (`!inputs.dry_run` retained). Candidate-mode betas, where `verify_beta_candidate` actually runs, behave as before. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
4af55ba6bd |
test(release-smoke): follow the onboarding wizard's chat-first rewrite (#11565)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system publishes a nightly build only after the release smoke suite passes against the newest canary > - The scheduled nightly run has failed every night since August 12, so no nightly, and therefore no beta candidate, has shipped for six days > - The failures are stale test locators, not a product regression: the chat-first onboarding rewrite (#11101) changed wizard copy and the post-launch destination > - This pull request updates the smoke spec to match the current wizard > - The benefit is a green nightly lane and an unblocked beta promotion ## Linked Issues or Issue Description **What happened?** The scheduled `Release` nightly run fails in `smoke_nightly / smoke` every night since 2026-08-12. The failing spec is `tests/release-smoke/docker-auth-onboarding.spec.ts`. Four assertions no longer match the product after the chat-first onboarding rewrite (#11101): - The step-1 heading is now "Name your organization", not "Name your company". - The step-4 hire button is now "Connect", not "Give it a heartbeat". - The seeded first task is now titled "Paperclip onboarding". - A successful launch navigates to the seeded task's thread (`/issues/<ref>`), not `/dashboard`. **Expected behavior** The smoke suite passes against a canary that contains the current onboarding wizard, and the nightly lane publishes again. **Steps to reproduce** Run `.github/workflows/release-smoke.yml` against `paperclipai@canary` (any version at or after the rewrite), or dispatch `release.yml` with `channel: nightly`. Example red runs: 32014452506 (Aug 17), 31938284401 (Aug 16). **Paperclip version or commit** `2026.817.0-canary.12` Related (not duplicates): #11190 updated this same spec for the mission-first wizard; this PR is the follow-up for the chat-first rewrite that landed after it. ## What Changed - Update the step-1 wizard heading locator to "Name your organization". - Update the step-4 hire button locator to "Connect" and reword the step comment. - Update `FIRST_TASK_TITLE` to "Paperclip onboarding" (the wizard's current `DEFAULT_TASK_TITLE`). - Assert the post-launch URL is the seeded task's thread (`/issues/`), not `/dashboard`. ## Verification - Local run of the exact CI harness: `scripts/docker-onboard-smoke.sh` with `PAPERCLIPAI_VERSION=2026.817.0-canary.12`, then `pnpm run test:release-smoke` against the container — 1 passed. - The suite's later API assertions (company, CEO agent, mission goal, seeded issue assignment, assignment-sourced heartbeat run) all pass unchanged against the current canary. ## Risks - Low risk: test-only change; no product code is touched. - The spec remains copy-coupled to the wizard. If wizard copy churn continues, a follow-up could add stable `data-testid` hooks to the wizard so the smoke spec stops breaking on wording changes. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
9b1fd42ac1 |
test(grok-local): isolate billing env in usage cost test (#11285)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters report run output, token use, and cost data. > - The Grok local adapter now reports real token use and cost data. > - Its new billing test must prove the no-key path and the API-key path. > - The no-key assertion used the caller environment without isolation. > - This made the test fail when `XAI_API_KEY` was already set. > - This pull request isolates that environment state in the test. > - The benefit is stable coverage for the cost gate from #10433. ## Linked Issues or Issue Description Refs #10433 **What happened?** The Grok local usage and cost test asserted subscription billing while it still used the ambient process environment. If `XAI_API_KEY` was set before the test ran, the adapter selected API billing instead. The subscription assertion could then fail on a developer machine or a CI runner with provider credentials. **Expected behavior** The test should prove the subscription path with no `XAI_API_KEY`. It should also prove the API billing path with a test key. **Steps to reproduce** 1. Start from `master` after #10433. 2. Set `XAI_API_KEY` in the shell environment. 3. Run `vitest` for `packages/adapters/grok-local/src/server/execute.test.ts`. 4. Observe that the subscription half can take the API billing branch without test isolation. **Paperclip version or commit** `master` after #10433. **Deployment mode** Built from source. ## What Changed - Isolated `XAI_API_KEY` with save, delete, set, and restore logic around both billing assertions. - Gave the subscription and API billing checks separate run ids and temp roots. ## Verification - `XAI_API_KEY=ambient-test-key corepack pnpm exec vitest run packages/adapters/grok-local/src/server/execute.test.ts` - `corepack pnpm --filter @paperclipai/adapter-grok-local typecheck` ## Risks Low risk. This changes test setup only. It does not change Grok local adapter runtime behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex local coding agent. The agent used shell tools, GitHub CLI, and local test execution. The context window size was not exposed in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
a09d7dcc06 |
feat(ui): bounce cold arrivals off archived company URLs, add Unarchive (#11302)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Archiving a company hides it from the sidebar switcher, but remembered last-visited paths, browser history, bookmarks, and restored tabs keep depositing users onto its URLs long after archiving > - Since the selection ping-pong fix (#11300) those arrivals render, but the user is stranded inside a workspace the sidebar refuses to show — and unarchiving had no UI anywhere, so the only way back was a hand-typed settings URL > - This pull request bounces cold arrivals at archived company URLs to an active company (with a toast naming why), lets deliberate visits stick, and adds an Unarchive action to the companies list > - The benefit is that stale URLs stop stranding users in retired workspaces, and archived companies become restorable from the one page that still lists them ## Linked Issues or Issue Description Follow-up to #11300. No existing issue for the remaining gap; description follows the enhancement template: **What happened?** After #11300, opening an archived company's URL (stale tab, history, bookmark, remembered path) renders that company's pages — but the sidebar switcher does not list it, so the user is stranded in a workspace they retired, and every stale URL pulls them back in. Separately, unarchiving a company has no UI: the archive button lives in company settings, which becomes unreachable through normal navigation once the company is archived. **Expected behavior** Arriving cold at an archived company's URL lands the user in an active workspace, with a toast explaining the redirect. Explicitly choosing the archived company (from the companies list) still works, so its pages remain reachable. Archived companies can be restored from the companies list. **Steps to reproduce** 1. Create two companies; archive one. 2. Open `/{archivedPrefix}/dashboard` directly — before: renders the archived workspace with no sidebar presence; after: bounces to the active company's dashboard with a toast. 3. On the companies list, open the archived company's row menu — before: no restore action anywhere; after: Unarchive. ## What Changed - `ui/src/lib/company-selection.ts`: `resolveArchivedCompanyBounce` — pure policy: bounce when the URL names an archived company that is not the current selection and an active company exists; prefer the currently selected active company as the destination. - `ui/src/components/Layout.tsx`: the route-sync effect applies the bounce (toast + selection + `replace` navigation) before syncing selection from the route. - `ui/src/pages/Companies.tsx`: Unarchive action (`PATCH status: "active"`) in the row menu for archived companies. - Tests: unit cases for the bounce policy; the e2e now drives all three behaviors (direct-load bounce with toast, re-arrival bounce, deliberate visit sticks) on top of the existing crash regression. ## Verification - `pnpm vitest run src/lib/company-selection.test.ts src/context/CompanyContext.test.tsx src/pages/Companies.test.tsx` in `ui/` — 20 tests pass. - `npx playwright test --config tests/e2e/playwright.config.ts archived-company-url` — passes, covering bounce, toast, and deliberate-visit paths. - `pnpm typecheck` in `ui/` — clean. ## Risks Low risk. The bounce only fires for archived-company URLs when the archived company is not already selected and an active company exists; all-archived instances render as before. Deliberate selection from the companies list is unaffected (selection equals the matched company, so no bounce). Unarchive reuses the existing `PATCH /api/companies/:id` status transition the server already supports. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution, Playwright e2e). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d24a79f741 |
ci(dependabot): surface major npm updates as one grouped weekly PR (#11307)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Dependabot keeps the npm dependency tree and the GitHub Actions workflows current with weekly update PRs > - The npm config ignores every major version bump with a wildcard `ignore` rule, and no other process reports pending majors > - Major-version debt grows silently, and ignore rules also suppress Dependabot security updates when the fix ships only in a newer major > - Individual major PRs are not a good replacement: the board decided in #7560 to keep the PR list mergeable, and a flood of breaking bumps works against that > - This pull request removes the blanket ignore and groups all pending majors into one weekly PR, while minors and patches keep one PR per bump > - The benefit is a standing, visible signal of pending major updates, at a cost of at most one extra PR per week ## Linked Issues or Issue Description **What existing behavior does this improve?** The Dependabot npm update flow configured in `.github/dependabot.yml`. **Current behavior** Dependabot opens weekly PRs for minor and patch npm updates. A wildcard `ignore` rule suppresses every major version update. No report or reminder replaces the suppressed PRs — the comment says "review those manually", but nothing triggers that review. Ignore rules also apply to Dependabot security updates, so a security fix that ships only in a newer major is suppressed as well. **Proposed behavior** Dependabot opens one grouped weekly PR that contains every pending major npm update. Minor and patch updates keep their current one-PR-per-bump flow. A deliberate hold on a specific major can use a targeted per-dependency `ignore` entry instead of the wildcard. **Reason and benefit** Silent major-version drift compounds: each skipped major makes the eventual upgrade jump larger and riskier, especially across peer-dependency families. A single grouped PR makes the backlog visible in the PR list without flooding it. When the grouped PR is green, it merges cheaply. When it is red, it is a visible standing task instead of invisible debt. **Breaking changes** None. This changes repository automation only. Runtime behavior, response shapes, and outputs are unchanged. **Additional context** Related history: #7483 grouped patch/minor updates by dependency type, and #7560 reverted that grouping because the resulting 26-package PR was hard to merge. This PR does not touch the patch/minor flow. It only groups majors, which currently produce no PRs at all — it adds a signal that does not exist today rather than replacing individually mergeable PRs. ## What Changed - Removed the wildcard `ignore` rule for `version-update:semver-major` from the npm ecosystem in `.github/dependabot.yml`. - Added a `major-updates` group (`applies-to: version-updates`, `update-types: ["major"]`, `patterns: ["*"]`) so all pending majors land in one weekly grouped PR. - Left the schedule, labels, PR limits, and the github-actions ecosystem unchanged. ## Verification - `npx js-yaml .github/dependabot.yml` parses cleanly and the `groups` stanza follows the Dependabot v2 schema (`applies-to`, `update-types`, `patterns`). - After merge: check Insights → Dependency graph → Dependabot for config errors. The next weekly run (Monday 06:00) opens a single `major-updates` grouped PR that lists the pending majors. - No code changed, so the test suite is unaffected. ## Risks - Low risk. This is CI/automation configuration only. - The first grouped PR may be large, and red if several majors break the build. That is the intended visibility mechanism, and it does not block other work. A noisy or deliberately held-back dependency can be excluded from the group with `exclude-patterns` or a targeted per-dependency `ignore` entry. - This does not regroup minors or patches, so it does not reintroduce what #7560 reverted. ## Model Used - Claude Fable 5 (Anthropic), model ID `claude-fable-5`, via Claude Code CLI, extended thinking and tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — N/A, YAML-only CI config change; validated with `js-yaml` - [ ] I have added or updated tests where applicable — N/A, no code changed - [x] I have updated relevant documentation to reflect my changes — none reference the Dependabot config - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — one e2e shard flaked on an unrelated MCP UI spec and passed on re-run with identical code - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
61a5b7c6f9 |
fix(ui): stop the selection ping-pong on archived company URLs (#11300)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The UI keeps a selected company in `CompanyProvider` with two writers: a bootstrap effect that repairs invalid selections, and a Layout route-sync effect that selects the company the URL prefix names > - The route-sync matches the URL against the full company list (archived included), while the bootstrap resolver only accepted companies from the sidebar-filtered non-archived list > - On any archived company's URL the two effects overwrite each other's selection in a synchronous loop until React throws error #185 ("Maximum update depth exceeded") and unmounts the root to a blank page — armed by remembered last-visited paths, back/forward navigation, or bookmarks, on first load and client navigation alike > - This pull request makes an already-selected company only need to exist, keeping the sidebar filter for fresh-boot resolution where no explicit selection exists > - The benefit is that archived company URLs render instead of blanking the entire app ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** Opening (or back-navigating to) a URL whose company prefix belongs to an archived company blanked the whole app with `Minified React error #185`. Console in dev mode: "Maximum update depth exceeded. This can happen when a component calls setState inside useEffect…". A workspace whose first/seeded company was archived hit this on every load of its remembered URL. **Expected behavior** An archived company's URL renders its pages (the company still exists and its API routes serve data). The sidebar simply does not feature archived companies, and fresh boots still land on a non-archived company. **Steps to reproduce** 1. Create two companies; archive one (`PATCH /api/companies/:id` with `status: "archived"`). 2. Navigate to `/{archivedPrefix}/dashboard` — direct load or client-side back-navigation. 3. Before this fix: React #185 and an unmounted blank page (reproduced deterministically by the new e2e test). ## What Changed - `ui/src/context/CompanyContext.tsx`: `resolveBootstrapCompanySelection` keeps an explicitly selected company that exists in the full company list; stored-id and default resolution still prefer sidebar (non-archived) companies. - `ui/src/context/CompanyContext.test.tsx`: resolver keeps an archived-but-existing selection; a truly deleted selection is still replaced. - `tests/e2e/archived-company-url.spec.ts`: end-to-end regression driving both field shapes (direct load and back-navigation onto an archived company URL); it failed with the exact #185 console errors before the fix and passes after. ## Verification - `pnpm vitest run src/context …` in `ui/` — 122 tests pass (includes the new resolver cases). - `npx playwright test --config tests/e2e/playwright.config.ts archived-company-url` — fails before the fix (captured "Maximum update depth exceeded" console errors), passes after. - `pnpm typecheck` in `ui/` — clean. ## Risks Low risk. The only behavioral change is that a selection naming an archived-but-existing company survives the bootstrap repair — previously that state was unreachable without crashing. Boots with no valid selection behave exactly as before (non-archived preferred), covered by the existing and new resolver tests. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution, Playwright-driven crash reproduction). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ff5fd62d07 |
Resolve the environment secret companyId context on first save (#11291)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments give agent runs an execution target, and their env vars can carry secrets, so writes need a company scope for secret bindings > - `PATCH /environments/:id` resolves that scope from a query param, existing bindings, or a single-membership actor — and fails closed otherwise > - A fresh environment has no bindings yet, so an admin whose memberships cannot pin one company gets a 422 on the very first env var save and can never create the first binding > - #11200 made this reachable from the UI: it opened the envVars-only edit on the managed sandbox row, but the UI never sends the company it already knows > - This pull request sends the page's selected company on environment updates and adds a server fallback to the instance's only company > - The benefit is that the first env var save works on fresh environments, while multi-company ambiguity still fails closed ## Linked Issues or Issue Description Refs #11200 **What happened?** On a fresh platform-managed sandbox environment, the first "Save environment variables" in the UI fails with HTTP 422: "Environment secret management requires a companyId context during the instance-scoped transition." The same failure hits any environment update that touches `envVars` or `config` when the environment has no secret bindings yet and the actor's memberships do not name exactly one company (for example an instance admin provisioned without a membership row, or a user in two companies). **Expected behavior** The save succeeds. The client tells the server which company scopes the secret bindings, and on a single-company instance the server can resolve the only possible scope by itself. **Steps to reproduce** 1. Provision a platform-managed sandbox environment (managed-config `environments` entry) so the row exists with no secret bindings. 2. Sign in as an instance admin whose memberships do not resolve to exactly one company. 3. Open Environments, edit the managed row, add an env var, and save. 4. The PATCH returns 422 with the companyId-context error. ## What Changed - `ui/src/api/environments.ts`: `environmentsApi.update` accepts an optional `companyId` and sends it as the `companyId` query param the server already reads. Renamed the `customImageCompanyQuery` helper to `companyIdQuery` since it now serves plain updates and probes too. - `ui/src/pages/CompanyEnvironments.tsx`: both the managed envVars-only save and the general environment save pass `selectedCompanyId`. - `server/src/routes/environments.ts`: `resolveEnvironmentSecretContextCompanyId` falls back to the instance's only company when exactly one exists, mirroring the existing fallback in `resolveCustomImageCompanyId`. Multi-company instances still fail closed. - Tests: three new route tests (explicit query context for a multi-company actor, single-company-instance fallback for a membership-less admin, fail-closed regression on a multi-company instance), and updated the SSH-probe and probe-ambiguity tests for the new fallback. ## Verification - `cd server && pnpm vitest run src/__tests__/environment-routes.test.ts` — 78 pass. - `cd ui && pnpm vitest run src/pages/CompanyEnvironments.test.tsx` — 22 pass. - Full `server` and `ui` vitest suites and `tsc --noEmit` on both packages pass locally. - Manual: on a managed instance, edit the managed sandbox environment, add an env var, save. Before: 422. After: the save persists and the PATCH carries `?companyId=`. ## Risks - Low risk. The explicit query param path already existed on the server; the UI now uses it. - Behavioral shift: on single-company instances, environment secret-context resolution (including `POST /environments/:id/probe`) now resolves the only company instead of returning no context. That lets probes resolve secret-backed config where they previously returned a 422 asking for an explicit companyId. Multi-company instances keep the fail-closed behavior, covered by a regression test. - Self-hosted single-company instances gain the same first-save fix; no schema or config changes. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c57c0f7498 |
fix(sandbox-providers): accept bsdtar listings in the syncOut tarball confinement check (#11289)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers (Daytona, Kubernetes) sync run results back to the host with a sandbox-authored tarball > - Before extraction, a confinement check parses the host `tar -tvf` listing and fails closed on unparseable lines > - The parser only understands the GNU tar listing dialect; macOS ships bsdtar, whose ls-style listing never matches > - Every sandbox syncOut on a macOS host therefore aborts with "refusing tarball with an unparseable entry listing", and the run fails at copy-back > - This pull request teaches the parser both dialects while keeping the fail-closed and traversal guarantees > - The benefit is that Daytona and Kubernetes sandbox runs work on macOS hosts, with no behavior change on Linux ## Linked Issues or Issue Description No existing issue. Bug description: **What happened?** On a macOS host, every Daytona sandbox run fails at syncOut. The adapter reports: `Daytona syncOut refusing tarball with an unparseable entry listing: -rw-r--r-- 0 daytona daytona 7560 Aug 11 21:43 AGENTS.md`. The Kubernetes provider has the same parser and fails the same way. **Expected behavior** The confinement check accepts a well-formed listing from the host tar, whichever dialect the host tar emits. It still rejects members that escape the extraction directory, and it still fails closed on lines it cannot parse. **Steps to reproduce** 1. Run Paperclip on macOS (system tar is bsdtar). 2. Configure an agent with the Daytona sandbox provider. 3. Trigger any run that syncs files back from the sandbox. 4. The run fails at syncOut with the unparseable-entry-listing error, because bsdtar prints `<perms> <links> <user> <group> <size> <Mon> <day> <time|year> <name>` while the parser expects the GNU `<perms> <owner>/<group> <size> <date> <time> <name>` shape. **Operating system** macOS (bsdtar 3.5.3). Linux hosts with GNU tar are unaffected. ## What Changed - Extracted the listing-line parse in both providers' `file-sync.ts` into an exported `parseTarVerboseListingLine` that accepts the GNU/busybox dialect and the bsdtar (libarchive) dialect. - The GNU shape now requires the slash-joined `<owner>/<group>` field. This keeps the two shapes mutually exclusive. Without it, a bsdtar line with numeric uid/gid satisfies the loose GNU pattern shifted by one field, which would hide a leading `../` from the traversal check. - Unparseable lines still fail closed. This includes device-node entries, whose size column is `major,minor` in both dialects. - Made the path-traversal fixture in the Daytona suite portable: GNU spells member renaming `--transform`, bsdtar spells it `-s`. - Added a Daytona test that refuses a sandbox-authored tarball carrying a symlink whose target escapes the extraction dir. - Added parser unit tests for both dialects (file, dir, symlink, hardlink, numeric owner, year-form dates, fail-closed lines) to both providers' suites. ## Verification - `pnpm test` in `packages/plugins/sandbox-providers/daytona`: 136/136 pass on a macOS host. On unpatched `master` the round-trip test fails there with the unparseable-entry-listing error. - `pnpm test` in `packages/plugins/sandbox-providers/kubernetes`: the new parser tests pass; no new failures against the `master` baseline on the same host. - `pnpm typecheck` passes in both packages. - CI runs the same suites on Linux/GNU tar and proves the GNU path is unchanged. ## Risks - Low risk. The GNU pattern is one token stricter (`<owner>/<group>` must contain `/`). GNU and busybox tar always print the slash-joined owner field, so accepted GNU listings are unchanged. - The bsdtar branch only widens acceptance on hosts that were failing 100% of syncOuts before, so no working deployment changes behavior. - The check still fails closed on anything neither pattern matches. ## Model Used Claude Fable 5 (`claude-fable-5`, Claude Code CLI, extended thinking + tool use). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
01112c350c |
fix(ui): always give respondWith a real Response in the sw fetch fallback (#11292)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The UI registers a service worker (`ui/public/sw.js`) with a
network-first fetch handler whose cache is an offline fallback
> - The fallback hands `event.respondWith` the result of
`caches.match(...)`, which resolves `undefined` on a cache miss — and in
the navigation branch, `caches.match("/") || offlineResponse` never uses
the fallback because `caches.match` returns a promise, which is always
truthy
> - When the network fetch rejects (server restart, deploy, brief
outage) and the cache misses, the browser fails the request with
`Uncaught (in promise) TypeError: Failed to convert value to
'Response'`, so navigation breaks outright instead of degrading to the
offline page
> - This pull request awaits the cache lookups and guarantees a real
`Response` on every path
> - The benefit is that brief server unavailability degrades to the
offline fallback instead of a dead navigation
## Linked Issues or Issue Description
No existing issue. Description follows the bug template:
**What happened?**
Navigating while the server was briefly unavailable (mid-restart)
produced `The FetchEvent for "…" resulted in a network error response:
the promise was rejected.` and `sw.js:1 Uncaught (in promise) TypeError:
Failed to convert value to 'Response'.` The navigation failed instead of
showing the offline fallback.
**Expected behavior**
A failed navigation serves the cached app shell when present, otherwise
the "Offline" 503 response. A failed asset fetch serves its cache entry
when present, otherwise a proper network-error response. `respondWith`
always receives a real `Response`.
**Steps to reproduce**
1. Load the app so `sw.js` is active; ensure `/` is not in the service
worker cache (fresh cache version).
2. Restart or stop the backend.
3. Navigate to any page: the fetch rejects, `caches.match` misses, and
the browser logs the conversion TypeError with a failed navigation.
## What Changed
- `ui/public/sw.js`: the fetch fallback awaits `caches.match(...)` and
returns the "Offline" 503 for navigations and `Response.error()` for
assets when the cache misses.
- `ui/src/lib/sw-offline-fallback.test.ts`: evaluates the real `sw.js`
in a sandboxed scope and covers the three fallback paths; the two
miss-path tests fail against the previous code.
## Verification
- `pnpm vitest run src/lib/sw-offline-fallback.test.ts` in `ui/` — 3
tests pass.
- Verified both miss-path tests fail against the unmodified `sw.js`.
## Risks
Low risk. The change only affects the fetch-rejection path; successful
fetches and cache hits behave exactly as before. `Response.error()`
mirrors what the browser would produce for an unhandled failed no-cors
fetch.
## Model Used
- Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code
CLI with tool use (code search, edit, test execution).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
2c53437fc9 |
fix(server): authenticate cloud-proxied browsers on the live-events websocket (#11290)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The UI receives live run/issue events over a websocket at `/api/companies/:id/events/ws`; the server authorizes upgrades with a bearer token or a Better Auth session > - On a cloud-managed deployment, browsers authenticate through trusted `x-paperclip-cloud-*` headers injected by the managing front door — they never hold a local Better Auth session, and the Express middleware lane that understands those headers is not consulted for websocket upgrades > - Every browser websocket upgrade behind the front door therefore resolves no identity and is rejected 403: the live-events socket has never connected on a managed instance, leaving permanent reconnect churn and console failure noise while the UI silently degrades to polling > - This pull request adds a cloud-actor lane to the upgrade authorization, reusing the same trusted-header resolver the HTTP middleware uses > - The benefit is working realtime updates on managed instances, an end to the reconnect churn, and unchanged self-hosted behavior ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** On a cloud-managed instance, the browser console shows `WebSocket connection to 'wss://…/api/companies/<id>/events/ws' failed:` repeating indefinitely for every company, on a healthy instance. The server rejects each upgrade with 403 because `authorizeUpgrade` in `server/src/realtime/live-events-ws.ts` only knows bearer tokens and Better Auth sessions, while cloud-proxied browsers authenticate via `x-paperclip-cloud-*` trusted headers (handled only by the Express `actorMiddleware` lane in `server/src/middleware/auth.ts`). **Expected behavior** A browser that authenticates through the trusted cloud headers can open the live-events websocket for any company in its membership scope, exactly as it can call the HTTP API for those companies. **Steps to reproduce** 1. Run Paperclip in `authenticated` mode behind a proxy that injects the `x-paperclip-cloud-*` headers with a valid `PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN`. 2. Load any company page in a browser (no local Better Auth session). 3. HTTP API calls succeed; every `/events/ws` upgrade is rejected 403 and the UI retries forever. ## What Changed - `server/src/middleware/auth.ts`: `resolveCloudTenantActor` now accepts a minimal `CloudActorHeaderSource` (`header(name)`) instead of an Express `Request` — `Request` satisfies it unchanged — plus `cloudActorHeaderSourceFromHeaders` to adapt raw `IncomingMessage.headers`. - `server/src/realtime/live-events-ws.ts`: `authorizeUpgrade` gains an injected `resolveCloudActor` lane, tried before the Better Auth session fallback in `authenticated` mode. A resolved cloud actor is authoritative: the upgrade is authorized only for a company in the actor's membership scope (`companyIds`, the same scope the HTTP lane grants). Absent/unresolvable cloud headers fall through to the session path. - `server/src/index.ts`: wires `resolveCloudActor` through `resolveCloudTenantActor` + the header shim. The resolver self-gates: without `PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN` and a matching trust token it returns null, so self-hosted deployments never take this path. - Tests: upgrade authorized for an in-scope company (session resolver not consulted), rejected for an out-of-scope company, fall-through to session auth when no cloud actor resolves; header-shim resolution from a raw lowercased header map including `string[]` values. ## Verification - `pnpm vitest run server/src/__tests__/live-events-ws.test.ts server/src/middleware/cloud-tenant-actor.test.ts` — 25 tests pass. - `pnpm typecheck` in `server/` — clean. - Not verified live end-to-end: that requires a managed instance running this build; the direct probe evidence (HTTP authenticated fine, every WS upgrade 403) matches the code path exactly. ## Risks Low risk. The new lane only activates when the deployment configures the cloud trust token and the request presents it; both checks already protect the HTTP lane. Authorization scope is the same `companyIds` set the HTTP middleware computes (primary stack company plus the user's real membership rows). The cloud resolver's user/company materialization writes are debounced (existing behavior shared with the HTTP lane), so websocket reconnect storms do not amplify database writes. Self-hosted instances see no behavioral change, covered by the fall-through test. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution; diagnosis included live websocket handshake probes against a managed instance). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f16071a290 |
fix(ui): use a fictional tailnet hostname in the vite proxy test (#11248)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The UI dev server proxies `/api` to the backend and injects `x-forwarded-host`; a unit test asserts that injection with a sample Host header > - The sample Host header is a contributor's real machine and tailnet hostname, committed to the public repository > - Real personal hostnames do not belong in a public codebase, and this one also contains the contributor's OS username, so `scripts/check-forbidden-tokens.mjs` (which forbids the local username) blocks `npm` publishing from that contributor's machine > - This pull request replaces the fixture with a fictional tailnet-style hostname > - The benefit is no personal identifiers in the test fixtures and a passing forbidden-token check for every contributor ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** `ui/src/lib/vite-api-proxy.test.ts` uses a real contributor dev-machine hostname as its `Host` header fixture. `node scripts/check-forbidden-tokens.mjs` fails on that contributor's machine because the hostname contains their OS username, blocking the publish flow. Introduced in #10718. **Expected behavior** Test fixtures use fictional hostnames. The forbidden-token check passes on every contributor machine. **Steps to reproduce** 1. On a machine whose OS username appears in the fixture hostname, run `node scripts/check-forbidden-tokens.mjs`. 2. The check reports the two lines in `ui/src/lib/vite-api-proxy.test.ts` and blocks with exit code 1. ## What Changed - `ui/src/lib/vite-api-proxy.test.ts`: the `Host` fixture is now `dev-box.tail1234.ts.net:3101` (fictional). The test only asserts that whatever host arrives is injected as `x-forwarded-host`, so the value is arbitrary. ## Verification - `pnpm vitest run src/lib/vite-api-proxy.test.ts` in `ui/` — 5 tests pass. - `node scripts/check-forbidden-tokens.mjs` — "No forbidden tokens found" on the previously affected machine. - Note: git history retains the old value; this removes it from the current tree only. ## Risks None. A test fixture string with no behavioral coupling. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d90f4d488e |
fix(ui): survive first load against a cold backend without a blank page (#11246)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The web UI holds live websocket connections for run events, coordinates cross-tab polling through a leader-election store, and renders app chrome (sidebar, providers) around a routed outlet > - When the backend is still cold-starting (managed hosting wake, server restart, reverse proxy up before the app), the event websockets refuse connections and the first SPA load mounts against a dead backend > - In that state the mount cascade can exceed React's nested update limit (minified error #185); the crash originates in shell hooks outside the routed error boundary, so React unmounts the entire root to a blank page, and the dead page keeps retrying the websocket on a flat 1.5s timer until the user hard-refreshes > - This pull request removes the wasted nested commits from the shared-polling subscription path, adds exponential backoff to the transcript websocket reconnect, and adds a last-resort app-shell error boundary > - The benefit is that a cold or briefly unreachable backend degrades to a recoverable state instead of a blank page that hammers the server ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** On the first load against a backend that was still starting, the app showed its loading animation and then a blank page. The console showed repeated `WebSocket connection to 'wss://…/api/companies/<id>/events/ws' failed` lines and `Uncaught Error: Minified React error #185` with a stack through the shared-polling coordinator's `subscribe`. The websocket retries continued indefinitely on the dead page. A manual refresh fixed it. **Expected behavior** A backend that is briefly unreachable degrades gracefully: websocket reconnects back off, the UI keeps rendering from cache, and even a worst-case crash shows a reload prompt instead of a blank page. **Steps to reproduce** 1. Serve the UI while the backend API is still starting (websocket upgrades and API calls refused). 2. Load any company page with several shared-polling consumers mounted (dashboard with sidebar). 3. Observe repeated websocket failures; on affected loads the page goes blank with React error #185. ## What Changed - `ui/src/hooks/useSharedPolling.ts`: coordinator snapshot notifications now keep the previous state object when leadership did not change, so React bails out instead of scheduling a nested re-render. `subscribe` invokes its listener synchronously from inside the mount effect with a fresh object each time; before this change every mount and notify burned nested-update budget even with no value change — the crash frame in the field report was exactly this `subscribe → setState` call. - `ui/src/components/transcript/useLiveRunTranscripts.ts`: the live event websocket reconnect backs off exponentially (1.5s → 15s cap, reset on successful open), mirroring `LiveUpdatesProvider`, instead of a flat 1.5s retry. - `ui/src/components/AppErrorBoundary.tsx` (+ wiring in `ui/src/main.tsx`): a dependency-free boundary above the router and providers. `RouteErrorBoundary` only guards the routed `<Outlet />`; a crash in the shell around it had no boundary, so React unmounted the root to a blank page. The boundary renders a reload prompt with the error message. - Tests: `useSharedPollingSnapshot.test.tsx` (mount costs no extra commit — fails against the previous code; a real leadership change re-renders exactly once and ticks stay quiet), a backoff test in `useLiveRunTranscripts.test.tsx` (delays grow 1.5s → 3s → 6s and reset after a successful open), and `AppErrorBoundary.test.tsx` (render throw, effect throw, healthy pass-through). ## Verification - `pnpm vitest run` in `ui/` over the touched suites (shared polling, cross-tab poll, transcripts, boundary): 34 tests pass. - `pnpm typecheck` in `ui/` — clean. - The snapshot regression test was verified to fail against the pre-change hook (extra commit per mount). - Not reproduced end-to-end: the exact 50-update cascade from the field crash needs a live cold backend; the change removes the identified per-mount/per-notify nested commits at the reported crash frame, bounds the reconnect load, and guarantees the shell can no longer blank the page. ## Risks Low risk. The snapshot change only suppresses re-renders whose state is value-identical; leadership changes propagate exactly as before. The backoff only lengthens retry delays after consecutive failures and resets on success. The new boundary renders children untouched unless an error reaches it; behavior on healthy loads is unchanged. Self-hosted deployments see the same code paths — the cold-backend window simply rarely occurs there. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution; diagnosis included mapping the production minified stack to source). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9adeb4a9d0 |
fix(ui): fetch the web app manifest with credentials (#11245)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The web UI is a PWA-capable SPA; `ui/index.html` links `/site.webmanifest` so browsers can read app metadata > - Browsers fetch `<link rel="manifest">` in "omit credentials" mode unless the link opts in with `crossorigin="use-credentials"` > - Self-hosted this is harmless, but when Paperclip runs behind an authenticating reverse proxy (a managed hosting front door), the cookie-less manifest request is rejected with 401 on every page load and logs a console error pair on each navigation > - This pull request adds `crossorigin="use-credentials"` to the manifest link so the request carries the same session cookies as every other same-origin asset request > - The benefit is a clean console and a servable manifest in proxied deployments, with self-hosted behavior unchanged ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** On every page load behind an authenticating reverse proxy, the browser logs `Failed to load resource: the server responded with a status of 401` for `/site.webmanifest`, plus `Manifest fetch from … failed, code 401`. The proxy rejects the request because the browser sends the manifest fetch without cookies. **Expected behavior** The manifest request carries the same session credentials as every other same-origin asset request, so the proxy can authenticate and serve it. No console errors. **Steps to reproduce** 1. Serve Paperclip behind a reverse proxy that requires a session cookie for all app routes. 2. Sign in and load any page. 3. Open the browser console: the manifest fetch fails with 401 while all other assets load. ## What Changed - `ui/index.html`: the manifest link now carries `crossorigin="use-credentials"`. - `ui/src/lib/pwa-install-mode.test.ts`: a regression test asserts the attribute stays on the link. ## Verification - `pnpm vitest run src/lib/pwa-install-mode.test.ts` in `ui/` — 2 tests pass. - Manual check of the rendered link tag in `ui/index.html`. ## Risks Low risk. The manifest is same-origin, so `use-credentials` only switches the fetch from "omit" to the include behavior all other same-origin requests already have. Self-hosted deployments see no change. Cross-origin manifest hosting is not used in this project. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d5bb396518 |
fix: pass sandbox provider credential env vars to plugin workers; hide Local default under managed-sandbox-only (#11244)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments give each agent run an execution target, and sandbox providers (Daytona, E2B, Novita, exe.dev) run as plugin workers > - A managed deployment provisions one platform-managed sandbox row with no credential in config; the provider is documented to fall back to its process env var (for example `DAYTONA_API_KEY`) > - Plugin workers spawn with a scrubbed environment, so that fallback never sees the host env var — probe and lease acquisition fail with "require an API key in config or DAYTONA_API_KEY" even when the deployment sets the var > - Separately, the managed-sandbox-only mode hides local rows from every list, but the instance Default picker renders a hardcoded synthetic "Local" option that no filter touches > - This pull request forwards each bundled provider's documented credential env var to its own plugin worker, and gates the synthetic Local option on the flag > - The benefit is that the documented host-env credential fallback works for plugin-backed providers, and managed-sandbox-only instances no longer offer Local anywhere ## Linked Issues or Issue Description **Subsystem affected** Plugin worker environment construction (`server/src/services/plugin-loader.ts`) and the environments UI (instance Default picker, agent form inherited-environment label). **Problem or motivation** Two follow-ups to the managed-sandbox-only mode (#11200), both found on a live managed deployment: 1. The deployment sets `DAYTONA_API_KEY` as a server env var and the managed sandbox row omits `config.apiKey` by contract. "Test Connection" fails with `Sandbox environment probe failed for provider "daytona". Daytona sandbox environments require an API key in config or DAYTONA_API_KEY.` A real agent run fails the same way at lease acquisition. The cause: sandbox providers run as plugin workers, and `buildPluginWorkerEnv` passes only model-provider keys and in-cluster Kubernetes vars. The provider's own documented credential env var never reaches the worker, so the in-plugin `process.env` fallback reads nothing. The self-hosted path has the same gap: the Daytona plugin README documents `DAYTONA_API_KEY` as a host-level fallback, and it does not work today. 2. With `enableManagedSandboxOnly` on, the instance Default environment picker still shows "Local". The server filters local *rows* out of the list, and the client filter mirrors that for cached lists, but this option is a hardcoded `<option value="">Local</option>` — not a list row — so no filter removes it. Selecting it writes a null default, which run selection then rejects fail-closed. **Proposed solution** Forward each bundled sandbox provider's documented credential env var into its plugin worker, keyed by the manifest's declared `environmentDrivers[].driverKey` so a worker only receives its own provider's credential (daytona → `DAYTONA_API_KEY`, e2b → `E2B_API_KEY`, exe-dev → `EXE_API_KEY`, novita → `NOVITA_API_KEY`). Keep the existing gate: only plugins that declare `environment.drivers.register` receive any passthrough. In the UI, render the synthetic Local option only when managed-sandbox-only is off; under the flag show a disabled "Select environment" placeholder only while no default is stamped yet, and stop the agent form's inherited label from reading "Local". **Alternatives considered** Adding `DAYTONA_API_KEY` to the existing `ADAPTER_ENV_PASSTHROUGH` list was rejected: that list goes to every environment-driver plugin, so each provider would receive every other provider's credential. A manifest schema field for declared credential env vars was rejected as heavier than needed: the bundled providers are known, and the mapping lives next to the two existing passthrough lists. ## What Changed - `server/src/services/plugin-loader.ts`: new `SANDBOX_PROVIDER_CREDENTIAL_ENV_PASSTHROUGH` map (driverKey → documented credential env vars). `buildPluginWorkerEnv` reads the manifest's `environmentDrivers` and forwards only the matching vars, after the existing `environment.drivers.register` gate. Blank values stay excluded. - `server/src/__tests__/plugin-database.test.ts`: the daytona worker receives `DAYTONA_API_KEY` and not another provider's key; a plugin whose drivers have no mapping (kubernetes) receives no credential var. - `ui/src/pages/CompanyEnvironments.tsx`: the Default picker's synthetic Local option renders only when managed-sandbox-only is off. Under the flag, a disabled "Select environment" placeholder renders only while the default is unset. - `ui/src/pages/CompanyEnvironments.test.tsx`: the Local option is present by default and absent under the flag; saved non-local environments stay selectable. - `ui/src/components/AgentConfigForm.tsx`: the inherited-environment label falls back to "Managed sandbox" instead of "Local" under the flag. ## Verification - `server`: `npx vitest run src/__tests__/plugin-database.test.ts -t buildPluginWorkerEnv` — 5 passed (3 existing, 2 new). - `ui`: `npx vitest run src/pages/CompanyEnvironments.test.tsx` — 22 passed (2 new); `npx vitest run src/components/AgentConfigForm.render.test.tsx` — 10 passed. - `tsc --noEmit` clean in `server` and `ui`. - Live managed deployment: confirmed the tenant service env carries `DAYTONA_API_KEY` while the probe fails with the exact message above, which pins the root cause to the worker env, not delivery. ## Risks - The worker env grows by exactly one var per matching bundled provider, only when the deployment sets it and only for plugins that declare a matching environment driver. Plugins without a mapping see no change. - Self-hosted behavioral shift is the fix itself: a host-level `DAYTONA_API_KEY` (or E2B/EXE/NOVITA equivalent) now reaches the provider as its README documents. Deployments that set the var but expected it to stay inert had no working configuration to preserve — the provider errored on every keyless probe and run. - UI change is inert unless `enableManagedSandboxOnly` is on (default false everywhere). ## Model Used Claude Fable 5 (`claude-fable-5`) via Claude Code — extended thinking, tool use, parallel read-only subagents for the two root-cause traces. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0aa743fc30 |
build(db): clean dist before drizzle generate (#11241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Database migrations are generated by drizzle-kit, which reads the schema from the db package's built `dist/schema/*.js` > - `tsc` never deletes stale outputs, so a long-lived checkout keeps compiled schema files whose sources were deleted long ago > - A `generate` run in such a checkout sees those ghost tables and sweeps phantom `CREATE TABLE` statements into an unrelated migration > - This pull request makes `generate` clean `dist` before building, so drizzle always diffs against exactly the current schema sources > - The benefit is that no contributor can accidentally resurrect deleted tables inside a new migration ## Linked Issues or Issue Description **What happened?** Running `pnpm --filter @paperclipai/db generate` in a months-old checkout produced a migration re-creating `cloud_upstream_connections`, `cloud_upstream_runs`, and `company_secret_pools` — tables whose schema sources were deleted in #10507. The compiled copies were still in `dist/schema/`, and `drizzle.config.ts` reads the schema from `dist`, so drizzle treated them as new tables missing from the snapshot. **Expected behavior** `generate` diffs the current schema sources only; deleted tables can never reappear in a generated migration. **Steps to reproduce** 1. Build the db package, then delete a schema source file without cleaning `dist`. 2. Run `pnpm --filter @paperclipai/db generate`. 3. The generated migration re-creates the deleted table. ## What Changed - `packages/db/package.json`: the `generate` script runs `pnpm run clean` before `tsc`, so the drizzle-kit input is always a fresh build of the current sources. ## Verification - In a checkout carrying stale `dist/schema/cloud_upstreams.js` / `company_secret_pools.js` artifacts, `generate` produced a phantom migration before this change; after a clean build it reports "No schema changes, nothing to migrate". With this change the clean happens inside `generate` itself. ## Risks - Low risk: `generate` is a developer-only script; the change only adds the existing `clean` step ahead of the existing build, at the cost of a full rebuild per generate. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0a95ada1be |
feat(server): chunked import preview endpoint and resumable upload in the Import page and CLI (#11224)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The previous pull request added server-side chunked resumable import transfers; without clients, large imports still ride the single fragile upload > - The Import page and the CLI need to slice large packages, upload parts with retry and progress, resume after interruptions, and preview before applying > - Preview is the missing server piece: the browser flow is preview-then-import, so a completed spool must be previewable without re-uploading > - This pull request adds the transfer preview endpoint, switches the Import page to the chunked path for zips over 48 MB, and teaches the CLI the same for oversized local imports > - The benefit is that large imports get progress, per-part retry, and resume in both clients, while small imports keep the exact single-shot path they have today ## Linked Issues or Issue Description **What happened?** With only the server transfer routes in place, users still upload large company packages as one request from the Import page and the CLI: no progress indication, no retry below the whole file, and no resume after a dropped connection or refresh. The preview-then-import flow also cannot run against an uploaded transfer, forcing a second full upload. **Expected behavior** A large package uploads once as verified parts with visible progress; preview and import both run against the uploaded spool; an interrupted upload resumes with only the missing parts re-sent; packages at or below 48 MB behave exactly as before. **Steps to reproduce** 1. Select a 500 MB zip on the Import page over an unreliable connection. 2. Watch the single upload fail near the end and restart from zero, twice — once for preview, once for import. 3. Same story headless via the CLI. ## What Changed - Server: `POST /import/transfers/:id/preview` runs the existing preview logic against the completed spool (shared assembly + whole-file verification helper with apply); preview neither completes the run nor deletes the spool, so the subsequent apply reuses it. Missing parts respond with the missing list. - UI: zips over 48 MB take the chunked path in both preview and import — the file is sliced into 32 MB parts hashed with WebCrypto (single ArrayBuffer, no second copy), the transfer is created or resumed (the create response's missing-parts list drives what uploads), parts upload sequentially with three attempts each and visible progress, then transfer preview/apply replace the multipart calls. The existing preview pane, collision handling, adapter overrides, and async job polling are unchanged; ≤ 48 MB keeps the single-shot path. - CLI: oversized local `.zip` or folder imports zip/slice/hash with node crypto, upload with resume and per-part retry and progress lines, and use transfer preview/apply. Small packages keep the inline path byte-identical. - Failure honesty: adapters/API errors fail open to existing behavior; a part failing all attempts surfaces a durable error panel with resume intact. ## Verification - Server: preview-then-apply on one spool (run stays open, spool intact, then apply completes), preview with missing parts rejected — added to the transfer route suite (embedded Postgres). - UI suite: large file takes the chunked path (manifest shape, part uploads, progress, apply on a resumed transfer, single-shot endpoints never called), small file stays single-shot, part failure after three attempts surfaces the error panel without running preview, resume re-uploads only the missing part. - CLI: manifest slicing/hashing, threshold behavior for zip and folder sources, folder-zip round-trip through the real zip reader, upload resume/retry/exhaustion/already-completed, full-command chunked and small-zip inline flows. - Server, ui, cli typechecks clean. Exact counts in the PR checks. ## Risks - The 48 MB threshold only routes between two verified paths; behavior below it is untouched. - Chunked CLI imports use the board-scoped transfer routes, so oversized CLI imports need board credentials (agent tokens keep the agent-safe small-file path). No privilege change — board actors already had the generic routes — but the two size regimes differ semantically; called out for review. - A CLI dry-run over the threshold uploads parts before previewing; the spool persists (24 h sweep) and a later apply resumes without re-upload — inherent to preview-against-spool. Stacked on #11223 — merge that first; this PR then shows only the preview endpoint and client changes. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
23a1b025c2 |
feat(server): chunked resumable company import transfers (#11223)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import moves large packages into an instance, and since the upload cap rose to 1 GB, the transport is the weak point: one HTTP request, buffered fully in memory, with no resume > - A dropped connection at 90% of an 800 MB upload starts the whole transfer over, and a server restart loses all progress > - This pull request adds the server side of chunked resumable import transfers: a durable run ledger and routes that accept the same import zip as verified ~32 MB parts spooled to disk > - An interrupted transfer resumes from the parts already uploaded — across dropped connections, page refreshes, and server restarts — and peak upload memory drops from the whole package to one part > - The benefit is that large imports become reliable on real-world connections instead of all-or-nothing ## Linked Issues or Issue Description **What happened?** Large company imports travel as a single HTTP upload. On a slow or flaky connection, any interruption discards all progress and the upload restarts from zero. The server buffers the entire compressed package in memory during upload. A server restart mid-upload loses the transfer entirely. With the upload cap now at 1 GB, these failure modes govern exactly the imports the cap was raised for. **Expected behavior** A large import upload survives interruptions: already-transferred data is kept and verified, only the missing remainder is re-sent, and the server's memory use during upload is bounded by a part, not the package. **Steps to reproduce** 1. Import a multi-hundred-MB company package over a connection that drops mid-upload. 2. The upload fails; retrying starts from byte zero. 3. Repeat on an unstable connection and the import may never complete. ## What Changed - New `company_transfer_runs` table (drizzle schema + migration) and `companyTransferRunService`: one row per transfer with a content-derived idempotency key, per-part completion recorded atomically and idempotently, resume scoped to actor and direction, completed runs short-circuiting retries of identical content. - New transfer routes beside the existing import routes, same authorization: declare a sliced zip (`POST /import/transfers` — validates cap, 64 MB part ceiling, contiguity, size sums, sha256 format), upload parts (`PUT .../parts/:n` — raw body, hash-and-size verified before an atomic write to a disk spool under the instance root; re-uploads are no-op successes), poll resume state (`GET .../:id` — missing parts recomputed from disk), and apply (`POST .../:id/apply` — requires all parts, re-verifies the assembled zip against the whole-file hash fail-closed, then feeds the existing import pipeline through factored helpers rather than duplicated logic). - Hourly sweep fails and cleans spools idle for 24 h; a swept transfer honestly reports all parts missing on resume. - Strict UUID gating on run ids before any filesystem path construction. - The existing single-shot upload path is untouched; clients arrive in the follow-up PR. ## Verification - Transfer route suite (embedded Postgres): create/upload/status/apply round-trip with a real imported company, out-of-order parts, wrong-hash part rejected and unrecorded, re-upload no-op, apply-with-missing-parts rejection, resume after failure with prior progress intact, assembled-hash mismatch failing closed with spool deletion, actor scoping 404s, async-job apply, sweep followed by honest resume. - Ledger suite (embedded Postgres): part idempotency, actor/direction scoping, completed-run short-circuit, cancelled runs staying cancelled. - Existing portability route suite unchanged and green; server + db typechecks clean. Exact counts in the PR checks. ## Risks - New routes are additive; the existing import path is untouched. The transfer routes carry the same board authorization as the import routes they sit beside. - Disk spool: bounded by the existing upload cap per transfer, cleaned on success, failure, hash mismatch, and by the 24 h sweep. Spool paths are strict-UUID-gated. - The apply step still materializes the assembled zip in memory once (same profile as today's single-shot import at apply time); upload-time memory drops to one part. - Known limitation, deliberate: transfers are keyed on content alone, so identical package content cannot be imported twice without re-exporting (surfaced explicitly to the caller). Acceptable for v1; noted for review. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0044fa8904 |
Let tenants edit env vars on managed sandbox environments; add managed-sandbox-only mode (#11200)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments give each agent run an execution target: the local host, SSH, or a sandbox provider > - A managed deployment can provision one platform-managed sandbox environment through the `PAPERCLIP_MANAGED_CONFIG` `environments` section > - That row is fully locked today. A tenant cannot add environment variables for their agents. There is also no way to hide local execution — run selection falls back to the local row > - A platform that manages the sandbox for its tenants needs both: the tenant adds env vars (and nothing else), and local execution is neither visible nor reachable > - This pull request opens exactly one tenant edit (env vars) on the managed sandbox row, and adds an `enableManagedSandboxOnly` mode that hides local and makes run selection fail closed > - The benefit is a complete managed-sandbox experience with no change for self-hosted instances ## Linked Issues or Issue Description **Subsystem affected** Environments (managed sandbox provisioning, environment routes, run environment selection) and the environments UI. **Problem or motivation** Platform-provisioned sandbox environments (`metadata.managedByPaperclip`) reject every write on cloud-managed instances. Agents often need environment variables inside their sandbox. The tenant has no way to set them on the managed row. Separately, an operator cannot remove local execution: the environment list always shows the local row, and run selection falls back to it when no default is set. **Proposed solution** Allow an envVars-only PATCH on the managed sandbox row, and echo those env vars back for editing. Add a managed-tier feature (`enableManagedSandboxOnly`) that hides the local environment from all read surfaces and redirects local-landing run selection to the managed sandbox environment, failing closed when it is unavailable. **Alternatives considered** UI-only hiding of the local row. This was rejected: it does not stop a run from resolving to local, so it is presentation without enforcement. Full unlock of the managed row was also rejected: name, driver, and config stay platform-owned so boot reconciliation cannot fight tenant edits. ## What Changed - `server/src/routes/environments.ts`: the platform-provisioned write floor admits an envVars-only PATCH on the generalized managed sandbox row (sandbox driver, `managedByPaperclip`, not legacy kubernetes-marker rows). Name, driver, config, status, metadata, and DELETE stay rejected. The read floor stops blanking env vars on that row; credential-shaped config keys stay redacted for every actor. Legacy kubernetes-marker rows keep the full floor. - Same file: under `enableManagedSandboxOnly`, the environments list and the by-id read omit the local row for every actor, including instance admins. - `server/src/services/execution-workspace-policy.ts`: `resolveExecutionWorkspaceEnvironmentId` gains the managed-sandbox-only inputs. A selection that lands on the local environment is redirected to the managed sandbox environment. With no active managed row it throws `ManagedSandboxUnavailableError` — never local. Non-local selections (ssh, user-created sandboxes) are untouched. - `server/src/services/heartbeat.ts`: the run path reads the flag, looks up the managed row (`findManagedSandboxEnvironment`, new read-only finder in `environments.ts`), and passes both to the resolver. Mirrors the forced-kubernetes precedent, which keeps precedence when both regimes are on. - `server/src/services/managed-environments.ts`: after a successful reconcile, the instance default environment moves to the managed sandbox row when the current default is unset, local, or dangling. A tenant-chosen custom environment is never overridden. - `packages/shared`: new `enableManagedSandboxOnly` key (schema default false, catalog tier `managed`, cloudDefault false, selfHostedDefault false) and the matching interface field. - UI: managed rows show a "Managed by Paperclip" lock badge; editing one opens a dedicated env-vars-only editor that sends the one PATCH shape the server admits (the old full form failed with a 403 on save). New `ui/src/lib/managed-sandbox-environment.ts` mirrors the local filter for cached lists (applied in the project picker; the agent picker already excluded local). The experimental settings page gains the toggle at its alphabetical card position. `environmentsApi.update` now declares the `envVars` field it already sent. - Tests: environment route floor coverage (envVars-only accepted, mixed bodies rejected, legacy rows still blanked and locked, local hidden and 404 under the flag, self-hosted unchanged), an embedded-postgres service test pinning that boot reconciliation never touches tenant env vars, resolver redirect/fail-closed cases, managed-environments default-stamping cases, and UI lib/settings tests. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/environment-routes.test.ts src/__tests__/environment-service.test.ts src/__tests__/execution-workspace-policy.test.ts src/services/managed-environments.test.ts` — all pass. - `pnpm --filter @paperclipai/shared exec vitest run` — 425 pass (catalog/schema default parity is pinned by an existing test). - `pnpm --filter @paperclipai/ui exec tsc --noEmit` and the affected UI suites (CompanyEnvironments, InstanceExperimentalSettings incl. card-order test, new lib test) — all pass. - Full workspace `pnpm test`: 3,414 passed. 17 files report failures on this machine; the identical 17 fail on a clean `origin/master` worktree in the same environment (git-worktree/skills/embedded-postgres environment dependencies and plugin-SDK zero-test collections). One additional file (`issue-monitor-scheduler.test.ts`) failed one timing-sensitive test in one of two full-suite runs and passes 7/7 in isolation on this branch — a flake in a domain this diff does not touch. The branch introduces no new failures. - Self-hosted zero-delta: every new behavior is gated on the cloud-managed instance check or the new flag, which defaults to false in schema and catalog; pinned by the "does not floor platform-marked rows on self-hosted instances" and flag-off tests. ## Risks - Behavior is opt-in twice over: the write-floor exception applies only to rows the managed-config provisioner stamps, and the hiding/forcing applies only when `enableManagedSandboxOnly` is on (default false everywhere). Self-hosted instances see no change. - The env-vars echo is scoped to the generalized managed sandbox row; legacy kubernetes-marker rows keep the blanket floor because pre-generalization builds may have written platform values there. - Fail-closed run selection means a managed instance with the flag on and an archived managed row (provider plugin down) refuses runs with a precise error instead of running locally. That is the intended posture. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
71e9d6bb0b |
feat(release): candidate-branch beta builds and the release checklist (#11209)
> Follow-up to #11208 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channels promote artifacts along canary → nightly → beta → stable, with the happy path being promotion of an existing build > - When one or two targeted fixes are needed before a beta or stable, the only options today are waiting for the next nightly or absorbing a whole day of unrelated master changes > - The channel model was designed with an escape hatch for exactly this: short-lived candidate branches carrying only cherry-picked fixes > - This pull request implements candidate-branch beta builds with full verification, documents the stable fix path through the soak-justification gate, and adds the release captain's checklist > - The benefit is that a surgical fix can ship forward without either delay or blast radius, with its provenance recorded ## Linked Issues or Issue Description Refs #11008 — completes the fix-path half of the channel model introduced there. **Subsystem affected** Release automation: `scripts/release.sh`, `.github/workflows/release.yml`, `doc/RELEASING.md`, new `doc/RELEASE-CHECKLIST.md`, tests. **Problem or motivation** Beta promotion only accepts commits that already shipped as a nightly, and stable promotion expects a soaked beta. There is no supported way to ship one or two cherry-picked fixes between lanes: an urgent fix must wait for the nightly cycle or pull in every unrelated master change from the day. The original channel design called for candidate branches to cover this, and they were deferred from the initial implementation. **Proposed solution** Candidate-branch beta builds: cut `candidate/beta-<target>` from a nightly's source commit, cherry-pick the fixes, and dispatch `channel: beta` with the new `candidate_branch` input. Selection enforces the naming convention, rejects heads that already shipped as a beta or predate the candidate tooling, and records the cherry-picked commits in the job summary. Because candidate heads never went through a canary or nightly, publication is gated on a full `release-verify` run (promoted nightlies keep skipping re-verification). The stable fix path (`candidate/release-<target>` as `source_ref`) works through the existing soak gate: the justification requirement is the deliberate, recorded trade-off for shipping unsoaked bits, and is now documented as such. ## What Changed - `scripts/release.sh`: `--from-candidate` flag (beta only) waives the shipped-a-nightly requirement while keeping the duplicate-beta guard - `.github/workflows/release.yml`: `candidate_branch` dispatch input; candidate mode in `select_beta` (naming validation, duplicate and tooling-era rejection, cherry-pick recording); new `verify_beta_candidate` job gating candidate publishes on full verification - `doc/RELEASING.md`: beta fix-path and stable fix-path sections - `doc/RELEASE-CHECKLIST.md` (new): the release captain's checklist for all four lanes as built - Tests: dry-run fixture coverage for `--from-candidate` (waives the nightly guard, keeps the duplicate guard, rejected outside beta) and wiring tests for candidate validation plus the verification gate ## Verification - `node --test` on the four affected suites: 42 pass in total (17 + 25 across the two runs), including the 5 new tests - `bash -n` on `release.sh`; YAML parse of the workflow - After merge: exercise the path end to end the first time a real cherry-picked beta is needed — dispatch with a `candidate/beta-*` branch and confirm the summary records the picks and verification runs ## Risks - Candidate builds bypass the smoke-tested-nightly provenance by design; the compensating controls are full verification before publish, the post-publish beta smoke, the human `npm-beta` gate, and recorded cherry-picks - The stable fix path rides the existing justification mechanism rather than adding a second bypass — one recorded escape hatch, not two ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2da6a248c3 |
fix(release): surface recovery commands when a lane tag push is rejected (#11208)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's promotion lanes publish to npm, then push a lane tag and dispatch the Docker image build at that tag > - The first nightly of the beta-tooling merge published to npm and then died at the tag push: GITHUB_TOKEN may not create refs pointing at workflow-modifying commits from dispatch or scheduled runs > - The failure was a bare `remote rejected` with no guidance, leaving the release half-finished (npm live, no tag, no images) until an operator reverse-engineered the recovery > - This pull request makes every lane's tag push degrade into exact recovery instructions in the job summary > - The benefit is that a rare platform-permission rejection becomes a two-minute runbook operation instead of a forensic exercise ## Linked Issues or Issue Description Refs #11008 — the incident occurred promoting that change's own merge commit, the first workflow-modifying commit to flow through the lanes it introduced. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31445344811 published `2026.811.0-nightly.0` to npm, then failed pushing `nightly/v2026.811.0-nightly.0`: `refusing to allow a GitHub App to create or update workflow .github/workflows/release.yml without workflows permission`. The tagged commit modifies workflow files, and GITHUB_TOKEN may not create refs pointing at such commits from dispatch or scheduled runs (push-event runs are exempt, which is why the canary tag on the same commit succeeded). The job failed with no explanation and the Docker dispatch never ran. **Proposed solution** Wrap the nightly, beta, and stable tag pushes: on rejection, write the exact recovery commands into the job summary — create and push the tag with maintainer credentials, dispatch `docker.yml` at the tag, and for stable also run `create-github-release.sh` — then fail the job. Document the cause and recovery in the failure playbooks and pin the three recovery blocks with a wiring test. ## What Changed - `.github/workflows/release.yml`: recovery-summary wrappers on the nightly, beta, and stable tag-push steps - `doc/RELEASING.md`: failure-playbook entry for the workflows-permission rejection - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting all three lanes carry the recovery summary ## Verification - Wiring tests: 6 pass; YAML parse of the workflow - The recovery commands are exactly the ones used to resolve the real incident (tag push + `docker.yml` dispatch for `nightly/v2026.811.0-nightly.0`) ## Risks - Low. The happy path is unchanged (a successful push skips the wrapper); the failure path trades a bare error for actionable output and still fails the job, since the release state is genuinely incomplete ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1ea2f0e2d6 |
feat(cli): add 'paperclipai channels' to show release lanes and the current one (#11210)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channels (canary → nightly → beta → stable) select by install target, and users discover them today only through maintainer-oriented docs > - A user who wants to know "which lane am I on, and what else is there" has no self-serve answer > - The channel rollout planned a read-only CLI command for exactly this > - This pull request adds `paperclipai channels`: every lane with the version its dist-tag resolves to, the install command for it, and which lane the running install follows > - The benefit is self-serve lane discovery without reading release documentation ## Linked Issues or Issue Description Refs #11008 — the user-facing discovery surface for the channel model completed there. **Subsystem affected** CLI: `cli/src/commands/channels.ts` (new), `cli/src/index.ts`, `doc/CHANNELS.md`, tests. **Problem or motivation** Channel selection is install-based (`@latest` / `@beta` / `@nightly` / `@canary`), but nothing in the product tells a user which channel their install follows or what the other lanes currently resolve to. The information lives in `doc/CHANNELS.md` and the npm registry, neither of which a running install surfaces. **Proposed solution** A read-only `paperclipai channels` command: prints each channel with the version its dist-tag currently resolves to (per-lane registry lookups that degrade to `unavailable` individually), the install command for each, and the running install's lane parsed from its version suffix — source checkouts carry the repository's placeholder version and are reported as unmapped rather than guessed. `--json` emits the same data for scripting. ## What Changed - `cli/src/commands/channels.ts` (new): channel table, lane parsing, registry resolution, human and `--json` output - `cli/src/index.ts`: registers `channels` - `doc/CHANNELS.md`: "Seeing where you are" section - `cli/src/__tests__/channels.test.ts` (new): lane parsing including unknown versions, full resolution against a fake runner, per-lane degradation, table/dist-tag sync ## Verification - `vitest run cli/src/__tests__/channels.test.ts`: 5 pass - `pnpm typecheck` in `cli/` - Live run against the real registry shows all four lanes with their current versions (`2026.722.0` / `2026.811.0-beta.0` / `2026.811.0-nightly.0` / canary) and correctly reports a source checkout as unmapped ## Risks - Low. Read-only command reusing the existing `resolvePublishedVersion` registry helper; no state, no auth, no publish surface ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d648becb90 |
refactor(ci): split workspaces-a into two Vitest native shards
Split the slow workspaces-a CI lane into two Vitest native shards and keep release verification in parity. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6601014898 |
fix(release): reject promotion sources that predate their channel tooling (#11197)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem promotes builds along canary → nightly → beta → stable, and each publish job checks out the promotion's source commit and runs that tree's release tooling > - The first beta dispatch failed with `unexpected argument: beta`: the selected nightly's source predated the beta channel, so its `release.sh` did not know the argument > - The failure was clean (argument parsing, nothing published) but cryptic, and the same trap waits for any promotion of a source older than its target channel's tooling > - This pull request makes the selection jobs reject such sources with an actionable error and documents the property > - The benefit is that a bootstrapping or old-source promotion fails in seconds with instructions, instead of mid-publish with a parser error ## Linked Issues or Issue Description Refs #11008 — the guard hardens the beta promotion flow introduced there, after its first dispatch surfaced the gap described below. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31444045044 (first beta dispatch) failed in `publish_beta` with `unexpected argument: beta`. Promotions deliberately build from the pinned source commit, which means they also run that commit's `scripts/release.sh` — and a source that predates the target channel's introduction cannot publish it. Nothing guards this today; the error surfaces deep in the publish job with no explanation. **Proposed solution** Guard at selection time: `select_nightly` requires the source canary's `release.sh` to know the nightly channel, and `select_beta` requires the source nightly's `release.sh` to know the beta channel. Each guard literally matches the channel case arm and fails closed with a clear message naming the remedy (promote a newer source). Document the tooling-era property in `RELEASING.md` and pin the guards with a wiring test. ## What Changed - `.github/workflows/release.yml`: tooling-era guards in `select_nightly` and `select_beta` - `doc/RELEASING.md`: documents that promotions run the source commit's release tooling - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test pinning both guards ## Verification - Wiring tests: 5 pass - Guard expressions exercised against real commits: accepts the beta-capable merge commit of the beta-channel change, rejects a pre-beta commit - YAML parse of the workflow - After merge: the next beta dispatch selects a beta-capable nightly and passes the guard ## Risks - Low. Selection-time check only; the guards match the channel case arm literally and fail closed (with the same actionable message) if that line is ever reformatted ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
35aaaa0bd0 |
feat(server): preserve task timestamps and hierarchy through company import/export (#11193)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company export/import moves a whole company — agents, tasks, comments — between instances as a portable bundle > - The bundle never carried task timestamps or parent links: the export writes neither, the importer lets database defaults stamp "now", and sub-tasks arrive flattened > - Boards sort by recency, so every imported task showing "created just now" collapses the task list into import order, and the task hierarchy the user built is gone > - This pull request adds created/updated/started/completed/cancelled timestamps and a parent link to the bundle (schema v7), preserves them end to end on import, and keeps comment imports from clobbering a preserved updated time > - The benefit is that an imported company reads like the company the user left: same recency order, same task tree ## Linked Issues or Issue Description **What happened?** After a company import, every task showed as created at import time. Recency sorting collapsed to import order, and parent/child task nesting disappeared. The user called out losing "the meaningful task hierarchy and recency sorting". Cause: the export bundle has no fields for task timestamps or parent links, the importer lets `defaultNow()` win on insert, and the comment importer bumps every touched task's `updatedAt` to now. **Expected behavior** An imported company preserves each task's creation/update/start/completion times and its position in the task tree, so sorting and nesting on the destination match the source. **Steps to reproduce** 1. On a source instance, create tasks over several days, including sub-tasks nested under parents. 2. Export the company and import it into another instance. 3. Every task shows the import moment as its creation/update time and all tasks are top-level. ## What Changed - Export writes `createdAt`/`updatedAt`/`startedAt`/`completedAt`/`cancelledAt` (ISO, only when set) and `parent: <taskSlug>` into each task's bundle extension; a parent outside the export selection drops the edge with an aggregate warning, mirroring the existing blocker-edge warning (`server/src/services/company-portability.ts`). - Bundle schema version 6 → 7. All new fields are optional: v5/v6 bundles import unchanged with a version-aware downlevel warning; bundles newer than the board still fail closed. - Manifest parsing validates the new timestamps like comment timestamps (invalid → warn and ignore, never a hard failure); shared types and the zod validator carry the new optional fields. - Import resolves parent slugs to pre-generated destination ids, drops self-references and cycles from tampered bundles with warnings, and orders rows parents-first because the self-referencing FK is checked per insert chunk. - `importIssues` writes the preserved timestamps (falling back to insert time when absent; `startedAt` stays null unless bundle-carried, per #11191's semantics) and `parentId`. - `addImportedComments` no longer blanket-bumps `updatedAt = now()`; it takes `GREATEST(updated_at, newest imported comment createdAt)`, so a preserved update time never regresses while unpreserved rows keep the old behavior. ## Verification - `pnpm vitest run server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts server/src/__tests__/productivity-review-service.test.ts` — 102 passed, 1 pre-existing opt-in benchmark skip. Includes: full round-trip with exact timestamp equality and a 3-deep parent chain against embedded Postgres; v6 back-compat (defaults + warning); forward-compat rejection (v8); cycle/self-reference/invalid-timestamp tampered-bundle handling; comment-bump preserve-awareness in both directions. - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/shared typecheck` — clean. ## Risks - **Rollout ordering**: a board on the previous build (max schema v6) refuses bundles exported by this build (stamped v7) — the existing newer-than-supported rejection, working as designed. Cross-instance moves need the importing board upgraded first. Called out here so operators aren't surprised during the transition window. - Parent edges from tampered bundles are dropped with warnings rather than failing the import; blocker relations already behave this way. - Timestamps are data-only; no destination schema migration. Stacked on #11191 (its commit is included here) — merge #11191 first; this PR then shows only the v7 changes. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4c062a0eb2 |
fix(ui): stop coercing imported agents to the destination CEO adapter (#11192)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import lets an operator bring a package of agents into an instance, and each agent declares which adapter runs it (Claude Code, Codex, and so on) > - The export and the server importer preserve each agent's adapter faithfully, but the Import page seeds an adapter override for every agent with the destination CEO's adapter before the user touches anything > - Every imported agent therefore arrives as the CEO's adapter (usually Claude Code) even when the source package holds a mix, and the picker shows the coerced value as if it were the source's, so nothing looks wrong > - This pull request makes the manifest adapter the default, sends overrides only for agents the user actually changed, and replaces the silent coercion with an explicit per-agent fallback warning when the destination truly lacks the source adapter > - The benefit is that a mixed Claude/Codex team imports as a mixed Claude/Codex team, and any real adapter gap is visible instead of silent ## Linked Issues or Issue Description **What happened?** A user imported a company package whose agents were a mix of Claude Code and Codex on the source instance. After the import, every agent was configured as Claude Code. The import preview showed no sign that anything had been changed. Cause: the Import page initializes its adapter-override map by assigning every agent the destination CEO's adapter type and sends that override for every agent, overriding the manifest's per-agent adapter server-side. For imports into a new company, the "CEO adapter" is read from whichever unrelated company is currently selected. **Expected behavior** Imported agents keep the adapter declared in the package. An override is sent only when the operator explicitly picks a different adapter, or when the source adapter is not installed on the destination — and in that case the page must say so per agent, not silently substitute. **Steps to reproduce** 1. On a source instance, create a company with one Claude Code agent and one Codex agent, and export it. 2. Import the package on another instance whose CEO uses Claude Code, changing nothing in the import dialog. 3. Both agents arrive configured as Claude Code; the Codex identity is gone. ## What Changed - The preview no longer seeds adapter overrides; the override map starts empty, and the picker displays each agent's manifest adapter (`ui/src/pages/CompanyImport.tsx`). - `buildFinalAdapterOverrides` sends an entry only when the effective adapter differs from the manifest or the agent's adapter config was edited — untouched agents flow through with no override. - The page fetches the destination's installed adapters (existing `adaptersApi.list()` client). When a manifest adapter is missing or disabled on the destination, only that agent defaults to the CEO's adapter, with a visible amber warning naming both adapters. If the adapters request fails, the page fails open: manifest adapters are kept and no coercion happens. - Tests: untouched mixed-adapter import sends no overrides; a user-changed agent sends exactly one; a missing destination adapter produces the fallback plus rendered warning for that agent only; an adapters-endpoint failure produces no coercion. ## Verification - `npx vitest run ui/src/pages/CompanyImport.test.tsx` — 19 passed (15 pre-existing + 4 new). - `pnpm --filter ./ui typecheck` (`tsc -b`) — clean. ## Risks - Behavior change: users who previously relied on the silent conversion (importing packages that reference adapters they don't have) now get an explicit per-agent fallback with a warning — same outcome, visible. The server's hard rejection of unknown adapter types remains the backstop for API callers. - UI-only change; no server or schema impact. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d816eb8095 |
fix(server): keep imported tasks quiescent under the productivity review sweep (#11191)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import brings a full company package — agents, tasks, routines — into an instance, with `pauseAutomations` promising a quiet landing > - The pause covers the imported entities, but the destination's own productivity-review sweep does not know the difference between imported rows and live work > - The importer stamps every imported in-progress task with `startedAt = now()`, so six hours later the sweep's long-active check fires on every one of them and floods the board with review tasks and agent wakeups > - This pull request stops fabricating `startedAt` on import and makes the sweep skip tasks whose assignee agent is paused > - The benefit is that an import lands quietly: no surprise review-task storm, and paused teams stay paused until the operator activates them ## Linked Issues or Issue Description **What happened?** After importing a company package with automations paused, a batch of "productivity review" tasks appeared roughly six hours later — one for every imported in-progress task — each with an owner-agent wakeup. The user described it as jarring and wasteful. Cause: `importIssues` fabricates `startedAt = now()` for imported in-progress rows, and `reconcileProductivityReviews` considers any assigned in-progress task without checking whether the assignee agent is paused, so its long-active-duration evidence (6 h threshold) trips on the fabricated timestamp. **Expected behavior** An import with paused automations must be quiescent: no destination sweep should generate work from imported rows until the operator unpauses the imported team. A paused agent must not accumulate review tasks it cannot act on. **Steps to reproduce** 1. Import a company package containing tasks with status `in_progress` assigned to agents, with "pause automations" enabled. 2. Wait for the productivity-review reconcile (runs at startup and on the heartbeat scheduler tick) more than six hours after the import. 3. Observe one new review task plus an owner wakeup per imported in-progress task. ## What Changed - `importIssues` no longer fabricates `startedAt` for imported `in_progress` rows; it inserts null (`server/src/services/issues.ts`). Audited every consumer of `issues.startedAt` — all are null-tolerant, and normal checkout/status-transition paths set the value when work really starts. - `reconcileProductivityReviews` skips candidates whose assignee agent is `paused`, counting them as skipped (`server/src/services/productivity-review.ts`). This is a general rule, not import-specific: a paused agent cannot act on a review. - Tests: paused-assignee candidate with an old `startedAt` creates no review, and creates one after unpausing; imported in-progress issue lands with null `startedAt` (embedded-Postgres import test); the pre-existing long-active regression test still passes. ## Verification - `pnpm vitest run server/src/__tests__/productivity-review-service.test.ts server/src/__tests__/company-portability-import-batching.test.ts` — 20 passed, 1 pre-existing opt-in benchmark skip. - `pnpm vitest run server/src/__tests__/company-portability.test.ts` — 78 passed. - `pnpm --filter @paperclipai/server typecheck` — clean. ## Risks - Behavior change beyond imports: tasks assigned to paused agents no longer receive productivity reviews anywhere. This is intended — the review would target an agent that cannot respond — and reviews resume on the first reconcile after unpausing. - Imported in-progress tasks now carry no `startedAt` until real work starts on the destination. The one sweep that read the fabricated value is the one this PR quiets; all other consumers fall back safely (audit in the commit body). - Low risk otherwise: no schema change, no API shape change. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8f7b8b3fda |
feat(release): add human-gated beta channel with stable soak enforcement (#11008)
> Follow-up to #11006 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem now publishes canary (every master push), nightly (scheduled, smoke-gated, added in #11006), and stable (manual) > - There is still no human-approved release-candidate lane between nightly and stable, and nothing enforces that a stable actually soaked anywhere before shipping > - Betas need a real approval gate, and stables need a soak policy that is data, not prose > - This pull request adds the beta channel: a manual promotion of a chosen nightly behind the `npm-beta` environment gate, re-smoked after publish, plus a stable preflight that enforces a 3-day beta soak with a written-justification bypass > - The benefit is a complete canary → nightly → beta → stable train where every stable shipped as a beta first, and emergencies leave a written trace ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** After #11006 the project has canary and nightly prerelease lanes, but no release-candidate lane. Stable promotion has no enforced soak: any ref can ship as stable directly. There is no approval boundary for a broader-audience prerelease, and no structured way to record why an emergency release skipped validation. **Proposed solution** Add a `beta` channel: a manual dispatch that promotes a chosen nightly's source commit, publishes behind the `npm-beta` GitHub environment (required reviewers are the gate), re-smokes the published beta, and tags `beta/vX`. Enforce in the stable path that the source commit shipped as a beta at least 3 days earlier (measured from the beta's npm publish time), with a `skip_soak_justification` input as the recorded emergency bypass. **Alternatives considered** Codifying the soak policy in docs only. Rejected: an unenforced policy decays; the preflight makes the policy executable while the justification input keeps the emergency path usable and auditable. ## What Changed - `scripts/release.sh` + `scripts/release-lib.sh`: `beta` channel — requires HEAD to carry a `nightly/v*` tag, publishes the package set as `YYYY.MDD.P-beta.N` under dist-tag `beta`, tags `beta/vYYYY.MDD.P-beta.N` - `.github/workflows/release.yml`: - `channel: beta` dispatch path: `select_beta` resolves the newest (or an explicit `source_version`) nightly and fails loudly on selection problems; `publish_beta` runs behind the `npm-beta` environment, pushes the tag, and dispatches `docker.yml`; `smoke_beta` re-runs the release smoke suite against the exact published beta version - stable path: new `preflight_stable` job enforces the 3-day beta soak from the beta's npm publish time; `skip_soak_justification` bypasses with the reason echoed into the job summary; dry runs report without blocking - `.github/workflows/docker.yml`: `beta/v*` tags publish `:beta` on both images, with exact version stamping - `.github/workflows/release-smoke.yml`: `beta` added to the dispatch choice list - Docs: `CHANNELS.md` beta entries; `RELEASING.md` beta lane, soak gate, and failure playbook; `RELEASE-AUTOMATION-SETUP.md` `npm-beta` environment setup, including the warning to create the environment before the first beta dispatch (GitHub auto-creates unprotected environments on first reference) - Tests: beta version-counting coverage in `scripts/release-registry-versions.test.mjs`; beta identity and nightly-tag guard coverage in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the two touched suites: 17 pass, including the 3 new beta tests - `bash -n` on both shell scripts and YAML parse of all three workflows - After merge, in order: create the `npm-beta` environment, dispatch `channel: beta` with `dry_run: true` to preview, then a real promotion of a published nightly through the approval gate, then a stable dry-run against a young beta to see the soak gate report ## Risks - If the `npm-beta` environment does not exist when the first beta dispatch runs, GitHub creates it with no protection rules and the beta publishes without approval. Mitigated by documentation and by creating the environment before merge (operator step) - Until the first beta exists, every stable dispatch requires `skip_soak_justification`. This is deliberate — the first beta ships immediately after this merges — but it is a behavior change to the stable dispatch - The soak clock reads the beta's npm publish time from the registry; a registry outage makes the preflight fall back to requiring justification (fail-closed) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5a0985f80a |
test(release-smoke): update onboarding spec for the mission-first wizard (#11190)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane gates every nightly publish on the release smoke suite, which drives real onboarding in a browser against the published artifact > - With the harness fixed (#11187, #11189), the gate reached the Playwright suite for the first time in CI — and the spec still walks the old onboarding wizard, so it fails at "Create your first agent" on every current build > - The wizard was redesigned to a mission-first five-step flow, and the spec rotted silently because the suite never ran in CI before > - This pull request rewrites the spec to drive the current wizard end to end > - The benefit is a smoke gate that actually tests today's product, verified against a real published canary ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `tests/release-smoke/docker-auth-onboarding.spec.ts`. **Problem or motivation** Nightly run 31431273139 failed in the smoke Playwright suite: the spec expects the old wizard step "Create your first agent", but current builds show the redesigned mission-first flow (front door → company → mission → team lead → connect model → review). The page snapshot in the run artifact shows the "Define your mission" step where the spec expected the agent step. Both retries failed identically — this is deterministic spec drift, not flake. **Proposed solution** Rewrite the spec for the current flow: fill the company name, define the mission directly (confirming creates the company), name the team lead, hire it through the adapter step — the adapter environment probe reports unhealthy in the CLI-less smoke container by design and must not block the hire — then launch to the dashboard. Assert the company, the ceo-role agent, and the company goal through the API. The first-task and assignment-run assertions are removed together with the wizard flow that created them. ## What Changed - `tests/release-smoke/docker-auth-onboarding.spec.ts`: rewritten for the mission-first wizard; sign-in and wizard-opening helpers and the company-name step are unchanged ## Verification - Full local run against the real nightly candidate: launched the smoke container for `paperclipai@2026.810.0-canary.3` via `scripts/docker-onboard-smoke.sh` (with the #11189 bind fix), then ran `pnpm run test:release-smoke` against it — 1 passed (4.5s) - After merge: dispatch `release.yml` with `channel: nightly` to run the full gate in CI ## Risks - Low. Test-only change. The spec now asserts less about first-task creation because the wizard no longer creates a first task; if a first-run trigger returns to onboarding, the spec should grow that assertion back ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (CI artifact forensics, UI source tracing, local Docker + Playwright reproduction and verification). All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f94f6003c6 |
fix(release-smoke): pin the smoke container to the lan bind preset (#11189)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane gates every nightly on the release smoke suite, which boots the published artifact in a Docker container and drives real onboarding > - The gate kept failing even after the readiness budget fix (#11187), and the new container-log dump revealed the server was healthy but listening on 127.0.0.1 inside the container, unreachable through Docker's port mapping > - `onboard --yes` without an explicit `--bind` prefers trusted-local quickstart defaults: it writes a loopback bind into the instance config and ignores the deployment env vars the harness passes, and that config outranks `HOST` at runtime > - This pull request pins the smoke container to the `lan` bind preset and adds a wiring test for it > - The benefit is a working nightly gate, verified end to end against a real published canary ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `docker/Dockerfile.onboard-smoke`, `scripts/__tests__/release-verify-workflow.test.mjs`. **Problem or motivation** Nightly run 31428558684 failed in smoke with the server unreachable at the mapped port for the full 420 second budget. The container logs (captured thanks to #11187) show a fully booted server with `Bind loopback (127.0.0.1)`. The harness sets `HOST=0.0.0.0` and the deployment env vars, but `onboard --yes` without `--bind` deliberately prefers trusted-local defaults, writes `bind: loopback` into the instance config, and the config outranks `HOST` at runtime. A loopback listener inside a container is invisible to the port mapping, so the health check can never pass. This behavior predates the current stable, so the harness was silently broken against every recent version — it only surfaced now because the nightly lane is the suite's first CI consumer. **Proposed solution** Pass `--bind lan` in the smoke container command (the flag is supported by `latest` and canary alike; it selects the all-interfaces preset and keeps the env-driven authenticated deployment), and pin the flag with a wiring test so it cannot regress silently. ## What Changed - `docker/Dockerfile.onboard-smoke`: the onboard command is now `onboard --yes --bind lan --data-dir ...`, with a comment explaining why the flag is load-bearing - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting the smoke Dockerfile pins a non-loopback bind preset ## Verification - Full local harness run against the real nightly candidate `2026.810.0-canary.1`: container healthy, bind banner shows `lan (0.0.0.0)`, authenticated bootstrap completed (admin created, bootstrap invite accepted, board session verified), `/api/health` returns `bootstrapStatus: ready` - `node --test scripts/__tests__/release-verify-workflow.test.mjs`: 4 pass - After merge: dispatch `release.yml` with `channel: nightly` to run the gate end to end in CI ## Risks - Low. The change only affects the smoke container. `--bind lan` inside a container exposes the port to the container network only; reachability from outside still goes through Docker's explicit port mapping ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (CI log forensics, upstream source tracing, local Docker reproduction and verification). All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
30f6999cbe |
fix(release-smoke): configurable readiness timeout and diagnostics for slow containers (#11187)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane (#11006) gates every nightly publish on the release smoke suite, which boots the published artifact in a Docker container > - The suite's first CI execution failed at the health readiness check: the harness hard-codes a 90 second budget, but a CI container cold-installs paperclipai from npm and initializes embedded postgres with no warm caches > - When the timeout expired with the container still running, the harness printed no container logs, so the failure gave no diagnostics > - This pull request makes the readiness budget configurable, raises it for CI, and dumps container logs on timeout > - The benefit is that the nightly gate measures the artifact, not the runner's cold caches, and a red smoke run is diagnosable from its logs ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `scripts/docker-onboard-smoke.sh`, `.github/workflows/release-smoke.yml`. **Problem or motivation** Run 31426044332 (first forced nightly after #11006) failed in `smoke_nightly` with `server did not become ready at http://localhost:3232/api/health` after exactly 90 seconds. The harness's readiness window is hard-coded to 90 attempts at 1 second. Locally that works because the npm cache is warm; in CI the container downloads the full package set and embedded postgres first. The timeout path also printed no container logs when the container was still running, so there was no way to see how far boot had progressed. **Proposed solution** Make the readiness budget an environment variable (`SMOKE_READY_TIMEOUT_SECONDS`, default unchanged at 90 for local use), set it to 420 in the CI workflow, and dump the last 150 container log lines when the readiness check times out on a still-running container. ## What Changed - `scripts/docker-onboard-smoke.sh`: `SMOKE_READY_TIMEOUT_SECONDS` env var (default 90) replaces the hard-coded readiness budget; timeout with a still-running container now prints the tail of `docker logs` - `.github/workflows/release-smoke.yml`: sets `SMOKE_READY_TIMEOUT_SECONDS=420` for CI runs ## Verification - `bash -n` on the harness and YAML parse of the workflow - The real proof is the next `channel: nightly` dispatch of `release.yml`, which re-runs this suite in CI with the new budget ## Risks - Low. The local default is unchanged; CI runs simply wait longer before declaring failure, and a genuinely broken artifact still fails (with logs now) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. Diagnosis from CI run logs; patch model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5ca752dc81 |
fix(server): raise company import zip upload limit to 1 GB and make it operator-configurable (#11184)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import/export lets an operator move a full company package between instances, with the Import page uploading the package as one compressed `.zip` > - The server caps that upload at 128 MB, and real company packages with attachments now exceed it — imports fail at the preview step > - The failure message tells the user to use the CLI folder import, but that path posts inline JSON capped at 64 MB, so the advice is a dead end for exactly these packages > - This pull request raises the zip upload cap to a 1 GB default, makes it operator-configurable through an environment variable, scales the decompression-bomb guards from the cap in effect, and replaces the misleading hint > - The benefit is that large real-world company packages import successfully, and operators with unusual needs can tune the cap without a code change ## Linked Issues or Issue Description **What happened?** A company import fails at the preview step with `Preview failed: Import package exceeds 134217728 bytes`. The package is a valid Paperclip export. Its compressed size is larger than the 128 MB server cap (one reported package is 257 MB). The error panel suggests the CLI folder import, but that path sends the package as one inline JSON body capped at 64 MB, so it also fails. **Expected behavior** A valid company package of realistic size imports successfully through the Import page. If a package is too large, the error must state the limit clearly and suggest a step that can work. **Steps to reproduce** 1. Export a company with enough attachments to make the compressed package larger than 128 MB. 2. Open the Import page and upload the `.zip`. 3. Click "Preview import". 4. The preview fails with `Import package exceeds 134217728 bytes`. **Deployment mode** Reported from a managed deployment; the limit applies to all deployment modes. ## What Changed - Raise `PORTABLE_ZIP_UPLOAD_LIMIT_BYTES` from 128 MB to a 1 GB default (`server/src/http/body-limits.ts`). - Add the `PAPERCLIP_IMPORT_ZIP_MAX_BYTES` environment override. Invalid or non-positive values fall back to the default. - Scale the zip decompression-bomb guard from the configured cap at the import route: the aggregate inflated ceiling is 4x the cap. The per-entry ceiling stays at 512 MB because V8's string length limit applies to an entry regardless (`server/src/routes/companies.ts`, `packages/shared/src/portability-zip.ts`). - Report the 422 limit error in MB instead of raw bytes. - Replace the "use the CLI folder import for very large packages" hint on preview failure with advice that works: re-export the package without large attachments (`ui/src/pages/CompanyImport.tsx`). - Update the stale comment in `ui/src/lib/import-preflight.ts` that made the same CLI claim. - Add tests for the new default, the env override, and the invalid-override fallback. ## Verification - `pnpm vitest run server/src/__tests__/body-limits.test.ts packages/shared/src/portability-zip.test.ts server/src/__tests__/company-portability-routes.test.ts server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts` — all pass. - `pnpm vitest run ui/src/pages/CompanyImport.test.tsx` — passes, including the updated failure-panel copy assertion. - `pnpm typecheck` — clean across the workspace. - Manual: upload a `.zip` larger than the configured cap; the preview fails with `Import package exceeds the 1024 MB upload limit` and the new hint. A package between 128 MB and 1 GB now previews and imports. ## Risks - Peak per-import memory rises with the cap: the upload is buffered in memory and unzipped in one pass. A 1 GB compressed package can use several GB transiently. Imports are instance-admin actions, so the exposure is a deliberate operator action, not anonymous traffic. Operators on small hosts can lower the cap with `PAPERCLIP_IMPORT_ZIP_MAX_BYTES`. - The aggregate bomb guard moves from a fixed 512 MB to 4x the configured cap. It still bounds expansion far below what a decompression bomb needs. - No migration and no API shape change. The 422 message text changes; no code matches on the old text. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (file edits, local test runs, live-instance inspection). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f9173782cd |
feat(release): add smoke-gated nightly channel and lane-separated Docker tags (#11006)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem publishes the `paperclipai` npm package set and the Docker images on two lanes: canary on every master push, and stable on manual promotion > - There is no middle ground between those lanes. Users must track every merge or wait weeks for a stable. Docker `:latest` also tracks master, so Docker users have no stable image at all > - A calm prerelease lane needs to exist, and it must never ship a build that failed its checks > - This pull request adds the nightly channel: a scheduled job that selects the newest master commit with a green canary publish, runs the full release smoke suite against that exact published canary, and only then republishes it as the nightly. It also separates Docker tags by lane, so `:latest` finally means stable > - The benefit is that users can follow prereleases at a nightly cadence with a smoke-tested guarantee, and Docker users get real `:canary`, `:nightly`, and stable image tags ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** The project publishes only `canary` (every master push) and `latest` (manual stable). Users who want prereleases without per-merge churn have no option. Docker has a second problem: master builds overwrite `:latest`, and CI-published stables never produced Docker images, because tags pushed with `GITHUB_TOKEN` do not fire the `v*` tag trigger in `docker.yml`. No stable-versioned image exists in ghcr today. **Proposed solution** Add a `nightly` channel. A scheduled job selects the newest canary-tagged master commit, smoke-tests that exact published canary, and republishes the same commit as `YYYY.MDD.P-nightly.N` under the `nightly` dist-tag. Separate Docker tags by lane (`:canary` for master, `:nightly` for nightly tags, `:latest` plus version tags for stable tags only), and have the release jobs dispatch `docker.yml` at the new tag so lane images actually build. **Alternatives considered** Moving the `nightly` dist-tag to the existing canary version without a republish. Rejected: the version string would say `canary` while the user is on nightly, which breaks at-a-glance lane identification in bug reports and `--version` output. ## What Changed - `scripts/release-lib.sh`: channel-parameterized `next_prerelease_version` and `prerelease_tag_name` helpers (canary helpers delegate to them), a `require_channel_tag_at_head` guard, and the no-provenance retry for Sigstore transparency-log duplicates now covers the `nightly` dist-tag as well as `canary` - `scripts/release.sh`: new `nightly` channel. It requires HEAD to carry a `canary/v*` tag, publishes the full public package set as `YYYY.MDD.P-nightly.N` under dist-tag `nightly`, and tags the source commit `nightly/vYYYY.MDD.P-nightly.N` - `.github/workflows/release.yml`: scheduled nightly chain (09:00 UTC) — select candidate, smoke it via `release-smoke.yml`, publish on green under the existing `npm-canary` environment, push the tag, dispatch `docker.yml`. New `channel` dispatch input (default `stable`, so existing stable dispatches are unchanged) with `nightly_source_version` and `dry_run` support for forced runs. The stable path now also dispatches `docker.yml` at the new `v*` tag - `.github/workflows/docker.yml`: lane tag mapping for both image jobs — master pushes publish `:canary` and no longer move `:latest`; `nightly/v*` tags publish `:nightly`; only stable `v*` tags publish `:latest` and the versioned tags. New `workflow_dispatch` trigger for the release-job dispatches. Build-version stamping uses the exact nightly version on nightly tag builds - `.github/workflows/release-smoke.yml`: `nightly` added to the dispatch choice list - `doc/CHANNELS.md` (new): user-facing guide to the channels - `doc/RELEASING.md`: nightly lane documentation, Docker tag mapping table, and a nightly failure playbook - `doc/RELEASE-AUTOMATION-SETUP.md`: note that nightly reuses `npm-canary` and needs no npm trusted-publisher changes - Tests: channel-parameterized version helper coverage in `scripts/release-registry-versions.test.mjs`, and nightly flow coverage (publish identity, notes not required, canary-tag guard) in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the release script suites: 68 pass, including 6 new tests. The only failure, `acpx-patch-packaging.test.mjs`, needs installed `node_modules` and fails identically on a pristine checkout of master in the same environment - `bash -n` on both shell scripts and YAML parse of all three workflows - Live fail-path check: `./scripts/release.sh nightly --print-version` from a master tip with no canary tag fails with `HEAD has no canary/v* tag` - Live success-path check: the same command from the `canary/v2026.806.0-canary.7` commit prints `2026.806.0-nightly.0` - Live selection check: the candidate-selection shell logic run against the real repository selects the commit of `canary/v2026.806.0-canary.7`, which matches the current npm `canary` dist-tag exactly - After merge: dispatch `release.yml` with `channel: nightly` and `dry_run: true` to preview, then a real forced run to validate end to end before the first scheduled run ## Risks - Docker `:latest` changes meaning from "latest master build" to "latest stable release". This is deliberate and will be announced. Users who want the old behavior pull `:canary`. Until the first stable release after this change, `:latest` stays at its current (master-built) image - The nightly is a rebuild of the same source commit, not the byte-identical canary artifact that was smoked. The lockfile pins dependencies, and the publish path's registry-visibility and clean-prefix install gates still run on the nightly artifacts - All npm publishing must stay inside `release.yml` because npm trusted publishing pins that workflow file per package. The nightly jobs were added to `release.yml` for exactly that reason; this constraint is now documented in `RELEASING.md` - The stable-lane Docker dispatch fails gracefully (a warning with manual instructions) when the source ref predates `docker.yml`'s `workflow_dispatch` trigger ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
75acc4650f |
feat(sandbox-providers): pre-fill environment form with default sizing and image values (#11004)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runs execute inside sandbox environments. Sandbox provider plugins (Daytona, Modal, exe.dev, and others) declare a JSON-Schema `configSchema`. The Environment configuration form renders from that schema. > - The sizing and image fields open empty. Users must guess working values. The Modal form cannot submit at all until the user types an app name and an image by hand. > - The form renderer already pre-fills every field that declares a JSON-Schema `default`. The manifests do not use this mechanism for sizing or image fields. > - This pull request adds optional `default` values to the Daytona, Modal, and exe.dev manifest schemas. > - The benefit is a form that opens with known-good values. Users can create a working environment without provider research. ## Linked Issues or Issue Description **Current behavior** The Environment configuration form opens with empty sizing and image fields for the Daytona, Modal, and exe.dev sandbox providers. Users must find working values in provider documentation. Modal declares `appName` and `image` as required with no default, so the form blocks submission until the user invents both values. **Proposed behavior** The provider manifests declare JSON-Schema `default` values. The existing form renderer pre-fills them: - Daytona: CPU `4`, memory `4` GiB, disk `10` GiB, image `daytonaio/sandbox:0.8.0` - exe.dev: CPU `4`, memory `4GB`, disk `20GB` - Modal: app name `paperclip`, image `node:22` Secret-ref fields (API keys, tokens) get no defaults on purpose. The form persists a raw string in a secret-ref field as a company secret on save. A placeholder default would become a stored secret with a bogus value. Each plugin test suite now guards this invariant. **Reason and benefit** New users can create a working sandbox environment without guessing. The defaults stay optional: users can clear or change every value, and the schema marks no new field as required. The Modal image default `node:22` satisfies the sandbox runtime contract in `SANDBOX-REQUIREMENTS.md` (`node`, `sh`, and `tar` on PATH). **Subsystem affected** Sandbox provider plugins (`packages/plugins/sandbox-providers/*`): environment driver `configSchema` manifests. **Breaking changes** None. Defaults only seed the create-mode form. Saved environments keep their stored config. E2B, Novita, Cloudflare, and Kubernetes manifests do not change: E2B and Novita already default to their base templates, and the Cloudflare bridge and Kubernetes cluster fields have no sensible universal value. ## What Changed - Add `default` values for `cpu`, `memory`, `disk`, and `image` in the Daytona manifest. Trim the memory description to match. - Add `default` values for `cpu`, `memory`, and `disk` in the exe.dev manifest. - Add `default` values for `appName` and `image` in the Modal manifest. Extend the image description with the runtime-contract rationale. - Add manifest tests in all three plugins: defaults match expected values, defaults satisfy their own schema constraints, and no secret-ref field declares a default. - Bump plugin versions: daytona and modal `0.1.0` → `0.1.1`, exe-dev `0.1.1` → `0.1.2`. ## Verification - Run `pnpm test` in `packages/plugins/sandbox-providers/daytona`, `.../modal`, and `.../exe-dev`. The new `* manifest form defaults` suites pass. - Run `./node_modules/.bin/tsc --noEmit` in each of the three packages. Typecheck passes. - Manual: rebuild the plugins (`pnpm build` in each package), let the plugin dev-watcher refresh the manifest, then open Environments → New environment. The Daytona form shows CPU 4, Memory 4, Disk 10, and image `daytonaio/sandbox:0.8.0`. The Modal form shows `paperclip` and `node:22`. The API key fields stay empty. - Verified live on a local instance: the served `configSchema` in the plugin registry carries the new defaults, and existing environments are unchanged. ## Risks - Low risk. The change touches only manifest schema metadata and tests. No runtime code path changes. - New environments created with untouched forms now request 4 CPU / 4 GiB / 10 GiB from Daytona instead of provider minimums. This can raise cost per sandbox for users who previously saved empty fields. - The Daytona image default pins `daytonaio/sandbox:0.8.0`. The default needs a manual bump when Daytona ships new sandbox images. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (Anthropic, model ID `claude-fable-5`), via the Claude Code CLI, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
00a24d7e8f |
ci: split general-server tests into five shards with refreshed durations (#10925)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The PR workflow runs the server vitest suite across sharded runners because the suite is pinned to one worker. > - In successful PR run 30930345729 (2026-08-04), shard `server (3/4)` took 311 seconds of wall time and was the slowest check in the run. > - The suite has grown to about 946 seconds of serial vitest time, but the duration manifest was last sampled on 2026-08-01 at about 882 seconds. > - This pull request refreshes the per-suite duration manifest from that run's logs and splits the lane into five shards. > - The benefit is a shorter PR critical path: each shard carries about 196 seconds of suite time, level with the other lanes. ## Linked Issues or Issue Description Refs #10663 (previous split of this lane into four shards). Related: #10923 splits the separate serialized-suites lane into five shards. Both PRs touch `.github/workflows/pr.yml` in different matrix blocks; whichever merges second needs a trivial rebase. **What existing behavior does this improve?** The `general-server` vitest lane runs in four shards with a duration manifest sampled on 2026-08-01. **Current behavior** In PR run 30930345729, shard 3/4 ran for 311 seconds (273 seconds in the test step) and was the longest check in the run. The suite now totals about 946 seconds of serial vitest time. **Proposed behavior** Run the same suite set in five shards, balanced with a per-suite duration manifest refreshed from that run's shard logs (279 suites measured by diffing consecutive completion timestamps). **Reason and benefit** The refreshed LPT partition balances at about 196 seconds of suite time per shard (about 240 seconds per job), level with the other PR lanes. No test coverage is lost. **Breaking changes** None. The change only alters the CI partition size and the duration manifest. ## What Changed - Bump the `general-server` shard matrix in `.github/workflows/pr.yml` from four to five shards. - Refresh `scripts/general-server-shard-durations.json` from the 2026-08-04 run's shard logs. - Update `SHARD_COUNT` in `scripts/__tests__/run-vitest-stable-shard.test.mjs` to five. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass, including the complete non-overlapping partition proof and the duration-balance check. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass. - `node --test scripts/__tests__/e2e-shard.test.mjs` — 7/7 pass. - A 5-way dry-run partition covers all suites exactly once with equal projected weights. ## Risks - Low risk. The change only alters CI partition size and duration weights; the suite set is unchanged. - One more runner is used per PR run for this lane. - Stale duration weights degrade gracefully: suites missing from the manifest get the median weight. ## Model Used - Claude (Anthropic), Claude Code CLI, model ID `claude-fable-5`, extended thinking with tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (workflow comments explain the new shard math) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude <claude@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b6e58019f2 |
ci: split serialized tests into five shards (#10923)
## Thinking Path > - Paperclip uses CI to keep control-plane changes safe and mergeable. > - The PR workflow splits serialized server tests across isolated runners. > - A recent successful run spent 305 seconds in serialized shard 2/4. > - That job was the slowest check in the run. > - The four shards reported about 739 seconds of Vitest suite time. > - This pull request adds a fifth serialized shard and keeps release verification aligned. > - The benefit is a shorter PR critical path with no loss of test coverage. ## Linked Issues or Issue Description **What existing behavior does this improve?** The PR and release verification workflows run serialized server tests in four shards. **Current behavior** Successful PR run 30876682788 spent 305 seconds in `Verify serialized server suites (2/4)`. The test step used 256 seconds and made this job the slowest check. **Proposed behavior** Run the same serialized suite set in five complete and non-overlapping shards. **Reason and benefit** The measured suites reported about 739 seconds of total Vitest time. Five runners reduce the expected average suite time from about 185 seconds to about 148 seconds before setup overhead. **Breaking changes** None. The change only alters CI partition size. ## What Changed - Split serialized server tests into five shards in the PR workflow. - Apply the same five-shard layout to release verification. - Add a partition test that proves complete and non-overlapping serialized coverage. - Update release workflow coverage tests for five shards. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` - `git diff --check` ## Risks - Low risk. CI uses one additional runner for the serialized lane. - Round-robin partition weights can still vary as suite timings change. > This change does not overlap with planned core work in `ROADMAP.md`. Related PR #10663 optimized the separate general-server lane. ## Model Used - OpenAI Codex, GPT-5, agentic coding with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Devin Foley <139239+devinfoley@users.noreply.github.com> |
||
|
|
ffd62a4cbb |
fix(adapter-utils): carry the workspace origin remote into transported git workspaces (#10873)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - When an agent runs on a different host (sandbox or SSH), the adapter transport copies the local git execution workspace to that host and syncs changes back after the run > - The transport materializes the remote copy with `git init` plus a depth-1 or bundle fetch, so the copy has no `origin` remote and its head reads as a parentless snapshot commit > - An agent asked to publish its branch (push it, open a pull request) sees "no remote, root snapshot" and must hand the publish step back to a human operator, even when the branch base is a commit the upstream remote already holds > - This pull request carries the workspace's `origin` URL (credential-scrubbed) onto the transported copy as metadata > - The benefit is that branches produced in transported workspaces stay publishable by any actor with credentials, while the transport itself still never fetches or pushes ## Linked Issues or Issue Description No public issue exists. Description follows the enhancement template: **What existing behavior does this improve?** The workspace transport in `@paperclipai/adapter-utils` already copies a git workspace to the execution host and back. This change improves the fidelity of that copy: the transported repo keeps the workspace's `origin` remote instead of losing it. **Subsystem affected** Adapter utilities — the sandbox transport (`withShallowGitWorkspaceClone` in `packages/adapter-utils/src/git-workspace-sync.ts`) and the SSH transport (`importGitWorkspaceToSsh` in `packages/adapter-utils/src/ssh.ts`). **Current behavior** The transported copy is built with `git init` plus a depth-1 (sandbox) or bundle (SSH) fetch. It has no remotes. `git remote -v` is empty and the head commit reads as a root snapshot with no visible ancestry. Agents and operators inside the execution host cannot fetch real ancestry or push a branch, even when the branch base is a commit the upstream remote already holds. **Proposed behavior** The transport reads the source workspace's `origin` URL, scrubs credentials from it, and configures it on the transported copy. The sandbox path adds the remote to the fresh clone. The SSH path sets or adds the remote in the remote setup script, which also covers reused workspace directories. A workspace with no `origin` transports exactly as before. **Reason and benefit** A branch committed in a transported workspace becomes publishable in place: the shallow boundary commit already exists on the remote, so a push pack closes without full local ancestry (a new test locks in this property). Fetching real ancestry also becomes possible for whoever holds credentials. Without this, agents must describe their change in a handoff document and a human must reconstruct the branch by hand. **Breaking changes** None. The URL copy is best-effort and metadata-only. The transport never fetches from or pushes to the remote. The no-remote-git contract holds: sync-back through the local cwd stays the only cross-run persistence path, and `packages/adapters/AUTHORING.md` gains a paragraph that makes the carried-remote nuance explicit. ## What Changed - `packages/adapter-utils/src/git-workspace-sync.ts`: new `sanitizeGitRemoteUrl` (strips http(s) userinfo, where tokens can be embedded; scp-like/ssh forms and filesystem paths pass through) and `readSanitizedOriginRemoteUrl`; `withShallowGitWorkspaceClone` configures the scrubbed `origin` on the fresh clone, best-effort. - `packages/adapter-utils/src/ssh.ts`: `importGitWorkspaceToSsh` sets or adds the scrubbed `origin` in the remote setup script, non-fatal under `set -e`. - `packages/adapter-utils/src/git-workspace-sync.test.ts`: four new integration cases (remote copied, credentials scrubbed, no-origin unchanged, push from the shallow clone to an origin that holds the base commit) plus `sanitizeGitRemoteUrl` unit tests. - `packages/adapters/AUTHORING.md`: documents that a transported copy may carry a credential-scrubbed `origin` as metadata, and why this does not weaken the no-remote-git contract. ## Verification - `npx vitest run packages/adapter-utils/src/git-workspace-sync.test.ts` — 12/12 pass (4 new integration cases + sanitizer unit tests). - `npx vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — 24/24 pass. - `npx vitest run packages/adapter-utils/src/ssh-fixture.test.ts` — 16/16 pass, including the `no-remote-git contract` case (a workspace without `origin` still round-trips with no remote introduced at any point). - `node scripts/check-no-git-push.mjs` — passes; this change adds no push or fetch to adapter/runtime code. - `pnpm typecheck` in `packages/adapter-utils` — clean. ## Risks - Low risk. The change is additive metadata on the transported copy only; failure to record the remote never fails the transport. - Credential exposure is the real hazard and is handled: http(s) userinfo is stripped before the URL leaves the host. Non-http forms (scp-like, `ssh://`) carry no secret in the URL and pass through. - A reused SSH workspace whose project `origin` changed now gets the current URL via `set-url` instead of keeping a stale one. ## Model Used Claude Fable 5 (`claude-fable-5`), Anthropic — extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f0b06d2de9 |
feat(claude): environment-aware test-environment probe and claude-local CI coverage (#10833)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The claude_local adapter runs Claude Code on sandbox execution targets, and operators verify an agent's configuration with the test-environment probe before running it > - Real runs merge the selected environment's env vars (secret refs included) under the agent's adapter config env, but the probe built its config from the adapter config alone — so environment-level auth worked in runs while the Test button reported missing auth, and a dropped secret binding passed silently > - The claude env-test hints also did not recognize `CLAUDE_CODE_OAUTH_TOKEN` even though the CLI accepts it, and a hello probe that hit the subscription usage limit reported a hard failure although authentication worked > - Separately, the claude-local package test suites were absent from the CI project list, so two suites drifted broken without notice > - This pull request makes the probe resolve the same layered env as a real run, adds the missing auth hint, classifies usage-limit probe results as a warning, repairs the drifted suites, and turns the claude-local project on in CI > - The benefit is a Test button that tells the truth about environment-level configuration, and a test suite that actually gates the claude-local adapter ## Linked Issues or Issue Description No public issue exists; related open PRs: Refs #9488 (recognizes CLAUDE_CODE_OAUTH_TOKEN in environment checks — overlaps with the auth-hint portion of this PR via a differently named check; it does not cover the environment-envVars probe merge, the usage-limit classification, or the CI coverage), Refs #9933 (live credential validation in environment checks — complementary, no file-level conflict with the route change). The underlying problem, following the enhancement template: **Current behavior** The test-environment route builds the probe config from the agent's adapterConfig only. Real runs merge the selected environment's envVars under the agent env, so environment-level env vars (including auth such as `ANTHROPIC_API_KEY` or `CLAUDE_CODE_OAUTH_TOKEN` bound as environment secrets) work in runs while "Test environment" cannot see them, and a missing secret binding passes silently. The claude env-test hints do not recognize `CLAUDE_CODE_OAUTH_TOKEN`. A hello probe that hits the subscription usage limit reports a hard `claude_hello_probe_failed`. The claude-local package test suites do not run in CI, and two of them are stale. **Proposed behavior** The probe resolves the selected environment's envVars (environment-consumer secret bindings included) and merges them under the agent config env with the run-path precedence; missing bindings surface as an explicit error check that fails the test. The env-test emits a `claude_oauth_token_configured` info check when that variable is set. Usage-limit probe results classify as a `claude_hello_probe_usage_limited` warning because auth works and only the usage window is spent. The claude-local suites run in CI. Docs state the resulting facts. **Reason and benefit** The Test button should tell the truth: it previously contradicted run behavior for environment-level configuration and hid broken secret bindings. Enabling the package suites in CI prevents further silent drift — two suites were already broken on master without anyone noticing. **Breaking changes** None. Runs are unchanged. The probe route only adds env layers and checks; setups without environment envVars behave exactly as before. ## What Changed - `server/src/routes/agents.ts`: the test-environment route resolves the selected environment's envVars (forbidden keys stripped, environment-consumer secret context) and merges them under the agent adapterConfig env, mirroring `resolveExecutionRunAdapterConfig` precedence. Missing secret bindings are skipped, reported as an `environment_env_binding_missing` error check, and fail the test — matching the `ConfigurationIncompleteFailure` a real dispatch would raise. - `packages/adapters/claude-local/src/server/test.ts`: new `claude_oauth_token_configured` info hint between the API-key warning and the subscription fallback; hello-probe classification gains a `claude_hello_probe_usage_limited` warning for provider-quota results (previously a hard `claude_hello_probe_failed`). - `scripts/run-vitest-stable.mjs`: add `@paperclipai/adapter-claude-local` to `nonServerProjects` so CI runs the package suites. - `packages/adapters/claude-local/src/server/execute.remote.test.ts`: assert both runtime asset syncs (skills and mcp-config); the suite predated the mcp-config asset. - `packages/adapters/claude-local/src/server/test.probe.test.ts`: usage-limit fixture now expects the usage-limited warning; new fixture covers the genuine transient path (529 overloaded); new tests cover the token hint and API-key precedence. - `server/src/__tests__/agent-test-environment-routes.test.ts`: new tests for the env merge (agent wins on conflict, forbidden key filtered), missing-binding reporting, and the no-execution-target fallback path. - `docs/adapters/claude-local.md`, `docs/adapters/overview.md`: state the auth-input facts (API key or oauth token wins over stored logins; snapshot-owns-auth applies when neither is configured) and describe the environment-aware Test behavior. ## Verification - `npx vitest run --project @paperclipai/adapter-claude-local` — 131 tests pass (both drifted suites repaired; they fail on master today). - `npx vitest run server/src/__tests__/agent-test-environment-routes.test.ts` — 7 tests pass. - `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` — passes with the added project. - `pnpm typecheck` in `server/` and `packages/adapters/claude-local/` — clean. ## Risks - Low risk. The run path is untouched; the probe route change is additive and inert when the environment has no envVars. - The probe now performs environment-consumer secret resolution at test time; access is authorized per binding exactly as at run time, and the audit consumer is the environment (as before for adapter-config resolution). - Enabling the claude-local project in CI adds about 2 seconds of vitest wall time to the general workspaces group and could surface future regressions in that package — which is the point. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking enabled, agentic tool use via Claude Code (CLI). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
cb52f0b750 |
db: env-configurable client options; parallelize attention feed queries (#10795)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The server stores all state in PostgreSQL through Drizzle and the
postgres.js driver
> - Self-hosted installs run Postgres on localhost, so per-query latency
is near zero; hosted installs often attach Postgres over a network,
sometimes through a transaction-mode pooler
> - The DB client passes no options to the driver, so operators cannot
disable prepared statements or tune the pool without a source edit, and
the deploy docs told them to edit `client.ts`
> - The attention feed also runs its related-data lookups one after
another, so its latency grows as queries × network round trip
> - This pull request adds optional environment configuration for the DB
client and batches the independent attention-feed lookups with
`Promise.all`
> - The benefit is that network-attached deployments get correct pooler
support and a much faster attention feed, while self-hosted behavior
does not change
## Linked Issues or Issue Description
No public issue exists for this; description follows the bug report
template:
**What happened?**
On deployments where PostgreSQL is network-attached (managed providers,
pooled endpoints), the attention feed endpoint is slow:
`attentionService.list()` awaits ~15–20 queries strictly in sequence, so
a 70ms round trip turns into more than one second of pure network wait
per call. Separately, connecting through a transaction-mode pooler
(pgbouncer, Supavisor port 6543, Neon `-pooler` hosts) requires
disabling prepared statements, and the only documented way was to
hand-edit `packages/db/src/client.ts` — which `doc/DATABASE.md` itself
tells operators not to do.
**Expected behavior**
The DB client is configurable from the environment (prepared statements,
pool size, timeouts) with driver defaults when unset, and hot read paths
do not multiply network latency by issuing independent queries
sequentially.
**Steps to reproduce**
1. Run the server with `DATABASE_URL` pointing at a Postgres instance
with ~70ms round-trip latency.
2. Open the attention feed (`GET /companies/:companyId/attention`) and
measure response time — it exceeds one second even with little data.
3. Try to connect through a transaction-mode pooler: there is no
supported configuration to disable prepared statements.
## What Changed
- `packages/db/src/client.ts`: `createDb` accepts a
`DatabaseClientOptions` argument and reads optional env config —
`DATABASE_PREPARED_STATEMENTS`, `DATABASE_POOL_MAX`,
`DATABASE_IDLE_TIMEOUT_SECONDS`, `DATABASE_CONNECT_TIMEOUT_SECONDS`.
When nothing is set, no option is passed to the driver and behavior is
identical to the previous bare `postgres(url)`.
- `packages/db/src/client-options.test.ts` (new): env parsing and
driver-option mapping tests, including malformed-value rejection.
- `server/src/services/attention.ts`: the independent related-data
lookups in each feed section now run under `Promise.all` (issue
summary/image/plan-document maps, decision bundle titles, blocked-issue
maps, the newer-runs scan). Section order, item assembly, and query
shapes are unchanged.
- `doc/DATABASE.md` and `docs/deploy/database.md`: the edit-source
pooling instruction is replaced with the env toggle, plus a short
client-tuning reference.
## Verification
- `pnpm --filter @paperclipai/db exec vitest run
src/client-options.test.ts` — 6 tests pass.
- `pnpm --filter server exec vitest run
src/__tests__/attention-service.test.ts` — 22 tests pass.
- `pnpm --filter server exec vitest run
src/__tests__/decisions-service.test.ts
src/__tests__/decision-training.test.ts` — 45 tests pass; this covers
the call path that runs `attentionService.list()` inside
`db.transaction`, where postgres.js serializes queries on the reserved
connection.
- `tsc` reports no errors in the changed files.
## Risks
- Low risk for self-hosted installs: with no env vars set,
`postgres(url, {})` receives an empty options object, which postgres.js
treats the same as no options — driver defaults throughout.
- The `Promise.all` batches only group queries that had no data
dependency on each other; on the transaction call path the driver still
executes them one at a time on the reserved connection, so transactional
semantics are unchanged.
- Malformed env values now fail fast at startup with a clear message
instead of being silently ignored; this is intentional and only affects
operators who set the new variables.
## Model Used
Claude Fable 5 (`claude-fable-5`), Anthropic — via Claude Code CLI,
extended thinking enabled, tool use (test execution, live latency
measurement against a network-attached Postgres to size the problem).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (searched "prepared statements", "pgbouncer", "pool",
"attention feed", "lockfile" — closest matches are #10573/#10787
lockfile chores, unrelated to this change)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
2c90cf0f2c |
fix(server): serialize managed-checkout materialization and stop misattributing clone failures (#10723)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Repo-only project workspaces are materialized by a server-side
managed `git clone` into a per-project directory (#10720 added
credentials for private repos)
> - Two issues on the same project routinely wake seconds apart, and
both runs race the same clone target
> - The loser fails with "destination path already exists", and its
failure cleanup removes the directory out from under the winner's
in-progress clone — both runs then fail every round
> - This pull request serializes materialization per target directory
and makes the clone land atomically via a temp sibling + rename, so the
shared target is never partial and never removed
> - The benefit is that concurrent runs on the same project converge:
one clone happens, everyone adopts it, and unrelated failures no longer
blame the GitHub credential
## Linked Issues or Issue Description
**What happened?**
With isolated workspaces enabled on a project whose only workspace is
repo-only, unblocking two issues at once produced lockstep mutual
destruction (observed live, two consecutive rounds): both runs called
the managed-checkout materialization concurrently; one clone created the
target directory, the other's `git clone` failed with `fatal:
destination path '…' already exists and is not an empty directory`, and
that run's failure cleanup deleted the directory while the first clone
was still writing into it (`fatal: could not set 'core…'`). Both runs
failed `workspace_validation_failed`; their staggered retries could race
again. The failure message also wrongly claimed the GitHub credential
"was rejected or lacks access" — the collision had nothing to do with
auth.
**Expected behavior**
Concurrent materializations of the same project checkout share one
clone; a completed checkout is never removed by a failing sibling; the
credential is only blamed for auth-shaped failures.
**Steps to reproduce**
1. Project with a repo-only workspace (private repo, isolated workspaces
on).
2. Move two issues on that project to `todo` at the same time so both
runs start within seconds.
3. Both runs fail workspace validation with "destination path already
exists" / "could not set 'core…'" instead of one clone succeeding.
**Paperclip version or commit**
`master` (
|
||
|
|
75f6256b76 |
fix(server): resolve duplicate scrubGitCredentialText declaration on master (#10722)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - #10719 and #10720 both hardened credential scrubbing for server-side git operations > - #10719 defined `scrubGitCredentialText` locally in `heartbeat.ts`; #10720 imported the identical function from the new `git-credentials.ts` > - The two merged cleanly textually, but together they leave `heartbeat.ts` with an import that conflicts with a local declaration (TS2440) > - This pull request keeps the `git-credentials.ts` copy as canonical, drops the heartbeat-local duplicate, and re-exports the import so existing importers are unchanged > - The benefit is that master typechecks, builds, and produces Docker images again ## Linked Issues or Issue Description **What happened?** After #10719 and #10720 merged, master fails typecheck, the server build, and both Docker image builds with `src/services/heartbeat.ts(76,3): error TS2440: Import declaration conflicts with local declaration of 'scrubGitCredentialText'`. The canary release run for |
||
|
|
e0c2448267 |
feat(server): authenticate server-side git clone and fetch with a company-secret GitHub token (#10720)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Repo-only project workspaces are materialized by a server-side `git clone`, and isolated `git_worktree` runs refresh their base ref with server-side `git fetch` > - Both operations run outside the agent process with no credentials, so private GitHub repositories can never be cloned or refreshed — agent-scoped credential env bindings do not reach them > - The company secret store already has a well-known GitHub token convention (`GITHUB_TOKEN` / `GH_TOKEN` / `PAPERCLIP_GITHUB_TOKEN`, consumed by the external-object provider for API reads), but nothing server-side consults it for git > - This pull request resolves that token per run and authenticates the managed clone and every base-ref refresh with it through an ephemeral credential helper > - The benefit is that isolated workspaces work on private repositories with one company secret, while public repositories and self-hosted ambient git configuration keep working unchanged ## Linked Issues or Issue Description **Subsystem affected** Server workspace materialization (`server/src/services/heartbeat.ts`) and execution-workspace realization (`server/src/services/workspace-runtime.ts`). **Problem or motivation** A project workspace configured with only a private GitHub `repoUrl` cannot be used for isolated `git_worktree` runs: the managed `git clone` runs with a sanitized, credential-less environment, and plain git cannot consume a bare token env variable without a credential helper. There is no way to give the server a git credential — storing a `GH_TOKEN` company secret has no effect on server-side git, and a credential-less private clone hangs on a terminal prompt until the ten-minute clone timeout. Base-ref refreshes (`git fetch`) during worktree realization have the same gap. **Proposed solution** A `git-credentials` module resolves a token per run — company secret by well-known name (`GITHUB_TOKEN`, `GH_TOKEN`, `PAPERCLIP_GITHUB_TOKEN`), then `GITHUB_TOKEN`/`GH_TOKEN` in the server process environment for self-hosted deployments, then none — and builds a git invocation that authenticates via an inline credential helper. The token travels in an env variable; it never appears in argv, URLs, or on disk. Only `https://github.com` remotes are authenticated; everything else keeps ambient behavior. The provider is a single factory seam so a future brokered credential source can replace it without touching call sites. **Alternatives considered** - A GitHub OAuth "connect your account" flow: heavier product surface, needs app registration and callback custody; out of scope for a server credential and better served by a dedicated connector later. The provider seam keeps that path open. - `gh auth setup-git`: writes helper configuration to disk and requires a global token env; rejected in favor of per-invocation config with no persistent state. - Embedding the token in the clone URL: leaks into argv, error messages, and `.git/config`; rejected. ## What Changed - New `server/src/services/git-credentials.ts`: `createGitRemoteAuthProvider` (memoized per run, one secret resolution and one audit event), `buildGitAuthInvocation` (helper-reset + inline helper, `x-access-token` username, `GIT_TERMINAL_PROMPT=0`), `isGitHubHttpsRemoteUrl` host gating (rejects ssh/GHES/http/other hosts/userinfo URLs), `describeGitAuthFailure`, and the canonical `scrubGitCredentialText`. Secret resolutions pass a `system` consumer access context so they are recorded as secret access events. - `ensureManagedProjectWorkspace` (now exported) accepts an optional auth provider; the clone env spreads the token after `sanitizeRuntimeServiceBaseEnv` (which strips `PAPERCLIP_*`), always sets `GIT_TERMINAL_PROMPT=0`, distinguishes "credential rejected" from "no credential configured — add a GITHUB_TOKEN or GH_TOKEN company secret" in the error, and removes the partially created directory on clone failure so a timeout-killed clone cannot be adopted as a broken checkout by the next run. - `refreshRemoteTrackingBaseRef` (now exported) captures the remote URL it already looked up, asks the provider for an invocation, and attributes failed authenticated fetches to the credential in a scrubbed warning. The optional provider threads through `detectDefaultBranch`, `resolveAuthoritativeBaseRef`, `inspectExecutionWorkspaceBaseDrift`, `realizeExecutionWorkspace`, and `ensurePersistedExecutionWorkspaceAvailable`; heartbeat builds one provider per run for both the anchor-resolution clone path and workspace realization/restore. - `github-external-object-provider.ts` imports the shared secret-name list; `isGitHubDotCom` is exported from `github-fetch.ts`. - Docs: "Private repositories and repo-only project workspaces" section in the execution-workspaces guide, cross-linked from the secrets deploy doc. ## Verification - `cd server && npx vitest run src/__tests__/git-credentials.test.ts` — resolution chain order and precedence, env fallback, memoization, audited access context, host-gating matrix, invocation shape (token absent from argv), scrubber, failure descriptions, and a real-git `git credential fill` round trip that proves the helper executes and answers with the env-carried token (no network). - `cd server && npx vitest run src/__tests__/heartbeat-managed-clone-credentials.test.ts` — clones behave byte-identically with no provider or a null-returning provider (local repos, no network), authenticated-failure errors name the credential, non-auth failures do not mention credentials, partial clone directories are removed, pre-existing non-git directories keep the "Using it as-is" path, and the sanitizer spread order keeps the token env alive. - `cd server && npx vitest run src/__tests__/workspace-runtime.test.ts` — new `refreshRemoteTrackingBaseRef` cases: provider offered the remote URL and null keeps behavior identical; failed authenticated fetch warning names the credential; unauthenticated failure warning stays credential-free. - `pnpm --filter @paperclipai/server typecheck` is clean. - Manual (optional, networked): store a `GH_TOKEN` company secret, configure a repo-only project workspace pointing at a private GitHub repository, run an isolated-workspace issue — the managed clone succeeds and the worktree run proceeds. ## Risks - Every new parameter is optional; with no provider the git invocations are byte-identical to before. Public repos and ambient credential helpers keep working whenever no token resolves. - Precedence change when a token exists: a stored company secret now wins over ambient helpers for `https://github.com` remotes (the helper list is reset for that invocation). The rejected-credential error names the secret so an operator can fix or remove it. - `GIT_TERMINAL_PROMPT=0` on the managed clone is the one always-on change: a credential-less private clone now fails fast with a clear message instead of hanging until the ten-minute timeout (it could only ever "succeed" interactively on a TTY dev server). - The token is scoped to the git process env for one invocation; it is never written to agent env, run context, disk, or logs, and error text is scrubbed of URL userinfo. - No migrations, no image changes (git ships in the image). ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, agentic tool use via Claude Code CLI). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
97590ff8c4 |
feat(dev): add pnpm dev:mobile and dev:both for prebuilt UI preview (#10718)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The board UI is a React SPA served by the paperclip server; the standard local dev flow is `pnpm dev`, which runs vite in dev mode with HMR and an unbundled module graph > - The unbundled dev bundle is hundreds of MB of JS across many requests, which is fine on a local machine but unusable from a phone or tablet on slow/lossy links (airplane wifi, mobile data, distant tailnet peers) > - Contributors who want to iterate on the board from a mobile device today have no supported way to preview a small production-shaped bundle without stopping the dev server and running a one-off `vite preview` with manual proxy plumbing > - This pull request adds `pnpm dev:mobile` — build the UI and serve `ui/dist` via `vite preview` on port 3101, with `/api` proxied to the running dev server on 3100 — plus `pnpm dev:both` to run both flavors together > - The benefit is a supported second flavor of the dev server for phones/tablets that runs alongside the normal one, without touching the primary `pnpm dev` flow ## Linked Issues or Issue Description **Subsystem affected** ui/ — React + Vite board UI **Problem or motivation** The vite dev server serves an unbundled module graph, which is fine on localhost but unusable from a phone or tablet on a slow link. Contributors testing responsive behavior on mobile devices have no supported way to serve a small production-shaped SPA against the running dev API. Running `vite preview` directly does not work either — the server's board mutation guard checks that the browser's Origin matches the request Host, and a preview on a second port would fail every mutation. **Proposed solution** Add two root scripts: - `pnpm dev:mobile` — build `ui/dist` and serve it via `vite preview` on port 3101, with `/api` proxied to the API server on 3100. - `pnpm dev:both` — run `pnpm dev` and `pnpm dev:mobile` together in a single terminal with prefixed output and shared signal handling. The vite preview config binds `0.0.0.0`, sets `allowedHosts: true` so it accepts arbitrary hostnames (LAN, tailnet, ngrok, etc.), and the shared `/api` proxy forwards the client's original Host header as `x-forwarded-host`. The paperclip server's mutation guard already prefers `x-forwarded-host` over `host` when computing trusted origins, so the browser's Origin becomes trusted automatically. **Alternatives considered** - Bespoke node proxy script — works but duplicates what vite preview already does. - Loosen the mutation guard to accept arbitrary origins — reduces security for the primary server for the sake of a dev-only workflow. - Second server config that binds a second port from the paperclip server itself — much larger change and mixes runtime concerns with a dev-tooling convenience. ## What Changed - New `pnpm dev:mobile` script — build UI then run `vite preview` on port 3101. - New `pnpm dev:both` script — run `pnpm dev` and `pnpm dev:mobile` together via `scripts/dev-both.mjs`, which prefixes each child's output, propagates SIGINT/SIGTERM, and exits when either child exits. - `ui/vite.config.ts` — add a `preview` block (port 3101, host `0.0.0.0`, `allowedHosts: true`, shared `/api` proxy). - New `ui/src/lib/vite-api-proxy.ts` — extracts the `/api` proxy factory shared by dev and preview, and forwards the client Host as `x-forwarded-host` (plus `x-forwarded-proto`). - New unit test `ui/src/lib/vite-api-proxy.test.ts` covering the header-injection behavior and the pass-through when no Host is present. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/lib/vite-api-proxy.test.ts` — 3 tests pass. - `pnpm --filter @paperclipai/ui typecheck` — clean. - `pnpm --filter @paperclipai/ui build` — clean. - Manual: ran `vite preview` against an echo listener and confirmed the request arrives with `x-forwarded-host` set to the client Host header and `x-forwarded-proto: http`. Then ran `pnpm dev:mobile` against the live dev server and verified board mutations (mark issue read, resolve recovery action, run routine) succeed from a second-port browser session that previously 403'd. ## Risks Low risk. Changes are limited to dev tooling — no runtime code paths, no server changes, no schema/migrations. The `apiProxy` refactor is a no-op behaviorally for the existing dev server (same target, same `ws: true`); the only new behavior is the two `x-forwarded-*` headers, and the server side already prefers those headers when trusting origins. `dev:mobile` and `dev:both` are additive; existing `pnpm dev` is untouched. ## Model Used Claude Opus 4.7 (1M context), extended thinking, tool use (bash, file edits). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
185515c97b |
fix(external-objects): refresh PR status labels (#10704)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue properties panel can show external objects such as GitHub pull requests. > - Those objects are resolved by external-object providers and then displayed as compact status labels. > - A GitHub pull request could remain in the fallback `unknown` state and appear as `Not yet resolved`. > - That label is confusing when the object is known but has not been refreshed yet. > - This pull request refreshes due external objects from the heartbeat scheduler and improves the unknown-status copy. > - The benefit is a properties panel that moves from pending refresh to the real pull request state without a manual refresh. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. I searched for related public issues and pull requests using the terms `Not yet refreshed`, `external objects refresh`, and `external PR status`, and did not find a duplicate implementation. **What happened?** The issue properties panel could show a GitHub pull request as `Not yet resolved` even when the referenced pull request was valid. The object stayed stale unless a manual refresh path ran. **Expected behavior** A known external object should show pending-refresh copy while it waits for provider data. When the scheduler refreshes it, the properties panel should show the provider status such as open, merged, or closed. **Steps to reproduce** 1. Create or view an issue that references a GitHub pull request. 2. Open the issue properties panel. 3. Observe the external object row before a manual refresh has run. **Paperclip version or commit** Current `master` before this pull request. **Deployment mode** Local dev and self-hosted server. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific. **Database mode** Not database-related. **Access context** Board view. **Privacy checklist** I reviewed this description and did not include logs, credentials, private URLs, internal issue IDs, or PII. ## What Changed - Added a heartbeat scheduler tick that refreshes due external objects for active companies. - Kept manual external-object refresh behavior on the same service path. - Changed display copy so known provider objects use liveness labels such as `Not yet refreshed`, while fresh unknown provider statuses show `Status unavailable`. - Added server and UI tests for scheduled refresh and label behavior. ## Verification - `corepack pnpm install --frozen-lockfile` - `pnpm check:token-gates` - `pnpm exec vitest run server/src/__tests__/external-objects-service.test.ts server/src/__tests__/server-startup-feedback-export.test.ts ui/src/components/ExternalObjectPill.test.tsx ui/src/components/IssueProperties.test.tsx ui/src/lib/external-objects.test.ts` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server build` - `pnpm --filter @paperclipai/ui build` - `pnpm run typecheck:build-gaps` - GitHub PR checks passed on head `0e7fcd30` - Greptile reported 5/5 on head `0e7fcd30` with no unresolved review threads Notes: - I ran recursive typecheck and build first. Both hit container resource limits with exit 137 during concurrent package work, so I reran the affected server and UI targets separately. - An unrelated workspace-runtime auto-port test fails in this container with a PID ownership mismatch. It is outside the files changed here. ## Risks Low to medium risk. The scheduler does more periodic external-object work, so the main risk is extra provider refresh load. The implementation bounds the work to active companies, due non-terminal objects, and 50 objects per company per tick. The path also stays behind the external-objects experimental setting. ## Model Used OpenAI GPT-5 Codex in the Codex execution environment, with shell and GitHub CLI tool use. The runtime did not expose a more specific internal model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7a3815eb9a |
fix(heartbeat): surface the real cause when a git_worktree base cannot be materialized (#10719)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Heartbeat runs prepare an execution workspace for each issue; with
isolated workspaces (or low-trust runs), the `git_worktree` strategy
needs a real git checkout as its base
> - For repo-only project workspaces, the server materializes the
checkout with a managed `git clone`; when that clone fails, the resolver
drops the error and silently falls back to the agent home directory with
`source: "project_primary"`
> - The pre-dispatch guard then reports
`git_worktree_base_not_git_checkout`, which hides the real cause; if the
fallback directory happens to be a git checkout, the run silently builds
worktrees off the wrong repository
> - This pull request records materialization failures on the resolved
workspace, marks the fallback explicitly, and fails the guard with a
truthful `git_worktree_base_materialization_failed` reason that carries
the clone error
> - The benefit is that operators see the real cause (a failed clone) in
the run error, the blocked-issue comment, and the recovery next action,
instead of a misleading symptom
## Linked Issues or Issue Description
**What happened?**
An issue configured for isolated `git_worktree` execution on a project
whose only workspace is repo-only (a `repoUrl` with no local path) fails
every run with `workspace_validation_failed` and reason
`git_worktree_base_not_git_checkout`, pointing at the agent home
directory. The message does not mention that the managed `git clone` of
the project repository failed (for a private repository the clone can
never succeed without credentials). The recovery flow then blocks the
issue with the same misleading explanation. Run warnings claim "Project
workspace has no local cwd configured" even though the workspace is
configured and the clone failed.
**Expected behavior**
The run failure, the blocked-issue comment, and the recovery next action
should state the real cause: the project workspace checkout could not be
prepared, including the clone error, so the operator can repair the
repository URL, clone access, or configured local cwd. A fallback
directory that happens to be a git checkout must not let the run proceed
against the wrong repository.
**Steps to reproduce**
1. Create a project whose primary workspace has a `repoUrl` pointing at
a private GitHub repository and no local path.
2. Enable the Isolated Workspaces experimental setting (or use a
low-trust run, which forces isolation).
3. Run any issue in that project.
4. The run fails with `git_worktree_base_not_git_checkout` on the agent
home directory; the clone failure appears nowhere.
**Paperclip version or commit**
`master` (
|
||
|
|
799973f26a |
fix(ui): hide empty inbox search sections (#10700)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Inbox helps operators scan issues that need attention. > - Inbox search can add supplemental sections for archived matches and other matches. > - The supplemental search section builder still sent empty sections into the grouped render path. > - That made the Archived and Other results dividers appear even when those sections had no rows. > - This pull request drops empty supplemental sections before rendering. > - The benefit is a cleaner near-empty inbox search view. ## Linked Issues or Issue Description No public GitHub issue exists for this report. Public GitHub search found no duplicate or related open issues or pull requests for this inbox search behavior. **What happened?** Inbox search could show Archived and Other results divider headers even when those supplemental sections had no rows. **Expected behavior** Empty supplemental search sections should not render divider headers. **Steps to reproduce** 1. Open the Inbox. 2. Search in a near-empty inbox with no archived matches and no outside-inbox matches. 3. Observe that empty supplemental divider headers can appear. **Paperclip version or commit** `master` before this change. **Deployment mode** Built from source. ## What Changed - Dropped empty supplemental inbox search sections before they reach the grouped inbox render path. - Added a unit regression test for empty Archived and Other results sections. - Refreshed the branch against current `master` to clear the merge conflict. ## Verification - `git diff --check origin/master...HEAD` passed. - Public diff is limited to `ui/src/lib/inbox.ts` and `ui/src/lib/inbox.test.ts`. - Local focused Vitest could not run in this execution checkout because dependencies are not installed and `corepack pnpm exec vitest ...` reports `Command "vitest" not found`. - Pull request CI is green for typecheck, build, server tests, e2e, security checks, policy checks, canary dry run, and aggregate verify. - Greptile Review passed on commit `dc2e224` with confidence score 5/5 and no comments. ## Risks Low risk. The Inbox change only filters empty supplemental search sections. Normal inbox sections and non-empty archived or other search results keep their current behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, reasoning-enabled with terminal tool use and code execution. The runtime context-window size is not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bd86dbe41b |
fix(codex-local): detect server-visible Codex credentials in ACP environment test (#10703)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The codex-local adapter runs Codex agents through the ACP lane and offers a "Test" button on the agent configuration page to validate the environment > - The environment test used its own ad-hoc credential probe, while real dispatch uses the shared `evaluateCodexCredentialReadiness` predicate in `codex-home.ts` > - The two paths disagreed: a user with valid Codex subscription auth in the shared, server-visible Codex home still saw "No Codex ACP credentials were detected" > - The old warning also suggested `codex login` without explaining that a `/login` in a separate Codex or chat session does not authenticate the Paperclip server process > - This pull request makes `testCodexAcpEnvironment` use the same shared readiness predicate as real dispatch and rewords the warning to name the server credential boundary > - The benefit is that the Test button now agrees with what dispatch will actually do, and the warning tells the user exactly which process needs the credentials ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Related credential-handling work, not duplicates: Refs #10160 (classifies OpenAI invalid-key 401s at dispatch time) and Refs #9598 (classifies Codex refresh auth failures). Both cover dispatch-time failures; this PR fixes the pre-dispatch environment test. Description follows the bug report template: **What happened?** A user configured a Codex ACP agent and had already authenticated Codex (subscription auth present in the shared Codex home visible to the Paperclip server). Clicking "Test" on the agent configuration page still reported: `warn: No Codex ACP credentials were detected. Hint: Set OPENAI_API_KEY or run codex login before starting a Codex ACP agent.` **Expected behavior** The environment test should detect the same credentials that real agent dispatch would use. When shared managed Codex auth is available, the test should pass with an informational check instead of warning. When credentials really are missing, the warning should explain that the Paperclip server process is the one that needs them. **Steps to reproduce** 1. Run the Paperclip server as an OS user whose shared Codex home contains valid subscription `auth.json` (no `OPENAI_API_KEY` in the adapter env or server env). 2. Configure an agent with the codex-local adapter using the ACP engine. 3. Click "Test" on the agent configuration page. 4. Observe the `codex_acp_credentials_missing` warning even though dispatch would succeed. **Agent adapter(s) involved** Codex **Additional context** The confusion was amplified by the hint: users had run `/login` in a Codex chat session and assumed the server was authenticated. That login lives in a different process and home directory, so the server never saw it. ## What Changed - `testCodexAcpEnvironment` (packages/adapters/codex-local/src/server/acp.ts) now calls the shared `evaluateCodexCredentialReadiness` predicate from `codex-home.ts` instead of a local ad-hoc `hasCodexNativeCredentials` probe, so the Test button and real dispatch agree. - An explicit empty `OPENAI_API_KEY` in the adapter config env no longer falls through to the server environment key. - An externally managed `CODEX_HOME` override is now reported as its own informational check (`codex_acp_external_home_configured`). - The `codex_acp_credentials_missing` warning now says the credentials must be visible to the Paperclip server, and the hint explains that a `/login` in a separate Codex or chat session does not authenticate the server. - Removed the now-unused `hasCodexNativeCredentials` helper. - Added two regression tests: shared managed Codex auth is detected (no false warning), and the missing-credentials warning carries the new server-boundary wording. ## Verification - `pnpm --filter @paperclip/adapter-codex-local test -- src/server/acp.test.ts` — the two new tests cover the shared-home detection branch and the new warning wording; the existing ACP lane tests cover the API-key and remote-target branches. - Manual: with subscription auth in the server-visible shared Codex home and no `OPENAI_API_KEY`, the agent configuration Test now reports `codex_acp_native_auth_detected` (info) instead of `codex_acp_credentials_missing` (warn). ## Risks - Low risk. The change only affects the environment test path, not dispatch. The readiness predicate is the same one dispatch already uses, so drift between the two paths is now structurally prevented. - Behavioral shift: an explicit empty adapter `OPENAI_API_KEY` no longer silently falls back to the server env key in the test result. This matches dispatch behavior and is intentional. ## Model Used - Implementation authored by OpenAI Codex (gpt-5.5) running through the Codex ACP lane with tool use. - PR preparation, rebase onto master, and review fix-up by Anthropic Claude (Claude Code CLI agent, extended thinking, tool use). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Cody <noreply@paperclip.ing> |
||
|
|
8b83d69e3c |
feat(heartbeat): serialize shared-workspace issue runs with bounded busy deferrals (#10699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service dispatches agent runs, and issues in a project can share one project workspace (one working tree on disk). > - Two runs can execute in the same shared workspace at the same time. Each run mutates the same uncommitted files and branches, and the runs corrupt each other's state. > - Multi-agent projects hit this as soon as two issues in one project become active together, so the platform needs to serialize shared-workspace execution instead of relying on luck. > - This pull request adds a pre-dispatch gate: a run whose issue targets a busy shared workspace is deferred with a bounded scheduled retry instead of dispatched. > - The benefit is that concurrent issue runs in one project take turns in the shared working tree, while isolated-workspace runs and unrelated workspaces stay fully parallel. ## Linked Issues or Issue Description Fixes #10645 ## What Changed - `server/src/services/heartbeat.ts`: - New pre-dispatch gate in the run executor. Before adapter dispatch, when the run's issue has a `projectWorkspaceId` and the effective execution workspace mode is `shared_workspace`, the executor looks for a holder: another `running` run whose context issue shares the same project workspace. The gate covers every run shape that reaches adapter dispatch with issue context — assignee execution runs, comment/mention interaction wakes, and review-participant runs. - When a holder exists, the run throws `WorkspaceBusyDeferral` instead of dispatching. The outer catch recognizes the deferral: it cancels the run with `errorCode: "workspace_busy"` (contention is not a failure), cancels its wakeup, schedules a retry through the existing `scheduleBoundedRetryForRun` primitive (`workspace_busy` reason, 60–120 s jittered delay), and returns the agent to idle. The issue execution lock transfers to the scheduled retry run, so the issue keeps an active execution path and stranded-issue recovery does not fire. - An adapter never dispatches alongside a live holder: deferral has no attempt ceiling, so a deferred run keeps rescheduling until the workspace frees. Deadlock safety comes from holder liveness, not a counter — a holder silent past `ACTIVE_RUN_OUTPUT_SUSPICION_THRESHOLD_MS` (recovery's own "suspicious silence" bar, 1 h) stops counting as a holder, so a zombie run can only delay work, never park it forever, and recovery's silent-run escalation is already reaping it in parallel. If no retry can be scheduled (agent paused, issue reassigned), the deferral releases the issue execution lock so the issue does not strand. - Holder detection honors isolation: when the isolated-workspaces experiment is enabled, holders whose issue settings select `isolated_workspace` / `operator_branch` (or the legacy `isolated` alias) are not counted, because they never touch the shared tree. A NULL or `agent_default` mode counts as a holder — over-serializing is the safe direction. - Non-assignee deferrals survive replay: the deferral stamps `workspaceBusyDeferredWhileAssignee` into the run context (inherited by the scheduled retry), and both the retry promotion gate and the claim-time staleness check exempt a non-assignee `workspace_busy` retry from the reassignment cancellation — for such a retry an assignee mismatch is the expected state, not a reassignment race. An assignee run's retry keeps the full protection: if the issue is reassigned while the retry pends, it still cancels with `issue_reassigned`. - `server/src/__tests__/heartbeat-workspace-busy.test.ts` (new): embedded-Postgres coverage of the full lifecycle plus unit coverage of the delay window. ## Verification - `cd server && pnpm vitest run src/__tests__/heartbeat-workspace-busy.test.ts` — 10 tests: - a run whose issue targets a busy shared workspace is cancelled with `workspace_busy`, its adapter never executes, a `scheduled_retry` run exists with the 60–120 s window, the issue execution lock points at the retry run, the holder run is untouched, and the agent returns to idle; - after the holder finishes, `promoteDueScheduledRetries` + `resumeQueuedRuns` execute the retry run to success; - a non-assignee comment-mention wake defers, does not touch the issue execution lock, and its retry promotes, survives the claim-time staleness check, and executes despite the assignee mismatch; - an assignee retry is still cancelled with `issue_reassigned` when the issue is reassigned while the retry pends; - a holder issue with `executionWorkspaceSettings.mode = "isolated_workspace"` does not cause deferral; - a running run in a different project workspace does not cause deferral; - a holder silent past the staleness threshold does not cause deferral (the run executes); - a retry with ten prior deferrals still defers again — never dispatches — while the holder is live; - delay jitter stays inside the base-to-base-plus-jitter window and clamps out-of-range random sources. - `cd server && pnpm vitest run src/__tests__/heartbeat-` — full heartbeat suite sweep. - `cd server && pnpm run typecheck`. ## Risks - Behavioral shift: shared-workspace runs that used to start immediately now wait for the workspace to free. Against a long-running live holder the wait is unbounded by design — the alternative is dispatching into a held working tree, which is the corruption this PR removes. Every deferral is visible in the run timeline (lifecycle event with the holder run, issue, and attempt number), and the wake is parked, never dropped. - A zombie holder (a `running` row whose process died) delays contending runs by up to the 1 h staleness threshold before it stops counting. Recovery's silent-run escalation targets the same run on the same clock, so this window matches what the system already tolerates for silent active runs. - The holder check and the dispatch are not atomic; two runs that pass the gate in the same instant can still race. The gate closes the common window (a second run waking while the first is mid-execution); the pre-existing sync-conflict handling remains the backstop for the rare simultaneous start. - No schema change, no API change, no new configuration. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, tool use, full repository access; implementation, tests, and verification runs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
6008482aab |
fix(decisions): remove clock race from sweep-expiry tests (#10701)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The decisions service lets agents propose decisions with a TTL, and a sweep expires them. > - Four sweep tests create decisions that expire 5 ms in the future, then race the service's own clock read. > - On a loaded CI runner more than 5 ms routinely elapse before validation, so `create` itself rejects the decision and the test fails on unrelated PRs. > - This pull request makes the expiry deterministic: create with a comfortable future TTL, then move `expiresAt` into the past directly in the store. > - The benefit is that the decisions suite stops failing intermittently and stops blocking unrelated PRs. ## Linked Issues or Issue Description No open issue exists; the defect is described here following the bug template. **What happened?** `decisions-service.test.ts` fails intermittently in CI with `expiresAt must be within 30 days` in `bounds expiration work to the configured batch size` and `falls back to the default sweep batch size for invalid configuration`. The failure hits unrelated PRs — for example the `PR` workflow runs for #10699 failed three times on this suite while the same suite passes locally. **Steps to reproduce** 1. Run `pnpm vitest run src/__tests__/decisions-service.test.ts` on a machine under load (or add a ~10 ms delay inside `decisionService.create` before the expiry validation). 2. The test builds `expiresAt: new Date(Date.now() + 5)`; by the time `create` validates, `expiresAt.getTime() <= Date.now()` is true. 3. `create` throws `expiresAt must be within 30 days` (the past-expiry branch of the validator) and the test fails before the sweep runs. **Expected behavior** The sweep tests exercise expiry deterministically and never depend on fewer than 5 ms elapsing between two clock reads in different modules. **Paperclip version** master (`717684ad8f`); the tests landed with the decisions desk workflow in #10672. **Deployment mode** Not deployment-specific — CI and local test runs. ## What Changed - `server/src/__tests__/decisions-service.test.ts`: added two helpers — `nearFutureExpiry()` (a 60 s TTL that passes validation with a wide margin) and `expireDecisionNow(id)` (moves the stored `expiresAt` into the past). The four affected tests create decisions with the future TTL, force-expire them through the store, and drop the 10 ms sleeps. The sweep observes the same expired state as before with no scheduler-timing dependence. ## Verification - `cd server && pnpm vitest run src/__tests__/decisions-service.test.ts` — five consecutive local runs, 31/31 passing each. - No production code changed; the diff is test-only. ## Risks - Low risk: test-only change. The force-expire helper writes the store directly, which is the same technique other TTL suites use to avoid sleeping through real time. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, tool use; diagnosis, fix, and verification runs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2ffebd4836 |
Test adapters in the environment a run would actually use (#10698)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents execute in environments — local, SSH, or sandboxes — resolved
at run time as agent environment → instance default → local
> - The Configuration page has a Test button that probes the adapter
(working directory, command, a model call) in the environment it will
run in
> - But the Test sent only the agent's own environment id, with no
instance-default fallback, so agents relying on the instance default
were probed on the Paperclip host instead
> - A sandbox image carrying an extra CLI then fails the Test with
"command not found" even though every real run would resolve to the
sandbox and succeed — the Test lies about a working setup
> - This pull request mirrors the run-time resolution in the Test call
via a small shared helper with tests
> - The benefit is that the Test button reports the truth about where
the agent actually runs
## Linked Issues or Issue Description
No existing public issue — inline description following the bug report
template:
**What happened?**
With the instance default environment set to a sandbox (whose image
includes the adapter CLI) and an agent that leaves its environment unset
("use instance default"), the Configuration page's Test fails with
`command not found` for that CLI.
**Expected behavior**
The Test probes the environment a real run would use — here the
instance-default sandbox, where the CLI exists — and passes.
**Steps to reproduce**
1. Set the instance default environment to a sandbox whose image carries
an adapter CLI not installed on the Paperclip host (e.g. `grok`).
2. Create a `grok_local` agent without selecting an environment.
3. Press Test on the agent's Configuration page → `command not found`,
while a real heartbeat run resolves to the sandbox and works.
**Paperclip version or commit**
Reproduced on `sha-53bcf38-cloud`-era master; root-caused in
`ui/src/components/AgentConfigForm.tsx` (`environmentId =
currentDefaultEnvironmentId || null`) versus the server's
`resolveExecutionWorkspaceEnvironmentId` (agent → instance default →
local).
## What Changed
- New `ui/src/lib/adapter-test-environment.ts`:
`resolveAdapterTestEnvironmentId` — agent environment first, else
instance default, else null (host probe) — documented as the mirror of
the server's run-time resolution.
- `AgentConfigForm` uses it in the Test mutation. The raw agent
environment id is now sent even when it points at the local environment:
the server already resolves the driver and probes the host for local, so
explicit-local behavior is unchanged, and the test-environment route's
remote paths (SSH/sandbox lease + custom-image template) engage exactly
as they do for the fallback environment.
- Tests pin the fallback (agent wins; instance default when agent unset;
null when neither).
Deliberately untouched: the onboarding wizard's adapter test still sends
no environment — during onboarding an instance default frequently
doesn't exist yet, and changing that flow deserves its own look.
## Verification
- `vitest run` on the new helper suite plus both `AgentConfigForm`
suites — 19 tests pass; `tsc` clean in `ui/`.
- Root cause verified against a live deployment: an agent with
`default_environment_id = NULL`, instance default = sandbox environment;
the Test posted `environmentId: null` and probed the host (no `Probing
inside environment: …` check in the result), which lacks the CLI that
the sandbox image carries.
## Risks
- Low. The change only widens which environment the Test probes,
matching run-time reality. Sandbox-backed tests boot a throwaway sandbox
(existing route behavior — lease, custom-image template,
archive-on-release), so Tests for instance-default-sandbox agents now
take sandbox-boot time instead of failing fast and wrongly.
## Model Used
Claude Fable 5 (`claude-fable-5`, extended thinking, via Claude Code
with tool use and code execution); diagnosis included live inspection of
a deployed instance's agent/environment configuration.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
(helper doc-comment carries the rationale; no user-facing doc covers the
Test button)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
8540ce2973 |
ci: shard general-server tests 4 ways (#10663)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request CI must give contributors fast and stable feedback. > - The `general-server` Vitest lane runs many single-worker server suites. > - A recent completed PR run showed this lane as the slowest completed check. > - Three shards still left one runner with the largest share of work. > - This pull request splits that lane into four duration-balanced shards. > - The benefit is a shorter critical path for the same server test coverage. ## Linked Issues or Issue Description No public GitHub issue exists for this CI maintenance change. **Pre-submission checklist** - I confirmed this improves existing behavior. It does not add a new command, endpoint, or concept. - I searched open public issues and pull requests for related CI sharding work. **What existing behavior does this improve?** The pull request workflow's `general-server` Vitest lane. **Subsystem affected** Cross-cutting. This affects GitHub Actions CI and the Vitest shard duration manifest. **Current behavior** The `general-server` lane uses three shards. The server suites now total about 880 seconds of serial Vitest wall time. The slowest shard was about 313 seconds in the measured run. **Proposed behavior** The `general-server` lane uses four shards. Each shard receives about 220 seconds of predicted suite weight from the refreshed duration manifest. **Reason and benefit** The slowest PR check controls how soon a reviewer can trust the PR. Four balanced shards reduce the slowest `general-server` shard while keeping the same suite selection rules. **Breaking changes** None. This only changes CI partitioning and duration data for existing test suites. **Additional context** Related public searches found no exact open issue or pull request for this `general-server` sharding change. ## What Changed - Split the `general-server` CI matrix from three shards to four shards. - Refreshed `scripts/general-server-shard-durations.json` with wall-time weights from a recent completed PR run. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check origin/master...HEAD` - Dry-ran the four `general-server` shards locally during implementation. The partition covers 300 unique suites with about 220.56 seconds of predicted weight per shard. - Ran a local sensitive-data scan before push. It found only test filenames that contain words such as `secret` or `token`, not credential values. ## Risks Low risk. The main risk is that the duration manifest becomes stale as suite costs move. Missing suites fall back to the median weight, so the lane still runs if the manifest is incomplete. ## Model Used OpenAI Codex, GPT-5, with tool use and local command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e4b0152ca3 |
fix(server): keep the agent invokable when a run fails on a workspace sync conflict (#10660)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents on the same project can share one `shared_workspace` clone, and each sandbox run reconciles its git history back into it > - When histories written by different runs genuinely diverge, reconciliation can hit a real merge conflict — a property of the *workspace*, not of whichever agent happened to run last > - That agent nonetheless finalized into a sticky `error` state, removing a healthy agent from rotation while the workspace stayed broken — and on a shared workspace this serially knocks out every agent that touches it > - This pull request classifies workspace-reconciliation failure signatures as workspace-scoped, so the run still fails with the full message but the agent stays invokable > - The benefit is that one bad workspace state no longer disables agents one by one ## Linked Issues or Issue Description Refs #10645 — this addresses the sticky-agent-error clause of that issue. Workspace run serialization / per-agent worktrees remain tracked there (design sketch on the issue). ## What Changed - New exported `isWorkspaceSyncConflictFailure(message)` matching the reconciliation failure signatures: `merge-tree` conflict ("Failed to merge concurrent remote git histories"), integrate-retry exhaustion ("Failed to integrate concurrent remote git history"), and bundle prerequisite failures ("did not send all necessary objects", "lacks these prerequisite commits"). - Both run-failure finalization paths (adapter returned a failed result; adapter threw) pass `keepIdleOnFailure` for these signatures — the same mechanism already used for provider-quota failures — so the agent finalizes to `idle` instead of `error`. The run itself still fails and carries the full message; nothing about run reporting changes. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-workspace-session.test.ts` — new signature matrix (4 positive signatures, negatives for unrelated adapter failures and null/empty); 121 tests total. - `cd server && pnpm run typecheck`. ## Risks - Low. The only change is which failure families put the agent into `error`; behavior for every other failure is untouched. A workspace stuck in conflict still fails every run against it (visible on the runs surface) — it just no longer takes agents down with it. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
592cade5a6 |
feat(server): trigger the push-capability preflight from the issue's stated PR deliverable (#10659)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs that must push to GitHub have a pre-dispatch credential preflight (`push_write_credential_missing`) so the missing token surfaces as a configuration-incomplete blocker instead of a late runtime failure > - The preflight only triggers when the issue mentions the GitHub PR workflow *skill* — routine-created issues and agent-to-agent handoffs rarely do, even when their text literally says "push the branch and open a PR" > - In practice the credential gap then surfaced only after implementation and review were complete, stranding finished work > - This pull request adds a conservative, verb-anchored text heuristic over the issue title and description as a second preflight trigger > - The benefit is that the credential ask reaches the human before any work is burned ## Linked Issues or Issue Description Fixes #10644 (completes the prevention set with #10648, #10650, #10658) ## What Changed - `issueTextImpliesPrDeliverable(text)`: matches verb-anchored deliverable statements — "open/create/raise/submit a (draft) pull request/PR", "push … branch/remote/origin/upstream". Verb anchoring deliberately ignores passing mentions ("the PR merged yesterday", "PR feedback addressed"). - `requiresPushCapabilityPreflight` takes the issue's title+description and ORs the text heuristic with the existing skill-mention trigger; adapter-type and issue gating are unchanged. The run-dispatch call site threads the already-loaded issue text — no extra query. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-workspace-session.test.ts` — new heuristic matrix (4 positive, 6 negative including null/empty) and preflight-by-text cases (text triggers, passing mention does not, no issue → no preflight); 122 tests total. - `cd server && pnpm run typecheck`. ## Risks - A false positive turns into a configuration-incomplete blocker asking for a GitHub token on an issue that didn't need one — the heuristic is intentionally conservative (verb-anchored) to keep that rare, and the blocker names the exact remediation. - No behavior change for issues that neither mention the skill nor state a PR deliverable. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ddbcf53e31 |
fix(server): refuse agent delegation cycles back to an open ancestor's creator (#10658)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents decompose work by creating child issues assigned to other agents > - When two agents each lack a capability the other assumed (e.g. neither can push to GitHub), each can "resolve" its blocker by delegating the same step to the other: A creates a child for B, B creates a grandchild back for A > - Nothing detects the cycle; the chain of blocked issues grows and no signal reaches the human who could actually fix the capability gap > - This pull request refuses agent-initiated child creation when the child's assignee is the creator of a still-open ancestor in the same chain — a mechanical, semantics-free cycle signal > - The benefit is that the hot-potato dies at creation time with an actionable error instead of growing a dead chain ## Linked Issues or Issue Description Fixes #10642 (write-time counterpart: #10648 refuses assignment to paused agents; the credential-gap *preflight* side is tracked separately in #10644) ## What Changed - `issueService.findOpenAncestorCreatedByAgent(parentIssueId, agentId, {maxDepth})`: bounded walk up the parent chain looking for a still-open (not done/cancelled) ancestor created by the given agent. - Agent-initiated issue creation with a parent (both the create-with-`parentId` route and `POST /issues/:id/children`) now refuses with a structured 409 (`code: delegation_cycle`, naming the ancestor) when the new child would be assigned to the agent that created a still-open ancestor: that agent delegated the work into this chain, so assigning it back is a cycle. The message states the alternatives — complete the work, leave the child unassigned, or escalate to a board operator. - Deliberately unaffected: human actors (deliberate re-routing is their call), closed ancestors (re-engaging the creator of finished work is normal), and accepted-plan decomposition (its children come from a human-approved plan). ## Verification - `pnpm vitest run server/src/__tests__/issue-assignee-invokability-routes.test.ts` — cycle refused with 409 and no create call; the same child allowed when no open ancestor matches; board actors never consult the guard. - `pnpm vitest run server/src/__tests__/issues-service.test.ts` — new embedded-Postgres coverage: ancestor found through the chain, closed ancestors ignored, depth bound honored (114 total). - `pnpm vitest run server/src/__tests__/issue-create-deduplication-routes.test.ts server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts` — unchanged (79). - `cd server && pnpm run typecheck`. ## Risks - Low-to-moderate: a new 409 for a creation shape that previously succeeded. The blocked shape (agent assigns new work to the creator of an open ancestor) is the cycle signature; the legitimate "hand a subtask to the parent's assignee" pattern is unaffected because it keys on assignee, not creator. Watchdog and plan-decomposition flows are exempt or unaffected as described. - The walk adds at most `maxDepth` (10) single-row lookups per agent child creation with an assignee. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
3dec88ce90 |
feat(agents): warn when an agent's escalation path routes to a paused manager (#10657)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents escalate work up the org chart (`reports_to`), and operators pause agents — notably, instance imports pause every agent by default > - A paused manager does not invalidate the chain (subordinates stay invokable), so nothing surfaces when an operator unpauses workers but leaves their manager paused > - Escalations then dead-letter silently: agent-created issues assigned to the paused manager sit in a queue nothing will ever run > - This pull request computes paused ancestors in the existing org-chain health model and surfaces a non-blocking warning on the agent read models and detail page > - The benefit is that the operator learns their escalation paths are dead before work vanishes into them ## Linked Issues or Issue Description Fixes #10647 (companion to #10648, which refuses agent-initiated assignment to paused agents at write time — this PR makes the standing hazard visible) ## What Changed - `AgentOrgChainHealth` gains two additive, optional fields: `pausedAncestors` (paused agents in the `reports_to` chain) and `escalationWarning` (human-readable, only set when the agent itself can work — a paused/terminated agent's escalation path is moot). Chain validity, invokability, and assignability are byte-identical. - No server route changes needed: the fields flow through every existing agent read model (list, detail, org chart) since they ride the same `getAgentWorkEligibility` computation. - Agent detail page shows an amber "Escalation path is paused" banner (same visual language as the invalid-chain banner, but non-blocking) with the warning text naming the paused manager and the two remedies. ## Verification - `pnpm vitest run packages/shared/src/agent-eligibility.test.ts` — 5 new cases: paused direct manager warns; paused grandparent through a healthy manager warns; the agent itself paused → no warning (but ancestors still reported); fully active chain → no warning, empty list; terminated ancestor keeps the invalid-chain classification without double-counting as paused. - Full `@paperclipai/shared` suite (392 tests) and `agent-eligibility-routes` (54) unchanged. - `tsc --noEmit` in shared, server, and ui. ## Risks - Low. Purely additive fields plus one UI banner; no behavior gates on the new data. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a0dbe21045 |
fix(server): stand down recovery while an operator-cancelled run is the latest activity (#10656)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Operators can cancel a running agent from the board when a run is unwanted — most acutely while cleaning up a runaway loop > - The recovery machinery treats a cancelled run like any other unsuccessful terminal run: the stranded-issue sweep classifies the issue as stranded, creates a recovery action, and wakes the agent again > - So cancelling runs to stop a loop *fed* the loop: each operator cancel spawned a recovery action that re-woke the agent the operator had just stopped > - This pull request stamps board-initiated cancellations with operator attribution and makes the sweep stand down while such a run is the issue's latest activity > - The benefit is that an operator's cancel is final until something new happens, instead of being fought by automation ## Linked Issues or Issue Description Fixes #10646 ## What Changed - `POST /heartbeat-runs/:runId/cancel` (board-only) now cancels with an explicit reason ("Cancelled by a board operator") and stamps `resultJson.cancelledByActorType: "user"` / `cancelledByUserId`. - `reconcileStrandedAssignedIssues` gains an early stand-down: when the issue's latest run is operator-cancelled (the new stamp, or the existing `operator_interrupted` error code from interrupt-by-comment), the issue is skipped entirely — no recovery action, no wake — and counted in a new `operatorCancelExempted` result field. The exemption is inherently self-limiting: any newer run or wake supersedes it because the gate only looks at the *latest* run. - System cancellations without operator attribution (lease expiry, assignee changes, terminal-status cancels, pause holds) keep today's recovery behavior unchanged. ## Verification - `pnpm vitest run server/src/__tests__/issue-recovery-actions.test.ts` (embedded Postgres) — 3 new cases: a stamped operator cancel produces zero recovery actions and zero wakes; an `operator_interrupted` cancel likewise; an unattributed system cancel still flows into pre-existing recovery (wake observed), proving the stand-down is scoped to operator attribution. - `pnpm vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/issue-scheduled-retry-routes.test.ts` — unchanged (109 tests). - `cd server && pnpm run typecheck`. ## Risks - Low. The only suppressed behavior is recovery of runs a human explicitly cancelled from the board; everything else is byte-identical. If an operator cancels and walks away, the issue stays quiet until any new activity — which is the intent (the operator owns the next step). ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0f12721ee9 |
fix(server): let board users cancel issues with an active review stage (#10655)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server enforces an issue execution policy. It gates status changes while a review or approval stage is active. > - A board user could not cancel a task while an agent reviewer held the active stage. The API returned "Only the active reviewer or approver can advance the current execution stage". > - Board users own the board. They must always be able to edit and cancel any task. > - This pull request adds a board override to the execution stage transition. A board cancel clears the pending stage state and proceeds instead of raising an error. > - The benefit is that board users can always stop work, even while a review is pending or the stored stage state has drifted. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Description follows `bug_report.yml`: **What happened?** A board user set a task to `cancelled` while the task had an active reviewer stage held by an agent. The PATCH failed with "Task Update Failed... only the assigned approver or reviewer...". The same failure occurred when the stored stage state had drifted: the server silently forced the task back to `in_review` instead of honoring the cancel. **Expected behavior** A board user can always edit and cancel any task. A board cancel must clear the pending review stage and apply the requested status. **Steps to reproduce** 1. Create a task assigned to agent A with agent B configured as reviewer in the execution policy. 2. Let agent A hand the task off so the review stage becomes active. 3. As the board user, set the task status to Cancelled. 4. The update fails with the reviewer-only error. Related work: #5487 touches the execution-policy approver UI. It does not address the board cancel path. ## What Changed - `server/src/routes/issues.ts`: the issue PATCH route now passes `allowBoardOverride` when the actor is a board user. - `server/src/services/issue-execution-policy.ts`: when `allowBoardOverride` is set and the requested status is not `in_review` or `in_progress`, the transition clears `executionState` and proceeds. This applies both while a stage decision is pending and when the stage state has drifted, so a board cancel is no longer rejected or silently flipped back to `in_review`. - Reviewer gating is unchanged for everyone else: a board user who is the active participant still uses the normal approve / request-changes flow, and non-participant agents still receive the 422 guard. - Assignee-only board updates on an `in_review` task keep the stage state coherent: reassigning to an eligible stage participant re-pends the stage with them as the current participant, while reassigning to a non-participant (or unassigning) dissolves the review back to `in_progress` instead of persisting an `in_review` issue with no execution state or an ineligible participant. - New unit tests and route tests cover board cancellation of an active review stage and of a drifted pending review, plus reviewer swap, non-participant reassignment, and unassignment during an active review. ## Verification - In `server/`: `pnpm exec vitest run src/__tests__/issue-execution-policy.test.ts src/__tests__/issue-execution-policy-routes.test.ts` — 2 files, 73/73 tests pass on top of current `master`. - In `server/`: `pnpm run typecheck` passes. ## Risks - Low risk. The override branch runs only for board actors and only for target statuses other than `in_review` and `in_progress`. Cancelling clears `executionState`, so a later reopen starts from a fresh stage state. Agent-facing flows and reviewer gating are unchanged. ## Model Used - Claude Fable 5 (Anthropic), model ID `claude-fable-5`, running in Claude Code (Claude Agent SDK) with extended thinking and agentic tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ada47be764 |
fix(server): refuse agent-initiated issue assignment to paused agents (#10648)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can create and assign issues to other agents, and commonly escalate to their org-chart manager (`reports_to`) when they hit something outside their authority > - Issue assignment already refuses terminated and pending-approval assignees, but accepts paused assignees from any actor > - A paused agent never runs, so agent-initiated escalations to a paused manager become invisible dead letters — accepted silently, never picked up, never surfaced > - This pull request refuses paused assignees when the assigning actor is an agent, at the single normalization helper all four assignment paths flow through > - The benefit is that agent-routed work can no longer silently vanish into a paused agent's queue ## Linked Issues or Issue Description Fixes #10641 ## What Changed - `normalizeIssueAssigneeAgentReference` (used by issue create, both child-create routes, and issue update) now throws a 409 when an **agent** actor assigns to a **paused** agent, with a message naming the alternatives: assign an invokable agent, leave the issue unassigned, or escalate to a board operator. - Board/user actors are unchanged and may still assign to paused agents deliberately — the pause state is visible in the UI, and staging work for a later unpause is a legitimate workflow. Terminated / pending-approval / invalid-org-chain refusals are unchanged for all actors. - This matches the existing precedent for watchdogs ("Cannot assign watchdog to an agent that is not invokable") using the same conflict-error shape. ## Verification - `pnpm vitest run server/src/__tests__/issue-assignee-invokability-routes.test.ts` — new coverage: agent PATCH → paused assignee 409 (no update call), agent child-create → paused assignee 409 (no create call), agent assignment to an invokable agent still 200, board assignment to a paused agent still 200. - Neighboring suites unchanged: `issue-update-comment-wakeup-routes`, `issue-agent-mutation-ownership-routes`, `issue-create-deduplication-routes`, `issue-watchdogs-routes` (97 tests). - `cd server && pnpm run typecheck`. ## Risks - Low. The only behavior change is a new 409 for agent actors assigning to paused agents — previously a silent dead-letter. Agents that relied on this (escalation flows) now get an actionable error instead; human workflows are untouched. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
27f8c8dbcf |
feat(server): cap agent review rounds and escalate exhausted reviews to the responsible human (#10650)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution policies let one agent implement and another review, cycling through changes-requested → addressed rounds > - Nothing bounds that cycle: no round counter, no escalation, no termination signal — two agents can ping-pong indefinitely, especially when the review's success criteria drift to something the implementer cannot satisfy > - On a real multi-agent instance this produced 6+ unattended rounds (~8 runs) that continued even after the human had merged the PR under review > - This pull request counts consecutive agent-initiated changes-requested rounds and, at a configurable cap, hands the still-pending review to the responsible human instead of bouncing back to the implementer > - The benefit is that unattended review loops terminate in a human decision instead of burning runs forever ## Linked Issues or Issue Description Fixes #10643 ## What Changed - `IssueExecutionState.changesRequestedCount` (schema + type, default 0): consecutive agent-initiated changes-requested rounds on the current stage. Carries through executor resubmissions, resets to 0 on approval, and resets when a **human** makes the changes-requested decision — the cap targets unattended agent↔agent ping-pong, never human review. - `IssueExecutionPolicy.maxReviewRounds` (optional, 1–50, default null → server default `DEFAULT_MAX_REVIEW_ROUNDS = 3`). - At the cap, the transition records the reviewer's changes-requested decision as usual but keeps the stage **pending** with the responsible human (`responsibleUserId`, falling back to `createdByUserId`) as the participant: the issue is assigned to that human and the pending review surfaces through the existing attention/review UI. The human then approves, requests changes (resetting the counter and handing back to the implementer), or re-scopes. - The escalated hold is sticky: transitions from anyone other than the escalated human no longer re-select a configured agent participant for the stage (which would have silently undone the escalation on the next unrelated PATCH). The escalated human's own decisions flow through the normal participant decision branch. - Issues with no responsible human keep today's hand-back behavior; the counter still accumulates so operators can see the churn. ## Verification - `pnpm vitest run server/src/__tests__/issue-execution-policy.test.ts` — 8 new cases: round counting on hand-back, count carried through resubmission, escalation at the default cap, sticky hold across unrelated transitions, human changes-requested resets the counter, human approval completes the stage, no-responsible-human fallback, and a `maxReviewRounds: 1` policy override. - `pnpm vitest run server/src/__tests__/issue-execution-policy-routes.test.ts` and the full `@paperclipai/shared` suite (387 tests) — schema additions are backward compatible (both fields optional with defaults; persisted states without the counter parse as 0). - `pnpm --filter @paperclipai/shared exec tsc --noEmit` and `cd server && pnpm run typecheck`. ## Risks - Behavior change: an agent-only review loop that previously ran forever now escalates to a human after 3 agent rounds by default. Instances that want longer loops can set `maxReviewRounds` per policy. Flows where a human participates are unaffected (human decisions reset the counter). - Escalation requires a `responsibleUserId`/`createdByUserId` on the issue; without one, behavior is unchanged. - Persisted execution states from before this change parse with `changesRequestedCount: 0` — no migration needed. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8444e5735c |
ci: pass e2e shard specs without separator (#10640)
## Thinking Path > - Paperclip uses pull request CI to test changes before merge. > - The e2e PR lane runs Playwright specs in a shard matrix. > - Each shard builds a list of spec files for its matrix entry. > - The workflow passed that list after a literal `--` separator. > - Playwright did not receive the list as file filters. > - This pull request removes the separator and adds a guard test. > - The benefit is that each e2e shard runs only its assigned specs. ## Linked Issues or Issue Description Refs #10629. **What happened?** The e2e shard step used `pnpm run test:e2e -- $specs`. The shard spec list was not applied as Playwright file filters. **Expected behavior** Each e2e shard should pass only its selected specs to Playwright. **Steps to reproduce** 1. Inspect `.github/workflows/pr.yml` at the merge commit for #10629. 2. Find the `e2e_shards` command that invokes `pnpm run test:e2e`. 3. See the literal `--` before `$specs`. **Paperclip version or commit** `86767951` **Deployment mode** GitHub Actions PR CI. ## What Changed - Removed the literal `--` from the e2e shard `pnpm run test:e2e $specs` invocation. - Added a regression test that checks the workflow passes `$specs` without that separator. ## Verification - `node --test scripts/__tests__/e2e-shard.test.mjs` ## Risks Low risk. This changes one CI command and one workflow guard test. The main risk is shell argument handling in the workflow, and the guard now covers the expected command shape. ## Model Used OpenAI GPT-5 through Codex. The run used shell and GitHub CLI tool access. The runtime did not expose a context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8676795188 |
ci: split e2e PR lane into three shards (#10629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The pull request workflow protects changes with a Playwright e2e lane. > - That lane already uses a weighted file partition so slow specs do not cluster by test count. > - Recent green PR runs showed the two e2e shard jobs were slower than the next slow required lane. > - The largest spec is indivisible, so a third shard lets that spec run alone and lets the rest split by duration. > - This pull request changes only the PR e2e shard matrix and the guard test. > - The benefit is a shorter expected PR critical path while the required `e2e` aggregate check name stays stable. ## Linked Issues or Issue Description Refs #9923 **What existing behavior does this improve?** The `pull_request` workflow Playwright e2e lane. **Subsystem affected** Cross-cutting: GitHub Actions CI and test scripts. **Current behavior** The PR workflow runs the weighted Playwright e2e partition across two jobs. Recent green runs showed those jobs as the slowest required checks. **Proposed behavior** The PR workflow runs the same e2e spec set across three weighted jobs. The aggregate required check stays named `e2e`. **Reason and benefit** The third shard lets the slow smoke-lab spec run alone while the rest of the catalog stays balanced. This should shorten the PR critical path. The win is bounded by fixed per-job setup time. **Breaking changes** None. The required aggregate check contract is preserved. ## What Changed - Change the PR e2e shard matrix from two entries to three entries. - Update the shard guard test to expect three shards. - Floor the balance bound at the largest single spec weight. - Assert that the workflow does not define more shard indexes than `SHARD_COUNT`. ## Verification - `node --test ./scripts/__tests__/e2e-shard.test.mjs` passes with 6 tests. - The recorded-weight partition is complete and non-overlapping: 168.0s, 116.5s, and 114.4s. - I checked `ROADMAP.md` and found no overlapping roadmap-level core feature. - I searched public GitHub PRs and issues for related e2e shard work. I found related PR #9923 and no open duplicate for this branch or change. ## Risks - This adds one extra GitHub Actions runner to the PR e2e lane. - The wall-clock win is bounded by fixed per-job setup. - Behavior risk is low because the aggregate required check remains named `e2e`. ## Model Used OpenAI Codex, GPT-5, tool-enabled coding agent in this repository. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> |
||
|
|
6401f4f78c |
fix(sandbox): bundle git copy-back against the merge-base so diverged/reset workspaces still import (#10601)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - An agent that runs in a sandbox has its workspace copied back to the host when the run ends, so its work persists — the copy-back ships a git bundle of the sandbox's commits > - The bundle is created as a thin delta, `git bundle create HEAD --not <baseSha>`, which records `baseSha` (the host workspace HEAD captured at export) as a prerequisite the host must already hold > - That assumption breaks when the sandbox HEAD has diverged from `baseSha`, or when a shared host workspace no longer holds `baseSha` at import time — then `git fetch` on the host hard-fails and the entire run is lost even though the agent finished its work > - This pull request bundles against the merge-base of `baseSha` and the sandbox HEAD (with a full-bundle fallback), which the host can satisfy in those cases > - The benefit is that copy-back no longer discards a completed run's work over a base the host can't reconcile ## Linked Issues or Issue Description **What happened?** A sandbox agent run completed its work, then failed during workspace finalize: ``` git -C <host workspace> fetch --force <git-delta.bundle> refs/…/export:refs/…/imported error: Could not read <baseSha> fatal: revision walk setup failed error: git-delta.bundle did not send all necessary objects ``` The run is reported as `adapter_failed` even though the agent produced output. The copy-back bundle names the host workspace's recorded HEAD (`baseSha`) as a prerequisite, but the host cannot satisfy it. **Steps to reproduce** Two independent triggers, both reproduced in tests: 1. The sandbox's HEAD has diverged from `baseSha` — e.g. the sandbox carries a local-only branch that forked from an older commit than the host's current HEAD. 2. The shared host workspace no longer holds `baseSha` at import time (it was reset / re-realized between export and import). In either case `git fetch` of the thin bundle fails with a missing prerequisite. **Expected behavior** Copy-back imports the sandbox's work as long as the host holds any common ancestor, instead of hard-failing and discarding the run. **Paperclip version** Current `master`. **Deployment mode** Any deployment running agents in sandbox environments with workspace sync (notably shared-workspace clones and custom images that carry a local-ahead branch). ## What Changed - `buildRemoteGitDeltaBundleScript` now computes `bundle_base = git merge-base <baseSha> HEAD` and bundles `HEAD --not <bundle_base>`. The merge-base is an ancestor of `baseSha`, so any host that holds `baseSha` (or an ancestor of it — e.g. after a reset) can satisfy the prerequisite, and the bundle stays a delta rather than a full-history transfer. - When `baseSha` is absent from the sandbox, or no merge-base exists, it falls back to a full, self-contained bundle (no prerequisites) so the import can always complete. - The existing empty-bundle no-op (no new commits) and the ordinary fast-forward path are unchanged; the `cat-file` base check no longer aborts the script under `set -e`. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — new cases: a diverged sandbox HEAD imports when the host holds only the merge-base (not `baseSha`), and the full-bundle fallback imports into a host that shares no history; existing thin-delta and empty-bundle cases still pass. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`. - Standalone shell repro confirmed the old thin bundle fails with "Repository lacks these prerequisite commits" in both trigger cases, and the merge-base bundle imports successfully. ## Risks - Low. For the common case (sandbox HEAD descends from `baseSha`) the merge-base is `baseSha`, so the bundle is byte-for-byte the same delta as before. The change only alters behavior when the old code would have hard-failed. - This makes the copy-back import succeed on a diverged base; the subsequent reconciliation of divergent histories (`integrateImportedGitHead`) is unchanged and still owns how the imported head is merged into the host branch. Where a workspace's history has genuinely diverged (e.g. a stale custom image carrying a local-only branch), a clean re-clone/re-capture is still the right operational fix — this change prevents work loss, it does not reconcile intentional divergence. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, shell repro, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ee9d907d01 |
fix(codex): do not inject a duplicate --skip-git-repo-check for sandbox runs (#10595)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A Codex agent runs `codex exec`, and the adapter assembles its argument vector from the agent's config plus execution-context options > - For sandbox execution the adapter injects `--skip-git-repo-check`, because a headless remote workspace has no git trust prompt to answer > - The adapter also appends the operator's `extraArgs` verbatim, so an agent that already lists `--skip-git-repo-check` in its config gets the flag twice on a sandbox run > - `codex exec` rejects a repeated `--skip-git-repo-check` and exits with code 2, which the adapter surfaces as `adapter_failed` before any work runs > - This pull request skips the sandbox injection when the operator's args already carry the flag > - The benefit is that a common, harmless-looking config no longer crashes every sandbox run ## Linked Issues or Issue Description **What happened?** A `codex_local` agent configured with `extraArgs: ["--skip-git-repo-check"]` fails on every sandbox run: ``` error: the argument '--skip-git-repo-check' cannot be used multiple times Usage: codex exec [OPTIONS] [PROMPT] ``` The adapter reports `stopReason: "adapter_failed"` (Codex exited with code 2). The flag appears twice in the argv: once injected by the adapter for sandbox execution, once from the operator's `extraArgs`. **Steps to reproduce** 1. Configure a `codex_local` agent with `extraArgs: ["--skip-git-repo-check"]` (or the legacy `args` field). 2. Point it at a sandbox environment. 3. Start a run — `codex exec` aborts immediately on the duplicate flag. **Expected behavior** The run launches with a single `--skip-git-repo-check`. An operator listing the flag the adapter already injects should be a no-op, not a hard failure. **Paperclip version** Current `master`. **Deployment mode** Any deployment running Codex agents in sandbox environments. ## What Changed - `buildCodexExecArgs` no longer pushes the sandbox `--skip-git-repo-check` when the resolved args (`extraArgs`, or the legacy `args` fallback) already contain it. The operator's copy stands; the argv carries the flag exactly once. Non-sandbox runs and configs without the flag are unchanged. ## Verification - `cd packages/adapters/codex-local && pnpm vitest run src/server/codex-args.test.ts` — new cases: `extraArgs` already carrying the flag (single occurrence), the legacy `args` field carrying it (single occurrence), and the operator's flag preserved when the sandbox injection is not requested. Existing "adds --skip-git-repo-check when requested" case unchanged. - `cd packages/adapters/codex-local && pnpm vitest run` — full package suite (218 tests). - `pnpm run typecheck` in the package. ## Risks - Low. The change only suppresses a duplicate of a single, idempotent flag; it never removes an operator-supplied argument and never adds one that was not already going to be present. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c0b875c46c |
fix(codex): let sandbox runs use the sandbox image's own Codex login (#10582)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Codex agents can run inside sandbox environments, and operators can bake a Codex login into the sandbox image during interactive image setup > - Two credential gates (the control plane's pre-dispatch configuration-incomplete gate and the adapter's execute-time fail-fast) required host-side Codex credentials — a usable `auth.json` in the managed home or a configured `OPENAI_API_KEY` — regardless of where the run executes > - On managed cloud hosts a local Codex login never exists, so every sandbox run of a Codex agent failed immediately with "configuration incomplete: no Codex credentials available for managed home …", even though the adapter's inbound auth merge already supports the image-login case end to end > - This pull request makes the execute-time gate probe the sandbox for its own `~/.codex/auth.json` before failing, and exempts sandbox-destined runs from the pre-dispatch host check > - The benefit is that a sandbox image signed in to Codex is a first-class credential source, matching what the auth-merge, precedence-warning, and copy-back machinery were already built for ## Linked Issues or Issue Description **What happened?** Running a `codex_local` agent in a sandbox environment whose image carries a Codex login failed instantly with `configuration incomplete: no Codex credentials available for managed home "…/codex-home". Sign in to Codex on the host with a ChatGPT subscription, or bind a per-agent OPENAI_API_KEY secret for this agent.` The host has no Codex login and never will on a managed cloud deployment; the sandbox's own login was never consulted. **Steps to reproduce** 1. Configure a sandbox environment and capture a custom image after signing in to Codex inside the interactive image setup. 2. Create a `codex_local` agent that uses that environment, on a host with no Codex login and no `OPENAI_API_KEY` bound. 3. Start a run: it fails pre-dispatch with the configuration-incomplete blocker above. **Expected behavior** The run launches and Codex authenticates with the sandbox image's own login, the same way the adapter's host↔sandbox auth merge already keeps the sandbox credential when the host ships none. A run should only fail fast when neither the host, a bound `OPENAI_API_KEY`, nor the sandbox has credentials. **Paperclip version** Current `master` (cloud image deployments). **Deployment mode** Managed cloud stacks (any deployment where the server host has no local Codex login). ## What Changed - Extracted the adapter's execute-time gate into `assertCodexCredentialsLaunchable`: when host readiness fails and the target is a sandbox, it probes `~/.codex/auth.json` in the sandbox (same command the auth-precedence warning uses) and proceeds with a log line naming the credential source; when the sandbox has no login either, the error now names all three remediation options (sandbox image sign-in, per-agent `OPENAI_API_KEY`, host sign-in). Non-sandbox targets keep today's strict behavior byte-for-byte. - The control plane's pre-dispatch gate in `resolveExecutionRunAdapterConfig` now takes the selected environment's driver and skips the host-credential check for sandbox-destined runs — only the adapter can probe the sandbox once it is up, so the execute-time gate is the authority there. Non-sandbox runs keep the early, well-attributed configuration-incomplete blocker. - The codex Test flow needed no change: it already seeds host credentials only when they exist and otherwise leaves the sandbox's `CODEX_HOME` alone; this aligns the run path with it. ## Verification - `cd packages/adapters/codex-local && pnpm vitest run` — 210 tests, including new gate cases: sandbox login present (proceeds + logs source), sandbox and host both credential-less (fails with the extended message), non-sandbox target (strict host requirement kept, no sandbox probe), per-agent API key (no probe at all). - `cd server && pnpm vitest run src/__tests__/heartbeat-project-env.test.ts src/__tests__/codex-local-adapter-environment.test.ts` — includes the new sandbox-exemption case next to the existing blocker tests. - `pnpm run typecheck` in `server` and `packages/adapters/codex-local`. ## Risks - Sandbox-destined misconfigurations (no credentials anywhere) now surface at adapter execute time instead of pre-dispatch, so they read as an adapter failure with a precise message rather than a configuration-incomplete blocker. The trade-off is deliberate: the sandbox must be up to know whether credentials exist, and the failure message names the exact remediations. - The sandbox probe adds one short (5s-capped) shell command to sandbox runs whose host has no credentials; runs with host credentials or a bound key are untouched. - Self-hosted behavior is unchanged for local and SSH targets. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
90ead239a8 |
feat(ui/server): name cross-company environment secret refs instead of calling them missing (#10577)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environment configs (sandbox providers, SSH) can bind stored company secrets through `format: "secret-ref"` fields, picked in the environment editor's secret picker > - Environments are instance-scoped and shared by every company on an instance, but the picker lists only the current company's secrets, so a ref pointing at another company's secret renders as "Missing secret (…)" in destructive styling > - That state is indistinguishable from a genuinely deleted secret, so operators "fix" a healthy binding by creating a duplicate secret in their own company — the exact sequence that used to corrupt bindings before #10576 > - This pull request adds an instance-gated metadata endpoint for an environment's secret refs and teaches the picker to name a cross-company secret and its owner honestly > - The benefit is that operators can tell a healthy cross-company binding from a broken one, and stop creating duplicate secrets ## Linked Issues or Issue Description **Is your feature request related to a problem? Please describe.** In the environment editor, a secret-ref field that points at a secret owned by a different company shows "Missing secret (22095402…)" in red, with "The previously selected secret is no longer available. Pick another or remove the binding." The binding is actually healthy — the current company's picker just cannot list the other company's secrets. Operators react by creating a duplicate secret and re-pointing the field. **Describe the solution you'd like** The editor should know the referenced secret's name, status, and owning company (metadata only, never the value) and present a cross-company ref neutrally, a deleted secret as deleted, and only an unknown id as missing. Related: #10576 (fixes the binding corruption this UI state used to trigger). ## What Changed - New `GET /environments/:id/secret-refs` returns `{ refs: [{ configPath, secretId, name, status, companyId, companyName }] }` for the environment's config-derived secret refs. Values are never returned. The route sits behind `assertCanAccessInstanceEnvironments`, the same gate as environment editing. - New `secretService.describeSecretRefs` loads that metadata across companies; unknown ids are omitted. - `SecretBindingPicker` reads an optional `SecretRefHintsContext` (keyed by secret id). With a hint, a ref the company list cannot show renders as `NAME — Owning Company` with neutral styling and the note "Owned by the … company. The binding keeps working; selecting a secret from this list re-points it here." A hint with `status: "deleted"` reports the secret as deleted. Without hints, behavior is byte-identical to before — agent editors and other picker users are unaffected. - `CompanyEnvironments` fetches descriptors for the environment being edited and provides them through the context. ## Verification - `cd server && pnpm vitest run src/__tests__/environment-routes.test.ts src/__tests__/secrets-service.test.ts` — new endpoint happy path, agent 403 (descriptors never computed), and embedded-Postgres coverage proving cross-company names resolve and unknown ids drop out. - `cd ui && pnpm vitest run src/components/SecretBindingPicker.test.tsx src/components/JsonSchemaForm.test.tsx src/pages/CompanyEnvironments.test.tsx` — hinted cross-company rendering, hinted deleted secret, and unchanged no-hint fallback. - `pnpm run typecheck` in `server` and `ui`. - Manual: edit an environment whose secret-ref field references another company's secret; the field names the secret and its owning company instead of "Missing secret". ## Risks - The endpoint exposes secret names and company names across companies to instance-level environment editors. Those actors already manage instance-shared environments (and instance admins are implicit members of every company), so this reveals no secret material and no new reach; the service method documents that callers must sit behind an instance-level gate. - UI change is additive and context-gated; pickers without a provider render exactly as before. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f51cba33fa |
fix(server): keep environment secret bindings consistent when re-pointing config secrets (#10576)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can run inside environments (SSH boxes, sandbox providers); a sandbox environment's config can reference stored company secrets (for example a provider API key) through `format: "secret-ref"` fields > - Environments are instance-scoped and shared by every company on an instance, but `company_secret_bindings` rows are company-scoped, and the environment routes synced config-derived bindings under one guessed "context company" resolved from the environment's existing bindings > - When a save re-pointed a secret-ref field at a secret owned by a different company, the binding sync threw after the config row had already been persisted: the config referenced the new secret, the binding still pointed at the old one, every later lease acquisition failed with `Secret is not bound to environment:<id> at apiKey`, and the stale cross-company binding made every later save fail with a company-context conflict — with no route-level way to recover > - This pull request makes config-derived bindings follow the company that owns each referenced secret, and makes the environment write and its binding syncs atomic > - The benefit is that environment saves can no longer strand an environment in a half-updated state that breaks all of its runs ## Linked Issues or Issue Description Refs #10577 (companion UX change: the editor state that nudges operators into this sequence). **What happened?** Saving an environment whose secret-ref config field points at a secret owned by a different company than the environment's existing binding partially applied: the config row updated, the binding sync failed server-side, and the environment was left referencing a secret it has no binding for. Every run that leased the environment then failed with `lease_acquire_failed: ... Secret is not bound to environment:<id> at apiKey`, and every later save of the environment returned 409 `Environment secret bindings already use a different company context.` — with no route-level way to recover. **Steps to reproduce** 1. On an instance with two companies, create a sandbox environment from company A with a picker-bound API-key secret owned by A (the binding lands in A). 2. From company B, create a new secret and re-point the environment's API-key field at it, then save. 3. The save persists the config but the binding sync throws, so no binding for B's secret exists. 4. Run any agent that uses the environment, or try to save the environment again. **Expected behavior** The save either fully applies (config and bindings consistent) or fully fails. Re-pointing a config secret ref to a secret owned by another company moves the binding with the secret. **Paperclip version** Reproduced on current `master` (also present on recent release images). **Deployment mode** Multi-company server deployment (any mode with more than one company on the instance). ## What Changed - New `secretService.replaceSecretRefsForInstanceTarget`: writes each config-derived binding under the company that owns the referenced secret, replaces all non-`env.*` bindings of the target across every company, and validates every ref (secret exists, not deleted, config-path and projection-class rules) before any row is written. `env.*` env-var bindings stay company-scoped and untouched. - The environment create and update routes now run the environment write and its binding syncs inside one `db.transaction`, threading the transaction through new optional executor seams on `environmentService.create/update` and the existing `SecretBindingDb` seam pattern, so an invalid ref rolls the whole save back instead of leaving a half-updated environment. - `resolveEnvironmentSecretContextCompanyId` no longer lets existing bindings veto the caller's context (the 409s above); it now only picks where new raw-pasted secrets are created and how env-var bindings and probes resolve: explicit route/query company first, then the single company the bindings live in, then the actor's company. ## Verification - `cd server && pnpm vitest run src/__tests__/environment-routes.test.ts src/__tests__/environment-instance-routes.test.ts src/__tests__/secrets-service.test.ts src/__tests__/environment-custom-image-routes.test.ts` (165 tests, includes new coverage below) - New embedded-Postgres tests prove: a re-point moves the binding to the new secret's company and deletes the stale row; refs across several companies each bind under their own secret's company; an unknown secret ref rejects without touching existing bindings; `env.*` rows survive config-ref replacement. - New route tests prove: a cross-company re-point that previously 409'd now saves, with the update and binding replacement on the same transaction executor; a failing ref surfaces as 422. - `cd server && pnpm run typecheck` ## Risks - Behavioral shift: environment saves no longer 409 on a company-context mismatch between the caller and existing bindings; bindings follow the referenced secret's company instead. Environment routes are instance-admin gated, and instance admins already had access to every company's secrets by passing the company explicitly, so this removes an ordering trap rather than widening access. - Runtime lease resolution is unchanged: a run still resolves environment secrets under the run's own company, so an environment referencing company B's secret still only leases for company B runs (fail-closed as before). - The delete route's per-company binding cleanup is unchanged. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ea0dd3917e |
build: make image layer caching actually hit (#10571)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Published container images are the deployable unit for self-hosted and managed instances, so merge-to-image latency bounds every deploy iteration > - The docker workflow configures BuildKit caching, but builds still ran ~12+ minutes essentially cold > - Two causes: the most expensive layer (four CLI toolchains + apt) is ordered after the always-changing app copy so it can never cache, and the type=gha cache's 10GB repo cap means the two multi-arch mode=max jobs evict each other > - This pull request reorders the tool layer above the app copy (with a weekly epoch so @latest tools keep advancing) and switches both jobs to registry-backed cache in ghcr > - The benefit is that warm builds shrink to roughly the app build + push, targeting the sub-5-minute range together with the amd64-only cloud variant ## Linked Issues or Issue Description No existing public issue — inline description following the feature request template: **Subsystem affected** CI / release publishing (docker workflow, Dockerfile) **Problem or motivation** Despite `cache-from/cache-to` being configured, image builds run effectively cold: (1) the production stage installs four CLI toolchains + apt packages *after* `COPY --from=build /app /app`, and since the app copy changes every commit, that most-expensive layer rebuilds every build, per arch; (2) the `type=gha` BuildKit cache is capped at 10GB per repository, and two multi-arch `mode=max` jobs overflow and evict each other's entries. **Proposed solution** Order the tool/OS layer before the app copy (it references nothing from `/app`), refresh it weekly via a `CLI_TOOLS_CACHE_EPOCH` build arg so the `@latest` tools don't freeze in the cache, and move both jobs to registry-backed BuildKit cache (`:buildcache` / `:buildcache-cloud` refs in ghcr, no size cap, separate refs so the parallel jobs don't clobber each other). **Alternatives considered** Pinning CLI tool versions instead of the weekly epoch — more deterministic, but adds a version-bump chore; the weekly epoch preserves current freshness semantics with bounded staleness. Keeping type=gha with `mode=min` — smaller cache but loses intermediate-stage reuse, which is where most of the win is. **Roadmap alignment** Not on ROADMAP.md; CI/publishing speed improvement only. ## What Changed - `Dockerfile`: the production stage's tool/OS `RUN` (npm --global CLIs, apt, `/paperclip` setup) moves above `COPY --from=build /app /app`; new `CLI_TOOLS_CACHE_EPOCH` arg consumed by that layer. The `cloud` stage is unaffected — it only layers plugin dists on top of the finished production stage. - `.github/workflows/docker.yml`: both jobs stamp the ISO week into `CLI_TOOLS_CACHE_EPOCH`, and both switch `cache-from/cache-to` from `type=gha` to `type=registry` with per-job refs. - Includes the one-line amd64-only cloud-variant commit from #10570 so the two PRs can't conflict; if #10570 merges first, this PR rebases down to a single commit automatically. ## Verification - Image content is unchanged by layer reordering: the moved `RUN` references nothing from `/app`, and Docker layer ordering only affects caching, not the final filesystem (tool installs and app copy touch disjoint paths). - The cache ref is written only by this workflow — `docker.yml` runs on master/tag pushes, never on PRs — so the workflow's existing "no shared caches into build inputs" supply-chain stance is unchanged (BuildKit layer cache was already accepted via type=gha; the registry backend has the same writer trust). - Runtime proof lands with the first two master builds after merge: the first warms the cache, the second should show the tool layer and deps/build stages as CACHED in the build log, with wall clock dropping accordingly. I'll be watching those as part of managed-deploy work. ## Risks - Low. Worst case the registry cache misses (cold-build behavior, same as today). The weekly epoch means CLI tools update at most a week late inside images; a release built mid-week ships the tools from that week's first build. Cache refs add two small artifacts to ghcr. ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, via Claude Code with tool use and code execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (no test-affecting changes) - [x] I have added or updated tests where applicable (n/a — build config and layer ordering only) - [x] I have updated relevant documentation to reflect my changes (in-file comments document both mechanisms) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ce40343b9c |
build: publish the cloud image variant for amd64 only (#10570)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Published container images are how both self-hosted users and managed-deployment hosts run it > - The repository publishes two variants: the self-hosted image and a cloud variant for managed deployments > - The cloud variant was built for amd64+arm64, but its only consumers are managed-deployment hosts, which run amd64 > - The QEMU-emulated arm64 half dominates the build's wall clock, delaying every merge-to-deployable-image cycle > - This pull request drops arm64 from the cloud variant only, keeping the self-hosted image multi-arch > - The benefit is roughly halving the time from merge to a deployable cloud image, with no change for any actual consumer ## Linked Issues or Issue Description No existing public issue — inline description following the feature request template: **Subsystem affected** CI / release publishing (docker workflow) **Problem or motivation** The cloud image variant builds for `linux/amd64,linux/arm64`, but the arm64 half runs under QEMU emulation and dominates the job's wall clock — while no consumer of the cloud variant runs arm64 (managed-deployment hosts are amd64). Every deploy iteration pays ~double the necessary build time. **Proposed solution** Build the cloud variant amd64-only. The self-hosted image keeps `amd64+arm64` so ARM users (Apple Silicon, ARM servers) are unaffected. **Alternatives considered** Keeping multi-arch but building arm64 on native arm64 runners with a manifest merge — faster than QEMU and worth doing for the self-hosted image if its build time becomes a pain point, but unnecessary complexity for a variant with no arm64 consumers. **Roadmap alignment** Not on ROADMAP.md; CI/publishing speed improvement only. ## What Changed - `.github/workflows/docker.yml`: the `build-and-push-cloud` job's `platforms` is now `linux/amd64` (with a comment explaining why). The self-hosted `build-and-push` job is untouched. ## Verification - Build-config-only change; the workflow runs on merge to master. The published `-cloud` manifest will be amd64-only, which its consumers already pull. - No test changes: nothing at runtime differs on any platform that actually runs the image. ## Risks - Low. If an arm64 consumer of the cloud variant ever appears (e.g. local `docker run` on Apple Silicon for debugging), it would fall back to emulation on the consumer's machine or need this reverted — a one-line change. The self-hosted image's platform matrix is unchanged. ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, via Claude Code with tool use and code execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (no test-affecting changes) - [x] I have added or updated tests where applicable (n/a — CI platform matrix only) - [x] I have updated relevant documentation to reflect my changes (in-workflow comment documents the rationale) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
521271ebb7 |
build: bake the build commit into published images (#10566)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instances commonly run from published container images, and operators need to observe which build a container actually serves > - `/api/health` now reports the running build commit, and server-info already falls back to `PAPERCLIP_BUILD_COMMIT` when git is unavailable > - But published images carry no `.git` and never received `PAPERCLIP_BUILD_COMMIT`, so containers report `commit: null` — verified live against a current image > - That leaves the new deployment-verification field inert exactly where it matters most: containerized deploys > - This pull request bakes the exact build commit into both image variants at build time, mirroring how `PAPERCLIP_BUILD_VERSION` is already stamped > - The benefit is that containers report their true commit on `/api/health`, so deploy tooling can verify a rollout actually shipped ## Linked Issues or Issue Description Companion to #10563 (which exposed the `commit` field on `/api/health`). Inline description following the bug report template: **What happened?** A container from a published image responds to `GET /api/health` with `"commit": null`. The image has no `.git` directory and the `PAPERCLIP_BUILD_COMMIT` fallback that `server-info` supports is never provided at build time, so git metadata resolves as unavailable. **Expected behavior** A container reports the commit it was built from, the same way it already reports its build version via the baked `PAPERCLIP_BUILD_VERSION`. **Steps to reproduce** Run any published image (e.g. `ghcr.io/paperclipai/paperclip:sha-c4f6264-cloud`) and `curl /api/health` — `commit` is `null` even though the build commit is known at image-build time. **Paperclip version or commit** `sha-c4f6264-cloud` (first image containing #10563). ## What Changed - `Dockerfile`: new `PAPERCLIP_BUILD_COMMIT` build arg, exported as an ENV in the production stage (the `cloud` stage inherits it), directly parallel to `PAPERCLIP_BUILD_VERSION`. Empty for local `docker build`, which keeps the normal fallbacks. - `.github/workflows/docker.yml`: both build jobs pass `PAPERCLIP_BUILD_COMMIT=${{ github.sha }}`. ## Verification - Reviewed the plumbing end-to-end: `build-commit.ts` reads `PAPERCLIP_BUILD_COMMIT` (validated as a full SHA), `server-info.ts` `readGitInfo` falls back to it when the git CLI fails, producing `available: true, fullSha` — which `/api/health` surfaces as `commit`. - Verified live that a current published image reports `commit: null`; this change repairs that on the next build. Post-merge, the first master image should report its commit — I'll be verifying that as part of managed-deploy validation. - No test changes: the fallback path is already covered by existing server-info tests; this PR only supplies the env at image build. ## Risks - Low. Two build-time stamps; no runtime code changes. A wrong SHA would only mislabel the build (same failure mode `PAPERCLIP_BUILD_VERSION` already carries), and `${{ github.sha }}` is the exact commit the workflow builds. ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, via Claude Code with tool use and code execution); diagnosis included live probes of a running container's `/api/health`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (no test-affecting changes; server suites unaffected) - [x] I have added or updated tests where applicable (n/a — build-time stamps only) - [x] I have updated relevant documentation to reflect my changes (Dockerfile comments document the arg) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
51bb41c7e3 |
Expose the running build commit on the unauthenticated health response (#10563)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instances run self-hosted or under hosting/deploy tooling, and operators need to observe what build a server is actually running > - `/api/health` carries the git SHA only inside `serverInfo`, which is gated to board/agent actors — anonymous callers get a redacted body with no version signal at all > - Deploy tooling that manages instances from outside (fleet rollouts, hosting providers, upgrade scripts) therefore cannot ground-truth that a deploy actually shipped without holding credentials > - A build commit is a plain git SHA of this public repository — it is not a secret, and gating it buys no security while blocking legitimate verification > - This pull request surfaces the running build commit as a top-level `commit` field on every `/api/health` response, including the redacted anonymous one > - The benefit is credential-free deploy verification: any operator or tool can confirm which commit an instance serves, while the fuller `serverInfo` block stays access-controlled as before ## Linked Issues or Issue Description No existing public issue — inline description following the feature request template: **Subsystem affected** Server (API, runs, routes) **Problem or motivation** An anonymous `GET /api/health` returns a redacted body with no version information; the running git SHA exists only in `serverInfo.git.fullSha`, which requires a board/agent actor. External deploy tooling (fleet rollouts, hosting providers, upgrade scripts) therefore cannot verify that an instance is actually serving the build it was just upgraded to — a rollout that silently keeps running the old image is indistinguishable from a successful one at the health endpoint. **Proposed solution** Surface the running build commit as a top-level nullable `commit` field on every `/api/health` response shape, including the redacted anonymous one, while keeping the fuller `serverInfo` block access-controlled as before. A build commit is a plain git SHA of this public repository — exposing it costs nothing and enables credential-free deploy verification, like the `version` endpoints on most server software. **Alternatives considered** Authenticating deploy tooling as a board actor to read `serverInfo` — rejected: it forces credential plumbing into infrastructure that only needs a public SHA, and adds a whole class of auth-misconfiguration failure to deploy verification. **Roadmap alignment** Not on ROADMAP.md; a small operational observability improvement, no overlap with planned core work. ## What Changed - `server/src/routes/health.ts`: derive `commit` from the server info snapshot (`serverInfo.git.fullSha` when git metadata is available, else `null`) and include it as a top-level field on every `/api/health` response shape — the redacted anonymous body, the full-details body, the no-db body, and the 503 database-unreachable body. - `serverInfo` itself remains gated to full-details responses exactly as before; only the bare commit is newly public. - `server/src/__tests__/health.test.ts`: updated exact-shape assertions to include `commit`, and added an assertion that `commit` is `null` (not omitted) when git metadata is unavailable. The redacted-response tests now pin that anonymous callers receive the commit. ## Verification - `pnpm vitest run src/__tests__/health.test.ts` in `server/` — 13 tests pass, including the redacted-anonymous shapes (which now pin the `commit` field) and the git-unavailable `null` case. - `tsc -p server/tsconfig.json --noEmit` — clean. - Manual: `curl -s https://<instance>/api/health` as an anonymous caller returns `"commit": "<full sha>"` alongside the existing redacted fields. ## Risks - **Version disclosure:** anonymous callers can now fingerprint the exact running commit. This is a deliberate trade-off: the builds are of a public repository (the SHA reveals no private code), the endpoint already responds to anonymous callers, and the operational value — verifying deploys actually shipped — outweighs the marginal fingerprinting surface. Operators who consider this sensitive are typically fronting `/api` with their own access controls already. - Otherwise low risk: no behavioral change to any gated field, no schema or API-surface removal; `commit: null` keeps the field shape stable when git metadata is absent (e.g. non-git installs). ## Model Used Claude Opus 4.8 (`claude-opus-4-8`, extended thinking, via Claude Code with tool use and code execution) authored the change and tests; finalized and PR'd under Claude Fable 5 (`claude-fable-5`). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (none needed beyond code comments — health endpoint has no standalone doc) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
dd1a7f5290 |
Ensure app-home ownership before the privilege drop, not only on remap (#10530)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Docker image persists all instance state (project checkouts, worktrees, run logs, uploads) under `PAPERCLIP_HOME`, and deployments mount a volume there for durability > - The entrypoint starts as root and drops privileges to the `node` user, but it fixes `PAPERCLIP_HOME` ownership only when it remaps the user's UID/GID > - A freshly mounted volume arrives root-owned and shadows the image's build-time `chown`, so a default-UID boot drops privileges onto an unwritable home and the server crashes on its first `mkdir` > - This pull request makes the entrypoint probe the home's ownership and chown whenever it does not match the runtime user, before the privilege drop > - The benefit is that the image works out of the box on any platform-managed volume, with the common already-correct boot staying chown-free ## Linked Issues or Issue Description No public issue exists — describing the bug inline (per the bug report template). **What happened?** Running the image with a freshly created volume mounted at `/paperclip` (a Docker named volume, a Kubernetes PV, or any platform-managed volume) and the default `USER_UID`/`USER_GID` crashes on boot: `Error: EACCES: permission denied, mkdir '/paperclip/instances/default/logs'`. **Expected behavior** The container boots and initializes its instance tree on the mounted volume, exactly as it does when `/paperclip` is the image's own (build-time chowned) directory. **Steps to reproduce** 1. `docker volume create paperclip-data` 2. `docker run -v paperclip-data:/paperclip ghcr.io/paperclipai/paperclip:<any current tag>` 3. Observe the EACCES crash on the first `mkdir` under `/paperclip`. **Root cause** `scripts/docker-entrypoint.sh` chowns `/paperclip` only inside its UID/GID remap branch (`changed=1`). A fresh volume mount is root-owned and shadows the image's build-time `chown node:node /paperclip`; with the default 1000:1000 no remap happens, so no chown happens, and `gosu node` drops onto an unwritable home. **Paperclip version or commit:** reproduces on `master` and any published image. **Deployment mode:** any; observed on managed-cloud volume mounts and reproducible with plain Docker named volumes. **Installation method:** Docker image (`ghcr.io/paperclipai/paperclip`). **Related PRs (dedup search):** no open or merged PR touches the entrypoint ownership logic; the entrypoint's privilege-handling tests were added previously and this extends them. No duplicate found. ## What Changed - `scripts/docker-entrypoint.sh`: the remap-conditional `chown` is replaced by an ownership probe — after any UID/GID remap, the entrypoint stats `PAPERCLIP_HOME` (default `/paperclip`) and runs `chown -R node:node` only when the owner does not match the runtime user, before `exec gosu node`. Covers fresh root-owned mounts and trees written under a previous UID mapping; the already-correct boot performs no chown. The unprivileged (non-root start) branch is unchanged. - `server/src/__tests__/docker-entrypoint.test.ts`: `stat` stub added to the harness; new cases for the fresh root-owned mount with default UID/GID and for `PAPERCLIP_HOME`-relative probing; the remap case now models the post-remap ownership mismatch. ## Verification - `pnpm vitest run server/src/__tests__/docker-entrypoint.test.ts` — 7 passed (5 existing behaviors unchanged, 2 new). - Live on a managed deployment: a container that crash-looped with the EACCES above boots cleanly once the home is chowned before the drop (the same effect this entrypoint change produces; forced there by a UID remap as an interim workaround). ## Risks - Low. Behavior changes only for boots where `PAPERCLIP_HOME` exists with mismatched ownership — exactly the boots that crash today. `chown -R` on a large previously-mismatched tree adds one-time boot latency; correctly-owned homes skip it entirely. Kubernetes restricted / OpenShift non-root starts keep the existing exec-directly path untouched. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic; Claude Code CLI with extended thinking and tool use; tests executed locally via Vitest). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (no duplicates; extends the existing entrypoint privilege tests) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
075951f6bd |
Fix import completion UX: inbox flood, false-failure message, stale company list (#10538)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company Import/Export (#10507, hardened in #10523 and #10531) now imports a large company end to end via an async job > - A real 1,418-issue import succeeded, but three rough edges showed up in that success > - Imported issues flooded the inbox, a completed import surfaced a false "failed" message after its in-memory result expired, and the new company didn't appear in the switcher until a manual refresh > - This pull request keeps imported issues out of the inbox, treats an expired-but-completed import as success, and refreshes the company list on completion > - The benefit is that a successful import looks and feels successful, and doesn't bury the user's inbox in historical tasks ## Linked Issues or Issue Description - Refs #10507 / #10523 / #10531 (Import/Export and its hardening). No open issue; three post-import bugs described above. ## What Changed - **Imported issues no longer flood the inbox.** The inbox "mine" tab is a query: an issue is "touched" if the user authored a comment on it, and import re-attributes bundled user comments to the importing user — so every imported issue appeared. Import now seeds a per-user `issue_inbox_archives` row for each imported issue (via a batched `issues.archiveImportedInbox`), the exact table the inbox visibility query excludes. Gated on an actor user id, so agent/system imports and normal issue creation are untouched; genuine new activity still resurfaces the issue. - **A completed import no longer shows a false failure.** The in-memory job's terminal retention was 5 minutes, so a poll after that 404'd and the UI showed "failed." Retention is extended to 60 minutes — the real mitigation for a user who steps away during a long import. `watchImportJob` additionally treats a *server-confirmed* success whose full result is no longer retained (a `succeeded` status carrying only the compact summary — a cloud tenant job, or a board job whose full in-memory result aged out) as a soft success ("import completed — open the company"), navigating by the summary's company id. A 404 while the job is still being watched is *not* treated as success: a running job is never dropped by the retention sweep, so its disappearance means a restart mid-import that may not have finished, and it surfaces the honest "may have restarted while the import ran" error. A first-poll 404 (the id never existed) is likewise a real error. - **The imported company appears without a refresh.** `onSuccess` now invalidates the companies/switcher query unconditionally (covering both the full-result and expired-but-completed paths) and navigates by the job's company id. ## Verification - shared/server/ui typechecks clean; 15 UI tests in the touched spec green, plus the embedded-Postgres import batching and portability-routes suites. - New tests: embedded-Postgres test that imported touched issues are archived for the actor and excluded from the inbox query while a normally-created issue still appears; job resolvable at the old window+1 and only 404s past 60 min; UI soft success on a server-confirmed `succeeded` job without a retained full result (no error, list invalidated, navigates by company id), a running-then-gone job → honest error (restart mid-import), and a first-poll 404 → error. ## Risks - Low and import-scoped: the inbox archive only affects imported issues for the importing user; normal issue creation and non-user (agent/system) imports are unchanged. Retention extension is a constant; the async job store remains in-memory by design. A restart mid-import still 404s and is surfaced honestly as a possible failure (never masked as success); only a server-confirmed success whose full result has expired is reported as a soft success. ## Model Used - Implementation: Claude Fable 5 (`claude-fable-5`, Anthropic). Review hardening (the confirmed-success narrowing): Claude Opus 4.8 (`claude-opus-4-8`, Anthropic). Both via the Claude Code CLI with extended thinking + tool use; root-caused against the live import. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |