mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 05:31:46 +02:00
0bff1d5cb60e5fc8ddaf3d5cd71c86745d10bdb5
1653
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0bff1d5cb6 |
refactor(server): simplify the wake-queue module (#13139)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server uses a deferred wake queue to release work at the correct time > - The wake queue had repeated decisions, helpers, queries, and recovery data > - This repetition made the module harder to read and left a drain invariant implicit > - This pull request moves pure decisions into the policy layer and simplifies the queue flow > - The benefit is a smaller, clearer module with the same behavior and direct test coverage ## Linked Issues or Issue Description Refs #13136 **Problem:** The deferred wake queue carried repeated logic across the application and database layers. **Expected behavior:** The queue keeps the same wake, release, recovery, and escalation behavior after the refactor. **Solution:** Move pure decisions into the policy layer, share repeated data and predicates, and state the drain invariant in the application layer. ## What Changed - Remove the unused database port method and carry the blocked release notice kind as a typed field. - Move four pre-drain decisions into pure policy functions with table-driven tests. - Resolve the responsible user once in the application layer. - Use the canonical helper for agent invokability checks. - Share string helpers and run predicates across the module. - Split the deferred-wake decision flow and share recovery facts with a discriminator. - Bound the drain loop and throw when it processes a wake identifier twice. - Rename queue ports to describe their behavior. - Share row-loading code between stranded-issue escalation adapters. - Restore the interaction-continuation integration test case. ## Verification - Run the three wake-queue module test files. They pass 61 cases locally. - Run `pnpm check:module-boundaries`. It passes locally. - Run the `server/` type-check and confirm that no error names the changed wake-queue module or `heartbeat.ts`. - Run continuous integration and confirm that `promotes an interaction continuation after removing a coalesced self-authored comment` passes. ## Risks The change refactors queue control flow and database adapter boundaries. The main risk is a behavior change in deferred wake release or recovery. The new policy tests and the restored integration case cover these paths. The local environment cannot run the integration test because the same dependency failure occurs on the base branch. ## Model Used OpenAI Codex, GPT-5, context window not exposed by the runtime, with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6dd48cad43 |
refactor(server): move the release half of the deferred wake state machine into a wake-queue module (#13132)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The heartbeat service releases deferred issue wakes and promotes the next run > - The release logic sat in a large service function, which made its decisions and database effects hard to test > - A promotion race could reopen an issue without creating the promoted run > - Company checks did not protect every read and write, and some readers used a second transaction connection > - This pull request moves the release logic into a layered wake-queue module and closes these race and company-scope defects > - The benefit is clearer decisions, safer writes, and focused tests while existing callers keep the same entry point ## Linked Issues or Issue Description Refs: #10195 ## What Changed - Move the release half of deferred issue execution from `heartbeat.ts` into `server/src/modules/wake-queue/`. - Add pure policy decisions with table-driven tests. - Claim a wake before the reopen write and advance to the next wake when the claim fails. - Add company predicates to guarded reads and writes. - Pass the transaction-scoped issue snapshot to reader ports. - Keep `releaseIssueExecutionAndPromote` as the public wrapper. ## Verification - Run `pnpm check:module-boundaries`. - Run `pnpm exec tsc --noEmit` inside `server/` and compare the result with the known baseline. - Run the wake-queue unit and adapter tests in continuous integration. - Check the promotion-claim ordering, guarded writes, transaction-scoped reads, and cross-company outcomes. - Search the pull request diff for internal issue identifiers. ## Risks - The refactor changes the transaction path for deferred wake release. - A stale or lost wake claim now skips that wake and continues with the next queued wake. - The public wrapper keeps its name and signature, which limits caller risk. - Continuous integration must confirm the full server test suite and build. ## Model Used OpenAI Codex, GPT-5, exact runtime model version not exposed, context window not exposed, with shell, Git, and GitHub tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018ca5daaf |
fix: verify ACP Stop and preserve safe continuation (#13119)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task controls coordinate provider execution and queued user messages. > - Stop could finish before an embedded ACP provider stopped its tools. > - A later request could be held for reconciliation without a clear task response. > - A restored provider could also retain the stopped run's API credential. > - This pull request verifies provider termination and preserves safe session continuation. > - Operators can continue known-safe work and see why uncertain work cannot start. ## Linked Issues or Issue Description **What happened?** Stop could leave an embedded ACP provider running. A queued follow-up followed by “go” could fail before it reached the provider. Task chat could show a generic missing-response message. Even a restored session could use the previous run's credential and fail its task update. **Expected behavior** Stop waits for confirmed provider termination. A later explicit wake continues the same compatible session only when recorded actions have known outcomes. It carries pending comments and the current run's environment. Uncertain actions retain a visible reconciliation hold. Composer Stop preserves the existing pause rule: conversation can continue while paused, but task work requires Resume. **Steps to reproduce** 1. Start an embedded ACP task. 2. Send a second request while the provider is running. 3. Interrupt the run, then send “go”. Also test composer Stop followed by Resume work. 4. Check that the request is delivered once and that the provider can complete the task through the current run's API credential. 5. Repeat with an unfinished write. Confirm that the write stops and that further execution stays blocked with a visible reason. **Paperclip version or commit** Built from source on master at `3bc60dd8b` plus this branch. **Deployment mode** Local source build with an isolated embedded PostgreSQL instance. Refs #11183. Refs #12552. Those changes address recovery after operator cancellation. This change also covers embedded ACP termination, session proof, pending-comment delivery, and task feedback. ## What Changed - Propagate Stop into embedded ACP and wait for bounded adapter cleanup and provider exit. Retain the actual ChildProcess object for forced termination on all platforms; never signal a recycled numeric PID. - Preserve interrupted checkpoints only for acknowledged, local, persistent sessions with settled reads or no tools. Keep writes, incomplete actions, and forced termination blocked. - Restore the same compatible provider session with the current run's environment. Reject fresh-session fallback for an interrupted checkpoint. - Adopt pending comments on the next explicit wake. Stop alone does not dispatch them. - Share the execution-blocker rule across dispatch, Resume, and task detail. Show Stopped or Couldn't start with the recorded reason. Resolve the stopped agent for the run link, including reviewer runs. - Keep execution reconciliation holds intact when generic recovery sees queued comments or healthy child tasks. - Add process, service, component, and browser regression coverage. Fix disposable database cleanup and React test settling exposed by the full suite. ## Verification - Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. - Passed all three `acp-stop-continuation.spec.ts` browser journeys. They use an actual ACP child process and require task completion through the agent API. - Passed 165 adapter execution, operator-stop, and child-process control tests, 17 queued-comment route tests, and 65 tests in the two adjusted UI suites. Earlier focused recovery, heartbeat, and task-control tests also passed. - Manually used the browser to queue a request, Stop, send “go” while paused, and Resume. The same session answered once and moved the task to Done with the current run's credential. - Manually interrupted an unfinished write. Its file size stayed fixed for five seconds. “Go” showed the reconciliation reason and did not start another provider prompt. - Separate live Claude ACP smoke checks confirmed that Stop ended a disposable local write and that a no-tool interruption could resume the exact provider session. The browser fixture does not call Drive or another external app. - Passed all 5,615 UI tests and 3,090 other workspace tests. The CLI and general server groups pass with targeted retries: two transient server failures passed together on retry, and two embedded-database startup failures passed after removing abandoned shared-memory segments from this task's completed browser fixtures. All 144 serialized server suites completed, with 2,189 tests passing after two transient HTTP socket failures passed on retry. - Passed all 135 heartbeat process/recovery tests, including a deterministic regression that failed before the recovery-sweep fix. - Passed 18 dispatch integration tests, including stopped-reviewer links, company boundaries, and malformed run IDs. - Greptile is 5/5 on `7dd170d83`, with zero unresolved review threads. The security scan and all required CI gates pass for the same commit. ## Risks - Safe continuation depends on complete tool reporting and a restorable local provider session. Unknown outcomes remain blocked and require reconciliation. - Provider cleanup can take time. A timeout does not grant replay permission. - The change adds optional adapter context fields and an optional issue projection. It does not change the database schema or require a migration. - Test cleanup truncates company data only in a disposable test database. ## Model Used OpenAI GPT-6, running as Codex with repository tools, code execution, and browser interaction. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
04a9f89ede |
fix(server): bundle the vendored paperclip-runner instead of hand-mirroring its deps (#13121)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server package is published to npm, but its native-runtime driver code lives in `packages/paperclip-runner`, a private workspace package that is never published > - So the server build vendors the runner's compiled code by copying it in directly, instead of taking it as a normal npm dependency > - But `cp -R` only copies code, not `node_modules`, so every npm package the runner imports has to be re-declared by hand in `server/package.json` to stay resolvable once vendored > - That hand mirroring step is silent and easy to forget: it missed `smol-toml` in #13110, and CI stayed green while production crash-looped 3 seconds into every start (#13116) > - This pull request keeps the proven `cp -R` vendor step exactly as it was, and adds a build check that derives the required dependency set from an esbuild scan of the vendored entry points, failing loudly and precisely if any package the runner actually needs isn't declared in `server/package.json` > - The benefit is the dependency list is now verified against the real module graph instead of hand-copied, so this exact class of bug cannot pass a green build again -- without changing how the runner's code is laid out on disk, which several of its modules depend on for unrelated filesystem lookups ## Linked Issues or Issue Description Refs: #13110 (introduced the `smol-toml` import that the vendor step could not resolve), #13116 (the follow-up fix for a different oversight in the same PR), #11813 (the same "vendored package installed outside the monorepo dependency graph loses a runtime dependency" failure shape, in the Kubernetes plugin installer instead of the server build) No issue exists yet for this specific incident, so per CONTRIBUTING.md option (B): **What happened?** `packages/paperclip-runner/package.json` added `smol-toml` as a runtime dependency in #13110. `server/package.json`'s existing convention (see `acpx`, `ajv`) requires mirroring every runtime dependency the vendored runner imports into `server/package.json` too, because the server build copies the runner's compiled `dist/` tree with `cp -R` -- code only, no `node_modules`. That mirroring step was missed. CI never runs the compiled server (`node dist/index.js`); it only builds it, type-checks it, and boots the app in dev mode via `tsx` against source, which never touches the vendored path. So the PR merged green, and the deployed server crash-looped in production: ``` Error [ERR_MODULE_NOT_FOUND]: Cannot find package 'smol-toml' imported from /srv/paperclip/app/server/dist/vendor/paperclip-runner/drivers/codex/codex-startup-trust.js ``` **Expected behavior** Any npm package the vendored runner code needs at runtime should either be guaranteed present by construction, or the build should fail with a clear, actionable error before the change ever reaches a PR -- not silently pass CI and fail only once deployed. **Steps to reproduce (the original incident)** 1. Add a new runtime dependency to `packages/paperclip-runner/package.json` (e.g. a TOML parser) and use it from a module reachable from the runner's `index.ts` export graph. 2. Do not add the same dependency to `server/package.json`. 3. Run `pnpm build` in `server/` -- it succeeds. 4. Run `node dist/index.js` -- it crashes with `ERR_MODULE_NOT_FOUND` for the new package. ## What Changed - **Revision note:** the first version of this PR replaced the `cp -R` vendor step with an esbuild bundle of the runner's entry points. Greptile's review correctly caught that this broke packaged ACPX/OpenCode provider startup: several runner modules resolve sibling build artifacts via `import.meta.url`-relative filesystem paths (not JS imports) at whatever depth their source file sits at, and bundling collapses/rearranges that layout. The current version keeps the file layout untouched and only adds verification. See the second commit's message for the full explanation. - `server/scripts/verify-runner-vendor-dependencies.mjs`: a new build step that runs esbuild with `write: false` (a pure module-graph scan -- nothing is written to disk) against the runner's two entry points server actually imports (`index.js`, `testing.js`), with `packages: "external"` so its metafile reports exactly which npm packages the code needs at runtime. It fails with a precise, actionable error if any of them isn't declared in `server/package.json`'s `dependencies`. This is deliberately more precise than "mirror every dependency the runner declares": running it against this repo's real manifests shows `packages/paperclip-runner/package.json` declares dependencies (`react-markdown`, the codex/opencode CLI packages, ...) that only its unrelated `./react` and `./browser` export subpaths use -- server never imports those, so a blanket mirror rule would demand dependencies server doesn't actually need. - `server/package.json`: added the new check into the `build` script (right after the runner is built, before the expensive `tsc`/copy steps, so it fails fast), and added `smol-toml` (`^1.4.2`, matching `packages/paperclip-runner/package.json`) to `dependencies` -- the actual missing piece from #13110. The vendor step (`cp -R ../packages/paperclip-runner/dist/. dist/vendor/paperclip-runner/`) is unchanged from before this PR. - Widened `server/vitest.config.ts`'s `include` to also run `scripts/**/*.test.mjs`, and added `server/scripts/verify-runner-vendor-dependencies.test.mjs` unit-testing the pure dependency-diff function (`findMissingVendorDependencies`) against the exact shape of the `smol-toml` incident, plus a case proving an unreachable dependency (like `react-markdown`) is correctly never flagged. - Updated `server/src/__tests__/server-package-build-script.test.ts`'s existing build-script assertions to match. ## Verification - `node --check` on the new script -- syntax OK. `node -e` JSON-parsed the edited `package.json` files after every edit. - Unit-verified `findMissingVendorDependencies` directly against: nothing missing, one missing (the `smol-toml` shape), and multiple missing with stable sort order. - Ran the actual check against this repo's real `packages/paperclip-runner/package.json` and `server/package.json` (via a standalone `node` invocation, since `pnpm build` needs a Rust toolchain this sandbox doesn't have -- see below) to see its real output. It correctly reported `smol-toml`, `acpx`, and `ajv` as already satisfied, and did **not** flag `react-markdown`, `remark-gfm`, `json-schema-to-ts`, `opencode-ai`, `@openai/codex`, or the `@agentclientprotocol/*` packages -- confirming the "reachable from index.js/testing.js" scoping works as intended and doesn't demand dependencies server doesn't need. - Built a fixture tree at a real filesystem location (not just in-process) mimicking `packages/paperclip-runner`: a manifest declaring both a reachable dependency (`smol-toml`, actually imported by the fixture's `dist/index.js`/`testing.js`) and an unreachable one (`react-markdown`, declared but never imported). Copied the real script next to a fixture `server/package.json` and ran it as its own process (`node server/scripts/verify-runner-vendor-dependencies.mjs`), twice: - `smol-toml` missing from the fixture's server dependencies → the script throws with the exact intended message and exits 1. - `smol-toml` present, `react-markdown` absent → the script exits 0, proving the unreachable dependency is correctly never flagged. - Not verified locally: the real `packages/paperclip-runner` build, and therefore the check running end-to-end against its true `dist/index.js`/`dist/testing.js`. This sandbox has no Rust toolchain (the runner's own build compiles a Cargo binary) and an incomplete workspace install. CI's `Build` job (`.github/workflows/pr-trusted.yml`) runs the real thing; I'll watch it on this PR. ## Risks - The check's precision (scoping to what's reachable from `index.js`/`testing.js`, rather than every declared runner dependency) means a dependency that becomes reachable through some *other* export subpath server starts importing later would need this check's entry-point list updated too. That list is a 2-line array in the script with a comment explaining why, and matches the only two paths server/src actually imports today (verified by a repo-wide search). - This only changes a build-time check; the actual vendored file layout (`cp -R` of the runner's whole compiled tree) is byte-for-byte the same as before this PR, so there's no behavioral change to the running server beyond `smol-toml` now being present as intended. - I could not exercise the real Rust-backed build locally (no Cargo in this sandbox); see Verification. I am relying on CI's `Build` job to confirm this end to end and will fix forward if it surfaces something the fixture-based testing didn't. ## Model Used Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code. Standard (non-extended) reasoning mode, with tool use (Bash, Read, Edit/Write, `gh`) for repository exploration, local esbuild-based verification against hand-built fixtures, and PR authoring. No extended thinking mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — see Verification: full local verification was not possible (no Rust toolchain, incomplete workspace install in this sandbox); watching CI's `Build` job on this PR to confirm. - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — no user-facing docs describe this internal build step; none needed updating. - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green — pending, will monitor. - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — addressed the first review round; watching for re-review. - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3b550c80fa |
fix(codex): correct startup trust, history reads, and resume usage (#13110)
## Thinking Path
> - Paperclip runs Codex locally and in remote sandboxes.
> - The runner must preserve startup configuration and session identity.
> - Missing project trust can disable repository configuration.
> - Full-history requests use deprecated provider fields.
> - Resume usage describes old work and must not become new run usage.
> - This change corrects startup trust, state reads, and usage
classification.
## Linked Issues or Issue Description
**What happened?**
Normal Codex runs could show repository-trust and history-deprecation
warnings.
Resume could report the preceding turn's token snapshot as a late-turn
warning.
The historical last-usage value could also be attributed to the new run.
**Expected behavior**
Trust the server-selected startup root in isolated configuration. Read
lightweight
provider state and paginated evidence. Use historical cumulative usage
as a
baseline without a new charge or user-facing warning.
**Steps to reproduce**
1. Start a native Codex task in a selected repository.
2. Finish the turn and resume the provider thread.
3. Inspect provider notices, history requests, and per-run usage.
4. Repeat startup and cold resume inside a Daytona sandbox.
**Paperclip version or commit**
Codex CLI 0.153.4 is the pinned runtime and reproduced baseline.
Replayed onto master at
|
||
|
|
ca96e1eb0a |
fix(runner): keep streaming after task completion tools (#13108)
Keep receiving provider events after paperclip_finish, drain pending event persistence, and select the final assistant answer after the provider turn ends. Preserve cancellation, failure, and governed-wait behavior. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6abeb67334 |
feat: add opt-in chat provider and data foundation (#13100)
Add dormant provider contracts, qualified patched adapters, tenant-scoped persistence and lifecycle ownership without activating chat routes. Preserve the experimental integration as dependent PR #13038. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
7d84b183fb |
test(heartbeat): pin 8 uncovered branches of the deferred issue-execution wake state machine (#13103)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server manages issue execution, queued comments, and agent wake state > - Deferred wakes can pass through several failure, hold, rollback, and steering branches > - These branches had no direct tests, so a later change could fail without clear evidence > - This pull request adds characterization tests for eight uncovered branches > - The benefit is clear test evidence for future changes to deferred issue execution ## Linked Issues or Issue Description **What existing behavior does this improve?** The server behavior for deferred issue-execution wakes and queued-comment steering transactions. **Current behavior** The code handles missing agents, cross-company agents, pause holds, rollback, steering outcomes, and invalid reorder sets. These branches had no direct tests. **Proposed behavior** Keep the current behavior and test each branch against the current implementation. **Reason and benefit** The tests make silent behavior changes visible. They also protect the wake queue while later work moves this logic into a dedicated module. **Breaking changes** None. This pull request changes no production code and no runtime behavior. Related public pull requests: #12671, #11168, and #10199. ## What Changed - Add tests for missing and cross-company deferred agents. - Add tests for pause-hold promotion and cancellation. - Add a test for atomic rollback when the responsible user cannot resolve. - Add tests for successful, timed-out, and rejected native steering. - Add a test for invalid queued-comment reorder input. ## Verification - Run `server/src/__tests__/heartbeat-comment-wake-batching.test.ts` against an embedded Postgres database. - Run `server/src/__tests__/issue-queued-comments-routes.test.ts` against an embedded Postgres database. - Confirm that all eight new tests pass. - Confirm that the pull request CI checks pass. ## Risks Low risk. The diff changes test files only. It does not change production code, database schema, API behavior, or runtime behavior. ## Model Used Codex, GPT-5, tool use and code review support. The model did not author production code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2991a59b17 |
fix(adapters): prevent engine fallback and preserve usable runtime defaults (#13105)
## Thinking Path > - Paperclip manages agents that must write work and report task outcomes through its API. > - Local adapters select an execution engine and its permission settings. > - A higher ACP Node requirement can make an unchanged installation lose access to its default engine. > - The adapter then silently selects CLI, which can change permissions and block API access. > - This pull request keeps the engine choice fixed and reports missing prerequisites before work starts. > - It also gives explicit Codex CLI runs usable defaults and keeps managed services on a supported Node runtime. ## Linked Issues or Issue Description Refs #12215. Related changes: #11792 raised the Node requirement; #13094 addressed separate runner networking behavior. This change fixes the engine-selection and managed-launcher paths. **What happened?** An unchanged agent could switch from ACP to CLI after an upgrade. Codex CLI then used read-only permissions with networking disabled. The run could finish without updating its task. Repeated recovery attempts used the same unavailable setup. Managed updates also skipped the Node check and did not refresh old launchers. **Expected behavior** An unavailable engine must fail with a clear setup error. It must not silently select another engine. Explicit CLI runs must be able to write workspace files and call the API unless the operator configures stricter settings. Managed updates must validate Node and keep child tools on that runtime. **Steps to reproduce** 1. Run an ACP-default agent under Node 22 after the ACP minimum rises to 24.11. 2. Leave the engine unset and disable the approval/sandbox bypass. 3. Observe the old adapter select CLI and fail to write task disposition through the API. 4. Start a managed service with an old launcher and a supervisor PATH that selects a different Node for child tools. ## What Changed - Remove automatic engine fallback for Codex, Claude, Gemini, and Kimi. Check prerequisites for default and explicit ACP selections. - Return a configuration error with proof that provider work did not start. Stop automatic continuation retries for this error. - Enable Codex ACP workspace networking at the actual turn boundary. Upstream mode presets otherwise force it off even when config.toml enables it. Preserve explicit network denial and read-only mode. - Set workspace-write and network access defaults for explicit Codex CLI runs. Preserve explicit sandbox modes, profiles, and network restrictions. - Pin the validated Node directory in managed launcher PATH. Refresh legacy launchers during installs and npm/Git updates. - Reject updates on unsupported Node. Keep update checks, dry runs, and rollback available. - Synchronize the qualified Codex ACP executable identity across server, TypeScript runner, Rust runner, and provider-pack launch paths. - Add regression tests and update engine and installation documentation. ## Verification - [Full CI passed on the final head](https://github.com/paperclipai/paperclip/actions/runs/34387099695): typecheck, build/native runner verification, all general and serialized test shards, all browser shards, release registry, canary dry run, and policy checks. - Greptile: 5/5 on `2c1d6e2815830a5cd39e36c8a082cc0c4441b6c0`, with no unresolved review findings. Security gates are green. - Full workspace typecheck and build also passed locally. The final deployed Linux build passed. - Full Codex, Claude, Gemini, and Kimi source test suites: 804 passed, 2 skipped. Installer, updater, and launcher tests: 47 passed. Installed ACP turn-boundary tests: 3 passed. ACP packaging tests: 14 passed. Focused recovery classification tests also passed. - Real Linux Codex CLI runs, both fresh and resumed, wrote a workspace file and reached the control-plane health API with the new defaults. - Explicit read-only and network-disabled control probes retained those restrictions. - A real ACP run on the final deployed Linux build wrote a file and reached the control-plane API with HTTP 200, without engine fallback. The same probe failed DNS before the turn-policy patch. - Executable-identity and installed-policy contracts: 12 passed. Affected native server tests: 197 passed. Runner factory tests: 21 passed. Rust qualification and native provider integration tests: 11 passed. - Deployed the production changes to a Linux service on Node 24.20 after a verified database backup. Health, bootstrap readiness, static UI, executable/cwd identity, and guarded restart checks passed. The restart lost no runs. - Corrected stale Kimi skill-default and Gemini remote-archive fixtures; both suites pass. ## Risks - Default or legacy auto engine settings now fail when ACP is unavailable. Operators who intend to use CLI must select it explicitly. - Codex CLI now permits workspace writes and networking by default, and ACP workspace-write turns permit networking by default. Explicit operator sandbox settings remain authoritative. - Old managed launchers keep their pinned Node until they are reinstalled under a supported runtime. An old updater cannot repair itself; the documentation gives the current installer command. - Custom service wrappers and global/source installations must configure their runtime PATH. No database migration is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, shell execution, and test tools. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cfd30fb07 |
feat(ui): add composer Stop and simplify task controls (#13104)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task composer is where operators direct running agents. > - Operators need to stop work without leaving the conversation. > - Existing pause controls already hold task trees and interrupt both runner types. > - This pull request connects the composer to those controls and removes repeated feedback. > - Operators can pause work quickly and still queue messages while agents run. ## Linked Issues or Issue Description **What existing behavior does this improve?** Task pause, resume, and cancellation in the task page and composer. **Current behavior** The empty composer cannot stop a running task. Task controls require extra confirmation and reason text. Pause can show several notifications for the task already on screen. **Proposed behavior** Show Stop while this task runs and the composer is empty. Text or attachments switch it to Send. Stop and the menu use the same manual pause hold. Parent pauses include descendants. Keep task cancellation in the menu with a compact confirmation. Show one quiet pause row and gray cancelled-run details. **Reason and benefit** Operators can interrupt execution with one click. Drafts and queued messages keep their existing behavior. The UI waits for actual termination, including native cancellation acknowledgment. **Breaking changes** No endpoint, schema, or task-status change. Pause no longer asks for confirmation or a reason. Resume now honors the existing wake-agents option. Task notifications are suppressed for the task and subtree currently in view. Related UI work: #8228 changes navigation and composer shortcuts. This PR covers execution controls. No duplicate Stop-button PR was found. The change improves existing controls and does not duplicate a roadmap milestone. ## What Changed - Add Stop, pending feedback, duplicate-click protection, and inline errors to the composer. - Share the pause mutation across the composer, active-run controls, and menu. - Poll affected runs after a pause request. Require native cancellation acknowledgment. - Remove pause confirmation and shared reason fields. Reduce cancel confirmation to its task count and actions. - Honor wake-agents for executable tasks only. Preserve the pause when recovery review is needed; show partial wake failures inline. - Preserve explicit legacy reconciliation decisions while their continuation waits for dispatch. - Suppress notifications for visible task trees. Use quiet pause and cancellation feedback. - Add interactive stories using production controls and native/legacy end-to-end tests. ## Verification - User reviewed the running feature and revised Storybooks in the browser. - Rebased focused checks passed: 295 original targeted tests, 161 updated route/page/notification/status tests, and 26 recovery integration tests. - Both isolated runner journeys pass on the final revision (1.7 minutes). Coverage includes queueing, parent and child interruption, persisted holds, no automatic continuation, reconciled resume, cancellation, terminal exclusions, and no Stop toast. - Native coverage uses real runnerd with a deterministic provider fixture. Legacy coverage checks actual process termination. Live hosted-provider execution was not tested. - Repository typecheck and build, Storybook build, and token gates passed after rebase. The final server typecheck/build also passed. - The broad local run completed its general-server stage with 7,219 passing tests, 48 skipped, and two failures from cached pre-fix source and a stale native provider fixture. Both failed tests pass in fresh final-head reruns after rebuilding the fixture; the script did not continue to its later local stages. CI runs all test groups on the final revision. - Final revision: all 31 applicable CI checks passed; Storybook visual regression was skipped by its workflow conditions. Greptile: 5/5, zero unresolved comments. - Review `Tasks / Execution Controls` in Storybook. Type and clear a draft, stop a run, expand cancellation details, and test the menu on desktop and mobile. ## Risks - Stop pauses descendants for a parent task. This is the existing pause contract. - A held task can remain active if interruption fails. The UI shows an error instead of claiming termination. - Resume can start multiple assignees when wake-agents is selected. Backlog, blocked, and terminal tasks stay excluded. Existing execution reconciliation remains mandatory where required; Resume never invents action-outcome evidence. - Notification suppression uses the visible task and cached subtree. Notifications for unrelated work remain enabled. ## Model Used OpenAI GPT-6 through Codex. The exact runtime snapshot and context-window limit are not exposed in this session. Used reasoning, tool calls, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fac07b42ad |
fix(runner): preserve durable native session authority across recovery (#13092)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner carries tool results and task output to the control plane. > - A lost connection must not change which run owns a result. > - A session must not become reusable while provider output is still pending. > - This pull request adds strict recovery evidence and bounded drain barriers. > - It preserves current PRP version negotiation and session-goal support. > - The benefit is safer reuse of native sessions after a transport failure. ## Linked Issues or Issue Description Refs #13038. This is the first of two stacked pull requests. It contains the native runtime prerequisites. The second pull request contains the experimental chat-channel integration. It preserves the provider identity and typed terminal-failure contracts in #13074 and the durable recovery work in #13075. **What happened?** Native session failures could leave retained provider events, incomplete tool results, or warm handoff state that was not safe to reuse. A later run could observe output from an earlier authority. **Expected behavior** Recovery must preserve exact run, tool, process, artifact, and lease evidence. Uncertain or corrupt state must fail closed. A successful close must prove that retained provider output is settled. **Steps to reproduce** Run the transport and control-plane regressions. They hold and drop authenticated frames, fail durable writes, and restart fresh controllers and runner processes with retained state. Provider executables are local test fixtures. ## What Changed - Preserve pending provider cleanup and semantic-result evidence across session close and restart. - Add an authenticated warm handoff with exact old and new identities, durable receipts, and completion acknowledgement. - Drain retained provider events under the cumulative acknowledgement fence. - Reject corrupt tool-result contracts without unsafe provider replay or reusable checkpoints. - Keep ordinary PRP v1 sessions and current session-goal behavior. Require negotiated PRP v2 and acknowledged native session evidence before warm authority rotation. - Preserve late semantic inputs and exact durable result receipts until close can prove settlement. - Add transport, crash-window, artifact, checkpoint, and final-output regressions. - Deduplicate resolved execution delivery under the current issue lock. Reuse the exact existing successor after concurrent scans or a lost acknowledgement. Preserve newer operator evidence. - Persist idle provider integrity/capacity failures before process retirement, retain permanent model-rejection classification, and keep external question identifiers out of task instructions. - Expose only the context source on native status events. Keep thin dispatch projections compatible without exposing the complete context. ## Verification - Review-fix revision: 128 runtime-context/native-session tests, five idle-failure/adjacent Rust cases, 24 warm crash-window cases, three startup-notification/close cases, and five attach/backlog cases passed. The security and idle-failure cases were first reproduced failing. - Prior merged revision: runner production build, TypeScript typecheck, complete Rust workspace tests and formatting passed; 272 focused runner tests and two real PostgreSQL regressions passed. - Earlier full runner runs and CI Build failed on missing semantic-result fixture receipts, stale local provider fixture bytes, startup-notification ordering, and a confirmation-loss fixture that could accidentally send its final ACK. Each cause was reproduced and corrected without relaxing production authority or close assertions. These earlier runs are retained as failures, not represented as passing verification. - The first local repository-wide run failed before later phases because the isolated install omitted PostgreSQL's native-library aliases; it also encountered an unrelated occupied-port fixture. Those results are retained, not represented as a passing run. - Exact `335b2ee52709afb3885d4d6ebb2a3ece4b5864d6`: the complete runner suite passed 1,888 tests, with 10 existing skips. The full Rust release workspace passed with serial test scheduling. The unchanged parallel Rust run hit the five-second 300-descendant fixture deadline; that failure is retained. No deadline or assertion was relaxed. - The resolved-execution regression suite passed 57 tests, including concurrent delivery, lost acknowledgement, superseded authority, and newer operator evidence. Plain server typecheck passed. The duplicate-delivery cases were first reproduced failing. - Prior exact `335b2ee52709afb3885d4d6ebb2a3ece4b5864d6` CI passed all required jobs and Greptile reported 5/5. Its local general-server run passed 7,208 tests but failed one responsibility fixture; later phases did not run. The fixture started the next wake while its bounded handoff was active. It also used nonexistent comment IDs, which hid the current stored-message-author identity rule. The updated tests use real message authors, preserve task ownership, and await exact automatic handoffs. No production identity policy changed. - Current head `aa39275a1f300f7d1a0b16cd0885eea567cff6b0` includes current master and the native context-source projection. The focused identity/status cohort passed 27 tests and plain server typecheck passed. Fresh full repository tests, types, build, required CI, and Greptile review are pending. Final results will be updated before merge. - This is deterministic local-provider evidence. It is not a claim of complete live-provider qualification. ## Risks - This changes authenticated recovery and close ordering. The TypeScript transport and runner binary must be built from the same revision. - Failed or incomplete evidence intentionally prevents reuse and can require a fresh run. - PRP v1 ordinary/cold sessions remain supported. A v1 connection lease cannot upgrade in place. A current v2-capable runner held on a v1 lease was qualified through owned-process retirement/join, fresh bootstrap on the same old authority, v2 observation/ACK, then warm rotation. Legacy binary replacement and adopted-owner migration are not qualified by that test; rollout must not present them as automatic same-lease upgrades. - This pull request has no database migration or chat-channel activation. The second pull request keeps the channel feature experimental. ## Model Used OpenAI Codex assisted with implementation, tool execution, tests, and reconciliation. The existing implementation records OpenAI `gpt-6-astra` assistance. The current environment does not report a context-window size. No private reasoning traces are included. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bf753b997a |
fix(connections): keep task context through Cloud enrollment and OAuth (#13098)
## Thinking Path
> - Paperclip lets people manage agent work from tasks.
> - Agents can request app access in a task card.
> - Some apps first require Paperclip Cloud enrollment.
> - Enrollment could leave the task, and OAuth could lose the task
interaction ID.
> - This PR keeps enrollment in a separate window and retains the
interaction ID through OAuth.
> - The task can then recognize the connection and continue
automatically.
## Linked Issues or Issue Description
**What happened?**
A first Gmail connection could leave the task dialog during Cloud
enrollment. Setup resumed on the Apps page. Gmail connected, but the
task card could remain pending because OAuth did not retain its
interaction ID.
**Expected behavior**
Keep the task open and preserve its access choices. Resolve the card
after the server verifies connection access. Continue the agent
automatically.
**Steps to reproduce**
1. Start a fresh source test-drive instance without Cloud enrollment.
2. Ask an agent to read Gmail.
3. Open Connect on the task card.
4. Complete Cloud enrollment and Gmail authorization.
5. Check whether the task card updates without selecting the connection
again.
**Paperclip version or commit**
Reproduced on
|
||
|
|
82f662656a |
fix(runner): restore legacy Git access and independent networking (#13094)
## Thinking Path
> - Paperclip runs agents for people with different GitHub accounts.
> - Managed operations must use the intended person's eligible
connection.
> - A failed duplicate connection must not hide a healthy grant for the
same account.
> - Legacy hosts also need their existing Git configuration when managed
access is not configured.
> - Runner networking and local Git operations must not depend on GitHub
broker availability.
> - This pull request separates those policies and improves failure
diagnostics.
## Linked Issues or Issue Description
**What happened?** New runs always cleared host Git credentials and
installed managed launchers. Network permission depended on GitHub
environment variables. A launcher failure could stop even local `git
status`. A newer unhealthy duplicate could take precedence over a
healthy connection, and generic health errors were shown as reconnect
requirements.
**Expected behavior:** Use a healthy eligible managed connection for the
intended account. Preserve host authentication only for unconfigured
standard-trust local or SSH execution. Permit local Git during broker
failures and keep network permission independent of GitHub credentials.
**Steps to reproduce:** Configure healthy and unhealthy grants for one
GitHub account, dispatch an agent, and execute Git commands. Separately
run an unconfigured legacy host with existing GitHub CLI authentication.
Stop the broker and run local `git status`.
**Paperclip version or commit:** Master at
|
||
|
|
35fdc0c66b |
fix: make task recovery durable and preserve current requests (#13075)
Make task recovery durable and preserve the latest user request across native and legacy continuations. Keep routine recovery quiet and prevent replay when action outcomes are uncertain. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5acf56658b |
feat(onboarding): first task opens as a chat with a chief of staff (#13068)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Onboarding ends by handing a new user to their first agent on a seeded first task > - Today the wizard asks for a mission up front, the UI composes what the agent is told, and the agent starts running before the user says anything > - New users get a cold, ticket-shaped start, and nobody can edit the agent's brief or persona without a code change > - This pull request makes the first task a short chat: a four-step wizard, a chief-of-staff persona, a greeting plus a two-option opening card, server-owned markdown texts, and no run until the user answers > - It also gives question cards one consistent action row (Cancel / Skip / Next), makes agent hires idempotent within a run, and turns the Paperclip Runner flag on by default for self-hosted instances > - The benefit is a first run the user steers, with texts a board operator can edit as markdown ## Linked Issues or Issue Description No public GitHub issue exists for this change. The feature request fields follow. Related PRs and issues: - Refs #11043 — an earlier draft of the first-task onboarding experience. This PR supersedes it. - Refs #11280 — a report about the onboarding first-task route test. This PR extends that test file. ### Subsystem affected Onboarding wizard, the seeded first task and its texts, task-chat question cards, agent hiring, and the instance experimental settings. ### Problem or motivation The onboarding wizard collects a mission through two extra steps and a questionnaire. The UI then composes the first agent's instructions and the first task description from those answers. The first task wakes the agent at once, so the agent runs and posts before the user types a word. Board operators cannot change the greeting, the brief, or the persona without editing TypeScript. Question cards in chat behave differently per adapter, and a single-select pick submits on click. A misread hire response could create a duplicate agent that the creating agent cannot remove. ### Proposed solution Reduce the wizard to four steps and stop the UI from authoring agent texts. Move the greeting, the brief, the chief-of-staff persona, and the opening question into markdown and JSON files that the server loads at runtime. Seed the persona onto the first agent through an explicit hire marker. Do not wake the first task until the user answers the opening card or types. Give every question card the same Cancel / Skip / Next actions. Add an experimental toggle that switches the single-task proposal between one confirmation card and a plan document with a checkbox card. Make agent hires idempotent within a run. ### Alternatives considered - Keep the mission questionnaire and feed it into the brief. Rejected: the agent asks better questions in chat, and the wizard gets shorter. - Keep the first task open-ended with a plain composer. Rejected: a two-option card gives the user a clear first move. - Derive the plan-document behaviour from the user's intent only. Rejected in favour of an explicit experimental toggle so operators can choose. - Key the "pick does not submit" behaviour off the presence of a submit label. Rejected: several adapters set a submit label on single-select cards, and their cards would change behaviour. ### Roadmap alignment `ROADMAP.md` lists no planned core work on onboarding or the first task. This change refines the existing flow and does not duplicate planned work. ## What Changed - Wizard: four steps (Name your organization, Create your first agent, Connect a model, Review). The front door and both mission steps are removed with their state and saved-progress keys. The UI no longer composes the first agent's instructions or the first task description. - Server-owned texts: the greeting, the brief with two proposal variants, the chief-of-staff persona, the opening question, and a README live in `server/src/onboarding-assets/first-task/` and load at runtime. The create route stores the assembled brief and ignores any client description. - Persona seed: an `onboardingFirstAgent` marker on the hire lets the server seed the chief-of-staff persona over the first agent's entry file. Board-authored hires only. The persona tells the agent the hire response shape and to list agents before it acts on an unclear result. - No auto-run: the first task does not queue an assignment wake. The stranded-assignment reconciler leaves it idle until a user comment or an answered card exists. - Opening card: the server seeds an `ask_user_questions` card right after the greeting with two options: "Interview me and propose a plan and an agent team to execute it." and "I have a task in mind" with free text. Answering wakes the agent. - Experimental toggle `enableFirstTaskPlanProposal` (default off): the single-task proposal is one confirmation card, or a plan document plus a checkbox card when on. - Question cards: every `ask_user_questions` card renders Cancel, Skip, and Next (the submit label on the last question). Skip hides on required questions. Picking an option no longer advances or submits by itself. - Wizard guards: the dashboard's agentless offer ignores a cached empty agent list while a refetch is in flight. The hire step adopts an agent that already carries the typed name instead of hiring "Name 2". - Agent hires are idempotent within a run: a retry of the identical request under the same run id returns the existing agent with `200` and `idempotent: true`. The fingerprint covers the whole validated request, so a corrected payload is a new hire. Lookup, create, and activity record run under one lock per company and run, so overlapping retries cannot both create. - The Paperclip Runner experimental flag defaults to on for self-hosted instances. Cloud keeps its declared default: a managed instance whose tenant row and managed overlay omit the flag resolves it to off. - Question cards: a send that finds an earlier required answer missing returns to that question with a message instead of failing silently. - The two onboarding e2e specs follow the new wizard: the front door and growth intake shots are gone, and the planning-mode spec dismisses the opening card before it reads the composer. - Docs: `docs/board-operator/editing-first-task-texts.md` explains how to edit the texts and the toggle. ## Verification Commands, run from the repo root: ``` pnpm -r --filter './packages/*' --filter '!@paperclipai/paperclip-runner' build pnpm --filter ./packages/shared typecheck pnpm --filter ./ui typecheck pnpm --filter ./server exec tsc --noEmit pnpm check:token-gates pnpm --filter ./ui exec vitest run OnboardingWizard onboarding QuestionForm InteractionCard ProtocolCard TaskChatComposer Dashboard feature PAPERCLIP_IN_WORKTREE=false pnpm --filter ./server exec vitest run onboarding-first-task heartbeat-process-recovery agent-hire-idempotency instance-settings agent-skills-routes issue-onboarding onboarding-greeting --testTimeout=90000 ``` Results on this branch: - Typecheck is clean for shared, ui, and server. - Token gates: 4 of 4 clean. - UI: 344 tests pass across 23 files. - Server: all suites pass. The first test in `agent-skills-routes` has its own 10 s cap and needs about 15 s on my laptop for the app cold start. It passes with a longer cap. This PR does not change that cap. Manual steps on a dev instance: 1. Open `/onboarding`. Confirm four steps: Name your organization, Create your first agent, Connect a model, Review. 2. Finish the wizard. Confirm the first task shows the chief-of-staff greeting and the opening card with two options. Confirm no run starts. 3. Pick "Interview me…". Confirm no run starts. Press Continue. Confirm a run starts and an interview card of 3–4 questions arrives. 4. On a fresh organization, pick "I have a task in mind", type a task, and press Continue. Confirm a proposal arrives as one confirmation card. 5. Turn on Settings → Experimental → "First task: propose with a plan document" and repeat step 4. Confirm a plan document and a checkbox card arrive. 6. Visit the dashboard after the hire. Confirm the wizard does not reopen and one agent exists. 7. Open any question card. Confirm Cancel returns the plain composer with the card still pending, Skip advances an optional question, and Next moves to the next question. Design reference with flow diagrams, chat mock-ups, and live captures: https://pages.paperclip.ing/first-task-flow/proposed/ ## Risks - `pnpm dev` now builds the runner daemon because the Paperclip Runner flag is on by default. Developers without a Rust toolchain must set `PAPERCLIP_RUNNER_BINARY` or turn the flag off. Self-hosted instances that never set the flag now let qualified agents use the runner. - The wizard drops the mission steps and their saved-progress keys. A user who is mid-wizard on an older build restarts at step 1 after an upgrade. Existing organizations are not touched. - The first task no longer runs on its own. A user who neither answers the card nor types sees no agent activity. This is intended. - The persona seed applies only to hires that carry the marker from the wizard. API hires are unchanged. - Hire idempotency is scoped to one run id and to the exact request. Retries across runs, or with a changed payload, still create a second agent. The lock is per server process, which matches how an instance serves its API. - Single-select question cards no longer submit on pick. Users of adapters that relied on that behaviour now press Next. - No database migrations. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic) through Claude Code. `claude-fable-5-1` with extended thinking, tool use, and code execution wrote most commits. `claude-opus-4-8` wrote the toggle, texts, wizard, and idempotency commits, as the `Co-Authored-By` trailers show. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e200104727 |
feat: review connection actions from tasks (#13063)
Bring governed connection reviews into task history and composer approvals. Share resolution with Connections, add scoped remembered permissions, and resume agents through durable outcome receipts. Keep cards compact, collapse raw results, isolate untrusted provider output, bound continuation payloads, and reconcile missed live events. Add Storybook coverage, browser journeys, and service regression tests. Verification: all PR CI gates passed, Greptile 5/5, security scans passed, five connection-review browser journeys passed, and real native Codex approval/continuation was verified against the local MCP fixture. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2043e0c735 |
fix: repair runner configuration, macOS execution, and artifact galleries (#13062)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters select a provider, a model, and a runtime. > - Runner conversion rejected existing Claude agents. The model list mixed providers. > - The native Claude runner rejected custom models and could not launch on macOS. > - This pull request fixes conversion, model selection, and verified macOS execution. > - It also groups configuration fields consistently across adapters and opens artifact images in the task gallery. > - Operators can change an agent configuration and run the selected model on their Mac. ## Linked Issues or Issue Description **What happened?** Converting an existing Claude agent to Paperclip Runner failed with a Codex-only restriction. ACPX Claude showed unrelated models and required `claude-sonnet-5`. Its native runtime rejected macOS. Configuration mixed common model settings with process controls. Artifact cards labeled “Open gallery” navigated to attachment URLs instead of opening the task gallery. **Expected behavior** Conversion keeps agent identity and compatible settings. ACPX Claude uses the normal Claude catalog and accepts typed model IDs. Codex uses the native runner. The verified Claude runtime can launch on macOS ARM64 and x64. Common configuration sections place the same fields together across adapters. Artifact images open in the shared task gallery with navigation and downloads. **Steps to reproduce** 1. Open the configuration of an existing Claude agent. 2. Convert it to Paperclip Runner. 3. Select ACPX Claude and a different catalog model or a typed model ID. 4. Save the agent and run a disposable task on macOS. 5. Inspect configuration and advanced run-policy controls across adapters. **Paperclip version or commit** The bugs were reproduced on `165ca56a22adb60e5fda56045442d9c8498116a8`. This branch was rebased onto `7ed122911`. **Deployment mode** Built from source. Local test-drive instance on macOS ARM64 with an isolated database. Related work: #11798 addresses unsupported ACP session options in the existing adapter path. #13048 addresses working-folder preservation. This change fixes native runner configuration and launch behavior. ## What Changed - Remove the Codex-only conversion restriction. Preserve agent identity, instructions, directories, credentials, and compatible model settings. Reset incompatible sessions while retaining history. - Show ACPX Claude and native Codex as distinct provider choices. Remove ACPX Codex from advertised configuration. Normalize legacy configurations before fresh runs without rewriting historical run descriptors. - Select model catalogs and cache entries by provider. Support refresh and typed model IDs. Pass exact Claude IDs through session creation, model changes, and recovery. - Add verified macOS ARM64 and x64 Claude SDK snapshots. Bound executable allocation and total snapshot size. Preserve package checks, dependency isolation, process ownership, cancellation, and Linux descriptor loading. - Probe local runtime readiness. Report remote platform checks as incomplete until the remote runner verifies its runtime. - Surface actual model rejection and allow correction and retry. - Repair missing ACPX goal-capability helpers exposed by the post-rebase live test. Persist and restore the optional capability without breaking session startup. - Put Agent identity first and intentionally remove the Capabilities editor, as requested. This is removal of UI editing, not relocation: preserve existing capability metadata and API compatibility without adding another editor. Use the themed select for configurable permission modes, with normal text instead of monospace. - Put model and provider under Adapter. Give environment variables their own section. Fold command and arguments under Configuration. Fold lifecycle, timeout, and interrupt grace under Advanced Run Policy. Hide single-option permission controls. - Open image and video artifact cards in the existing task gallery, including cards in the artifacts panel. Chat attachment images use the same gallery. Preserve standalone media previews and download links. ## Verification - Rebased focused UI/API/database suites: 293 tests passed. - Rebased native runtime and ACPX suites: 242 passed, 7 skipped. - Repository typecheck, build, and token gates passed for the runner changes. Gallery follow-up UI typecheck, build, and token gates also passed. - Follow-up UI suites passed (86 tests), packaging checks passed (14 tests), and the final focused runtime suites passed (126 passed, 7 skipped). - Linux container isolation and lifecycle fixtures passed before rebase (57 passed, 2 skipped). Rust ACPX provider-session tests passed after rebase (8 tests). - Browser tests completed actual Claude and native Codex tasks on macOS ARM64. They covered conversion, catalog refresh, a non-default catalog model, a typed `haiku` ID, save/reload, cancel, follow-up session continuity, invalid-model errors, and recovery. - Final-revision live tests completed a typed Claude task, a follow-up with the same provider session, and a native Codex task on macOS ARM64. - Browser tests confirmed the moved interrupt-grace field saves and survives reload. Cross-adapter tests cover Claude, Codex, Gemini, process, gateway, and schema forms. - Full local run: 7,080 passed, 30 skipped, and two timeouts. Both timeout suites passed on isolated rerun (84 tests); the failures were the plugin login-worker exit diagnostic and the runner real-server vertical slice. - Final follow-up checks: 50 registry tests and 45 snapshot/installation tests passed (6 platform-specific skips). Oversized executable rejection is covered before allocation or reading; unsupported-platform tests invoke the real installation probe. - Runner head `ddb5101c483a297f74875ab96b3c66035b002d50`: all CI gates green, including full runner verification, repository build, typecheck, general/serialized server suites, browser tests, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34286178670). - Greptile: 5/5 on that runner head. All four review threads resolved. Superagent, Socket, and Snyk checks green. - After snapshot hardening, another real Claude task completed on this Mac using the rebuilt runtime. - Gallery follow-up: 148 focused tests passed, covering artifact selection, shared attachment collections, deduplication, image/video cards, standalone previews, downloads, and closing. Live browser verification completed on the settings follow-up: artifact selection, 6-image pagination with wrapping, download action, and closing all stayed on the same task URL. All checks passed on gallery head `96136da58ff195bf6ca00b281eb3022ad12d7bd8`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34287987536). Greptile returned 5/5 on that exact head with no unresolved threads. - Final settings polish: 96 focused tests, UI typecheck/build, and token gates passed. A real browser walkthrough verified readable permission options, identity placement, Capabilities removal, and permission save/reload. Original test-agent permission mode restored. All 31 checks passed on final head `e46540d6bf32bfb0566dca16b2f4a75ba437618c`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34292797886). Greptile returned 5/5 with no unresolved threads. ## Risks - Capabilities intentionally has no editable UI field after this change. Existing values remain readable and API-compatible; removing the field does not erase stored metadata. - macOS launch now copies verified package files into private snapshots. The implementation must retain isolation and clean up snapshots on exit. - Runtime provider or model changes reset the current session. Historical runs remain available. - The macOS x64 SDK executable digest was verified, but a live Intel Mac run was not available. Linux verification used container fixtures, not a real Claude task. - Remote environment tests report a warning when only the platform has been checked. They do not claim package readiness from the server host. ## Model Used OpenAI Codex, based on GPT-6. The exact served model identifier and context-window limit are not exposed in this session. Used reasoning, repository inspection, code execution, Rust and TypeScript tests, and browser automation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites and both timeout suites on rerun; full-run counts above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7ed122911b |
Add end-to-end session goals to Paperclip Runner
Add capability-aware slash-goal controls, durable provider goal state, PRP v2 negotiation, autonomous goal execution, and safe local session recovery. Integrate with current master, preserve provider session identity, and verify the browser goal/chat/replacement/clear workflow and unsupported-agent rejection. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6019e2bd6e |
feat(adapter-utils): carry binary bodies and attachment routes over the HTTP/2 sandbox bridge (#12923)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agents in local and remote sandboxes through adapter utilities > - The HTTP/2 sandbox bridge decoded every body as UTF-8 text and rejected non-JSON content > - This stopped agents from uploading or downloading issue attachments through that bridge > - This pull request carries raw bytes, permits the two attachment routes, and enforces a shared body limit > - The benefit is correct attachment transfer with a process-wide memory guard ## Linked Issues or Issue Description **What existing behavior does this improve?** The HTTP/2 sandbox bridge forwards request bodies between an agent sandbox and the Paperclip host. It now supports binary bodies and the issue attachment routes. **Current behavior** The bridge decodes each body as UTF-8 text. It returns HTTP 415 for content types outside the JSON route list. An agent cannot upload or download an issue attachment through this transport. **Proposed behavior** The bridge carries raw bytes through the forward path. It permits the attachment upload and content routes. The queue transport and file gateway keep their existing route behavior. A shared 10 MiB body limit and process-wide byte reservation protect memory use. **Reason and benefit** Attachment clients need byte-preserving transfer. The shared limit keeps the gateway and host aligned. The reservation prevents concurrent streams from exceeding the accepted process memory ceiling. **Breaking changes** The HTTP/2 bridge accepts two attachment routes and permits binary content. The queue transport and file gateway keep their previous route lists and HTTP 415 behavior. No schema or external endpoint changes. ## What Changed - Carry request and response bodies as raw bytes through the HTTP/2 bridge. - Permit attachment upload and attachment content routes on the HTTP/2 bridge only. - Raise the resolved per-body limit to 10 MiB and share it between the gateway and host. - Reserve body bytes before allocation and release each stream reservation on every terminal path. - Document the body limit, process ceiling, and reservation behavior. ## Verification - Run `pnpm exec vitest run packages/adapter-utils/src/http2-bridge-server.test.ts packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapter-utils/src/sandbox-callback-bridge.test.ts`; 226 tests pass. - Run `pnpm --filter @paperclipai/adapter-utils typecheck`; it passes. - Run the direct server TypeScript check with `tsc --noEmit` in `server/`; it passes with zero errors. - Verify multipart upload and binary download round trips over HTTP/2 without corruption. - Verify the queue transport and file gateway return HTTP 415 for the same routes. - Verify the host rejects bodies over the resolved limit. - Verify a denied reservation returns HTTP 503 and allocates no copy. - Verify stream cleanup releases reservations after completion, error, abort, timeout, and close. ## Risks The bridge now accepts larger bodies and binary content. The process-wide reservation limits total live body bytes to 1 GiB. Route behavior changes only for the HTTP/2 bridge. The security review found no blocking issue for this commit range. ## Model Used OpenAI Codex, GPT-5. The runtime used tool calls and code execution. The runtime did not expose the context window size. No model-generated code changes were made for this pull request. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e095b84dab |
feat(connections): connect services from native task feeds (#13058)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents use connections to reach external services. > - A fresh native task can have no service tools installed. > - The agent needs a way to discover services and ask the responsible person for access. > - This pull request brings the existing connection-intent flow into native task execution. > - The person can connect from the task, and the agent can continue with updated tools. ## Linked Issues or Issue Description **Subsystem affected** Native runner tool authority, connection intents, task interactions, and shared connection setup. **Problem or motivation** A task that needs an unconnected service cannot finish its work. Leaving the task to configure access also loses context. A resolved request must survive a restart and resume the correct agent once. **Proposed solution** Expose connection discovery and access requests as server-owned native tools. Render a durable task card and use the shared setup dialog. Persist outcome delivery and start a fresh provider session after access is ready. **Alternatives considered** Sending the person to the Connections page adds navigation and does not solve continuation. Polling for authorization consumes runs and can create duplicate requests. **Roadmap alignment** This extends the existing connection-intent runtime and setup experience. It reuses the shared access model and the native runner. Related: #12345, #12347. The service-slug fix in #12906 is related but separate. Companion evaluation PR: https://github.com/paperclipai/paperclip-evals/pull/21. ## What Changed - Expose `connections_search` and `connection_request` with server-bound company, task, agent, and responsible user. Preserve the legacy entry points. - Discover catalog services and authorized custom connections. Check installation, identity, health, and executable permissions before reporting ready. - Keep pending cards through ordinary messages. Reuse requests and retire stale ownership. Put Connect at the right of Not now. - Reuse the shared setup flow in a task dialog. Keep access additive and default to the requesting agent. Recover from cancelled or blocked OAuth windows with a new-tab fallback. - Persist outcome delivery with an idempotent wake key. Resume in a fresh session and recheck ownership before dispatch. - Add native browser fixtures, offline Storybook states, server contracts, and evaluation fixtures. Update guidance and documentation. ## Verification - `pnpm build`: passed after replaying the change on current master. - `pnpm -r typecheck`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed. - New continuation-policy regression cases: 16 passed. - Docker-backed PostgreSQL regressions passed for requester-only OAuth access, assignment-only expiry, terminal expiry, and credential-free setup metadata. - Shared setup and task-card UI tests: 121 passed, including configured MCP reconnect URL recovery and preserving user edits across refetch. - Storybook browser checks: all 119 passed on the latest reconnect fix. - `pnpm test:run`: 4,734 tests passed in the first server group, but embedded PostgreSQL startup failures and resulting cleanup errors prevented a complete local pass. All Linux CI lanes passed on the latest reviewed commit. One external-object route test returned an unexplained 500 on the first run; it passed twice locally and the failed shard passed on retry without code changes. - Earlier feature-checkout evidence: three deterministic native browser journeys passed, including restart delivery and an actual fixture tool result. Legacy scripted coverage also passed. All 59 added stories were inspected in light and dark themes. - Live Notion testing recorded successful provider reads. The manual test used a local-trusted instance. It does not prove authenticated/cloud deployment or every provider journey. - Native browser rerun reached the embedded PostgreSQL startup limit before bootstrap, so the latest checkout’s full native browser journey remains unverified. Both OAuth page/task regression cases passed against isolated Docker-backed PostgreSQL 17. They verify no premature task access, requester-only completion, additive retries, and reconnect preservation. - Applied both new migrations twice to isolated PostgreSQL 17. Foreign keys remained intact, duplicate active delivery keys were rejected, and failed delivery records did not block retries. Reviewer path: start a fresh test drive, enable the native runner, use an agent that can perform work directly, and ask it to summarize a Notion page. Connect from the card, then verify the resumed provider call and source-linked answer. The default test-drive CEO is instructed to delegate, so it can introduce an unrelated hiring step. ## Risks - Two additive migrations create durable deliveries and a partial unique wake index. They are idempotent. The wake index can require a maintenance window on large tables because migrations run in a transaction. - OAuth and continuation cross asynchronous boundaries. Tests cover ownership changes, retries, additive access, and restart delivery; live provider behavior still varies. - The latest requester-scope fix has not yet been exercised through live OAuth. GitHub, API-key, authenticated-user, and all recovery journeys are not claimed as verified. ## Model Used OpenAI GPT-6-based Codex assisted with implementation, tests, and review using tools and code execution. The runtime does not expose the exact model version, context window, or reasoning setting. Live evaluation used `gpt-5.6-luna`; manual native testing used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used and disclosed unavailable runtime details - [x] I have checked ROADMAP.md and confirmed this extends existing connection work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run all required tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5128b4f323 |
fix(claude): default unset models to Opus 5 (#13055)
Resolve unset Claude models to Opus 5 across CLI and ACP execution, preserve explicit and provider-specific overrides, and show the default in agent configuration. Verified 212 focused tests after merging master, UI typecheck and token gates, and all CI checks. Greptile reviewed the final head at 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ebaeba40ee |
feat: simplify agent onboarding and configuration (#13011)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators create agents and configure their runtimes in the board UI. > - The old creation flow presents several choices and a large form before an agent can start. > - The existing onboarding controls already provide clear provider connection steps. > - This pull request uses those controls in a new-agent wizard and organizes the full configuration pages. > - Operators can connect, test, save, and assign a first task while keeping the existing configuration tools. ## Linked Issues or Issue Description Related: #10974. That earlier open PR also reorganizes agent configuration. This PR follows the reviewed Storybook designs for agent creation and the current configuration tabs. **What existing behavior does this improve?** Agent creation, provider connection, runtime tests, and full agent configuration. **Current behavior** The creation dialog leads to a large manual configuration form. Provider login controls differ from onboarding. Environment variables and secret access appear in separate places. **Proposed behavior** Choose a name and adapter. Connect Claude or Codex through the existing onboarding controls. Configure and test the runtime, save the agent, and open a task dialog with that agent assigned. Use the same design on the existing configuration tabs. **Reason and benefit** The first setup asks for fewer decisions. The full editor keeps instructions, skills, runtime controls, secret access, permissions, keys, and revisions available in clear sections. **Breaking changes** The board creation and configuration layouts change. The test-environment API adds an optional, allowlisted `testCredentials` field for one-shot probes. Database contracts stay the same. Native ACPX tests now reject unsupported local platforms before a CLI login can mask the runtime restriction. ## What Changed - Added a new-agent wizard with numbered steps, adapter branding, provider connections, editable model choices, runtime tests, and confirmation. - Added Codex app-server, Claude ACPX, and OpenCode runner choices. - Stored API credentials through existing secret APIs and persisted references in agent configuration. New setup keys are isolated from credentials used by existing agents. - Preserved external-agent invitations beside the wizard, including optional messages, one-time prompts, and clipboard fallback. - Added OpenRouter provider and secret bindings for Pi and OpenCode. - Added adapter-specific prerequisite fields for Cursor, Gemini, Kimi, and Hermes. Cursor Cloud keys are saved as new organization secrets. - Fixed Cursor Cloud repository field mapping, omitted empty remote environment values, and added useful model and repository error messages. - Preserved complete MCP assignments when multiple valid profiles contain more than 250 tools in total. Generated profiles retain exact tool selectors. - Added service branding and deployment-aware adapter choices. Cloud setup offers Claude, Codex, and OpenCode; local native runners require the experimental setting. - Made the agent list responsive at intermediate widths. - Applied the reviewed design to the real agent configuration pages. Kept the instruction editor, skills, and existing mutations. - Combined secret access and environment variables under one Save and Discard action. - Added interactive Storybook screens for setup, configuration, confirmation, authentication, and test results. - Fixed Pi provider-error parsing and thinking-effort persistence. Native ACPX validates Linux x64 on the actual local, SSH, or sandbox target. - Redacted the complete transient probe-credential field from HTTP error logs, including rejected provider names. ## Verification - Current head `df0292fe6` has a fresh Greptile 5/5 review with no unresolved findings. All 31 executed CI checks passed, including the aggregate verification gate and all browser E2E shards. Storybook visual regression is skipped by its workflow; the local Storybook build passed. - Browser tests completed real assigned tasks with direct Codex, Claude, OpenCode, Pi, and native Codex. - Verified external-agent invitation generation and automatic prompt copying in the live browser. - Pi and OpenCode used an existing OpenRouter secret. Browser checks covered save and reload, instruction edits, skill selection, environment-variable Save and Discard, and assigned task creation. - Invalid Claude API credentials remained on the connection step with an error. A live Pi/OpenRouter invalid-key probe returned a provider failure and left the user-secret inventory unchanged (zero entries before and after). - Full workspace typecheck and build passed after rebasing onto current master. After review fixes, server and UI typechecks, token gates, and the full build passed again. Storybook built successfully. - All 5,542 local UI tests passed. The Cursor Cloud and Pi adapter regressions passed all 24 tests. Review regressions passed 69 server tests and all 18 agent-list tests. - The local full test command ran 6,971 general server tests successfully. Editing review fixes during that long run caused nine tests to use stale modules; fresh isolated runs passed. An unrelated embedded-Postgres fixture hit the host shared-memory limit; its 15 affected tests passed when the fixture groups ran separately. - Local workspace groups passed after rerunning 18 CLI tests sequentially to avoid host database limits and parallel-load timeouts. The local full command stopped at the general server phase, so serialized server verification comes from the five passing CI shards. - Browser testing at 390px confirmed that the agent action menu opens and the page has no horizontal overflow. CI browser E2E shards passed. - Review the `Onboarding / New agent` and `Agents / Configuration refresh` Storybook groups. In the real app, create an agent, run its connection test, save it, assign a task, and reload its configuration. ## Risks - This changes the main agent setup and configuration UI. Regression tests cover routing, persistence, secret bindings, and form actions. - Native Claude ACPX requires Linux x64. Direct Claude works on macOS. Remote checks execute a bounded platform probe and reject unsupported or unverified targets. - A native OpenCode task reached the provider context limit because of its tool payload. Its provider connection test passed. Direct OpenCode completed a task. This existing native execution limit is not fixed here. - Claude and Codex connection keys use the existing user-secret store. Other runtime setup keys use distinct organization secrets. Existing credentials are never rotated. Probes do not store entered keys. Failed agent creation removes newly staged credentials. - Cursor Cloud has not completed a live task. Its authenticated account still needs GitHub repository access. The live run passed MCP provisioning, remote environment validation, and explicit Auto model selection before the repository prerequisite blocked execution. - Generated runtime MCP profiles can exceed the public profile-edit request limit. They still contain exact catalog selectors and preserve permission boundaries. - No database migration, dependency, lockfile, or workflow changes are included. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell execution, and browser automation. The runtime did not expose the exact model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f65991a5f1 |
refactor(server): move scheduled-retry and queued-run dispatch into a run-dispatch module (#12920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server heartbeat service dispatches scheduled retries and queued runs. > - The service kept policy decisions and database writes in one large file. > - This layout made policy branches harder to test and transaction boundaries harder to inspect. > - This pull request moves the policy rules and database transactions into a run-dispatch module. > - The benefit is a smaller service, pure policy tests, and clear transaction ownership. ## Linked Issues or Issue Description **What existing behavior does this improve?** The server heartbeat service promotes scheduled retries and cancels stale queued runs. **Subsystem affected** server/ — REST API and orchestration services. **Current behavior** The heartbeat service contains the policy rules and the database writes for these dispatch paths. **Proposed behavior** A run-dispatch module owns pure policy functions and semantic database transactions. The public service contracts stay unchanged. **Reason and benefit** The new layout separates branch rules from database effects. It makes each policy branch easier to test and keeps each operation’s row writes in one transaction. **Breaking changes** None. The public service contracts stay unchanged. ## What Changed - Move scheduled-retry promotion and queued-run staleness rules into pure functions. - Add table-driven unit tests for each policy branch. - Move promotion and cancellation writes into semantic transactions. - Keep row locking, company isolation, and post-commit effects unchanged. ## Verification - `node scripts/check-module-boundaries.mjs` passes. - `tsc --noEmit` from `server/` reports no errors. - The focused server test command passes 228 tests in six files. - Full pull request CI passes. - Greptile reports 5/5, and all review threads are resolved. ## Risks The main risk concerns changed transaction boundaries in scheduled-retry promotion and queued-run cancellation. The focused tests retain coverage for locking, transactionality, company isolation, and transport contracts. The public service contracts do not change. ## Model Used Codex, GPT-5, with code execution and tool use. The context window is not provided by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I have addressed all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
be6bb768b1 |
fix(ui): restrict company navigation to accessible memberships (#13039)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board selects a company before it loads that company's inbox and tasks. > - Instance administrators can list companies where they have no membership. > - The board treated that directory as a list of companies the user could enter. > - This pull request gives navigation a list based on the existing company access check. > - Users can select their companies without landing on an inbox that rejects their access. ## Linked Issues or Issue Description Fixes #6090. Refs #4855 for the related account-recovery case; this PR does not grant company membership. **What happened?** An instance administrator can select a company where they have no membership. Its inbox then shows “User does not have access to this company.” Company directory visibility and access to company contents use different rules. **Expected behavior** Company navigation should show only companies the current user can enter. A stored selection for an inaccessible company should fall back to an accessible company. A direct link to an inaccessible company should use the existing unavailable-company page. **Steps to reproduce** 1. Create Company A and Company B with separate owners. 2. Sign in as an instance administrator who belongs only to Company A. 3. Select Company B through a stored selection or a link with its prefix. 4. Observe that the board accepts the company selection, but company-scoped requests return 403. Related: #10524 lets cloud users enter additional companies where they hold memberships. This fix preserves that access and excludes companies where they have no membership. ## What Changed - Added `scope=accessible` to `GET /api/companies`, using the existing `hasCompanyAccess` predicate. - Changed the board navigation list to request that scope. Instance Access uses a separate unscoped, account-keyed directory so administrators can manage all companies. Membership edits refresh navigation. - Reject empty, unknown, and repeated scope values with 400. Directory loading errors offer a retry before access controls are shown. - Added route tests for cloud, session, board-key, local trusted, non-member, and agent access. - Added client and component tests for navigation/admin request isolation, grants outside the navigation list, self-membership refresh, directory failure recovery, and forbidden administration. - Updated the API guide and OpenAPI document. ## Verification - Latest commit: 36 focused UI tests passed. The broader UI shard passed all 281 files / 2,533 tests after correcting an asynchronous test assertion. - Server authorization and OpenAPI regression suites: 31 tests passed. - UI typecheck and `pnpm check:token-gates`: passed. - `pnpm -r typecheck` and `pnpm build`: passed after review fixes. - Full local test runner: exercised the supported shards. Several unrelated suites hit embedded PostgreSQL startup failures or startup timeouts under local load. The UI regression issue found in the broad run was corrected and its full UI shard passed. These local limitations are not reported as a green full-suite result. - [GitHub CI](https://github.com/paperclipai/paperclip/actions/runs/34233416473): all checks green on `4e1698cd4` — all server and workspace test shards, all browser end-to-end shards, typecheck/release registry, build, canary dry run, policy, and Docker context integrity. Security checks also passed. - Greptile: 5/5 on the latest commit; both initial findings addressed and all review threads resolved. ## Risks - The UI now excludes companies visible only through instance administrator status. Company membership continues to control access to contents. - Additional companies with active memberships remain available. - The client and server changes must ship together. An older server ignores the new query parameter and retains the previous behavior. - No database migration or permission grant changes. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and test execution. The exact served model identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5752d6bd93 |
fix(heartbeat): block runs on a stuck sandbox plugin and re-enable errored bundled plugins at boot (#12957)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents run inside environments. A sandbox environment gets its
sandbox from a provider plugin (for example the bundled
`paperclip.kubernetes-sandbox-provider`), and every run starts by
acquiring a lease through that plugin.
> - When a plugin activation fails once (on a hosted deployment: one
`RPC call "initialize" timed out after 15000ms`), the loader calls
`markError`. That persists `status = error` on the plugin row and
switches off worker auto-restart. Boot activation (`loadAll`), the
bundled-plugin bootstrap and the lazy worker recovery all consider only
`ready` plugins, so the plugin stays in `error` across restarts until an
operator enables it by hand.
> - Every run that needs the provider then fails before dispatch with
`Sandbox provider "kubernetes" is installed via plugin "...", but that
plugin is currently error.` That message matches neither the retryable
classifier (`... but its worker is not running`) nor any configuration
classifier, so the run is recorded as a plain `setup_failed`, the issue
is released, and the scheduler dispatches the same failing run again on
the next tick. On the hosted deployment one company produced about
11,300 identical failed runs, one every 30 seconds, for a week (#12953
is a customer's report of the same condition).
> - Two gaps cause this: the heartbeat treats a condition that only an
operator can change as a transient setup failure, and the bundled-plugin
bootstrap never gives a plugin in `error` another chance even though the
bundle ships with the release image.
> - This pull request classifies the "installed but not ready" lease
failure as `configuration_incomplete`, so the existing recovery path
moves the issue to `blocked` with one recovery action and an actionable
notice; and it re-enables a bundled plugin found in `error` once per
boot, so the next server restart heals the plugin.
> - The benefit is that a stuck provider plugin surfaces as one blocked
issue per task with clear next steps, instead of an endless stream of
identical failed runs, and a restart repairs the plugin without an
operator having to know the plugin API.
## Linked Issues or Issue Description
- Refs #12953 — hosted report: "that plugin is currently error" on every
run for six days, including runs that were retried by hand. This PR
stops the retry loop (issue goes to `blocked`) and makes a server
restart re-activate the bundled plugin. It does not change how a managed
Kubernetes environment is provisioned for a company, which the same
report also mentions.
- Related PR: #9760 pauses the agent for the permanent `Adapter "..." is
not in the configured adapter registry` setup failure. This PR handles a
different permanent condition (plugin not `ready`) and routes it through
the existing `configuration_incomplete` recovery path (issue-level block
with a recovery action) rather than an agent-level pause, because the
gap is on the plugin, not on the agent. The two do not overlap in code
paths.
- No existing issue covers the bundled-plugin re-enable. Bug
description:
**What happened**
A bundled sandbox provider plugin went to `status = error` after one
failed activation. It stayed in `error` across every later server
restart. Every run for every agent on that provider failed lease
acquisition in under a second with `... but that plugin is currently
error.` (`setup_failed`), and the heartbeat kept dispatching new runs
that failed the same way.
**Expected behavior**
A run that fails because its provider plugin is not `ready` is recorded
as a configuration gap and the issue is moved to `blocked` with a notice
that names the plugin and its status, so no further runs are dispatched
until an operator acts. A bundled plugin left in `error` gets a fresh
activation attempt on the next boot.
**Steps to reproduce**
1. Install a sandbox provider plugin and create a sandbox environment
that uses it; make it an agent's default environment.
2. Set the plugin row's status to `error` (or make its worker fail
`initialize` once so the loader does it).
3. Assign an issue to the agent and let the heartbeat run it.
4. Observe: the run fails with `... but that plugin is currently error.`
as `setup_failed`, the issue is released, and the next tick dispatches
another run that fails the same way. Restart the server: the plugin is
still `error`.
**Paperclip version**
master at
|
||
|
|
023e640a7e |
fix(db): reap idle pool connections, name the pool, and end it on shutdown (#12956)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server keeps one postgres.js pool (`packages/db/src/client.ts`, `createDb`) for every query it runs. #10795 made the pool tunable from the environment, but the defaults stayed at the driver defaults: an idle connection never closes, the pool reports itself as `postgres.js`, and no code path ever calls `sql.end()`. > - On a hosted Paperclip deployment the server entered a restart loop (a bundled plugin failure that #12953 describes made every run fail, and the pool saturated). Each generation opened its ten connections, died, and left the backends open on the PostgreSQL side until TCP keepalive reaped them hours later. After about 20 generations the backends exceeded `max_connections`, and every later boot died on its first bootstrap query with `sorry, too many clients already`, before `server.listen()`. The loop could not heal itself. #9555 describes the same shape on a launchd-supervised self-hosted install. > - Three properties of the pool combine to make this possible: idle connections are never reaped, the pool is never ended on any exit path, and an operator cannot even find the leaked backends in `pg_stat_activity` because they carry the generic driver name. > - This pull request gives the pool a 60 second idle timeout and the `paperclip` application name by default, exposes `max_lifetime` and `application_name` through the same `DATABASE_*` environment contract that #10795 introduced, and ends the pool on the orderly SIGINT/SIGTERM path and on the fail-loud startup path. > - The benefit is that a restarting or crash-looping server releases its backends instead of accumulating them, and an operator can see and count Paperclip's connections. ## Linked Issues or Issue Description - Refs #9555 — database connection pool leak causes an infinite restart loop under load. This PR closes the "pool never ends, idle connections never close" part of that report. - Refs #12953 — hosted outage report. The pool exhaustion is the second half of that incident; the first half (a stuck sandbox provider plugin) has its own PR. - Related prior PRs: #9597 and #8780 both propose hard-coded `idle_timeout` / `max_lifetime` values in `createDb`. Both predate #10795 (merged), which made these options environment-driven; this PR builds on the merged shape and adds the shutdown `end()` that neither covers. #4006 and #7481 are closed earlier attempts in the same area. ## What Changed - `packages/db/src/client.ts` - New `resolveDatabaseClientOptions()` applies Paperclip defaults on top of the environment: `idleTimeoutSeconds` defaults to 60 (`DEFAULT_DATABASE_IDLE_TIMEOUT_SECONDS`) and `applicationName` to `paperclip` (`DEFAULT_DATABASE_APPLICATION_NAME`). `createDb` uses it for both the environment path and explicit options. - `DATABASE_IDLE_TIMEOUT_SECONDS` now accepts `0` to restore the driver default (keep idle connections open). Negative or non-integer values still throw. - New environment variables: `DATABASE_MAX_LIFETIME_SECONDS` (positive integer, maps to `max_lifetime`) and `DATABASE_APPLICATION_NAME` (non-empty string, maps to `connection.application_name`). - `postgresJsOptions()` maps the two new options. - `server/src/shutdown.ts` - `finalizeServerShutdown` gains two optional ordered steps: `closeHttpListener` runs first, before the application services stop; `closeDatabase` runs after the application services and before the embedded PostgreSQL stop. A failure in either is logged and does not stop the teardown. Final order: listener → application services → database pool → embedded PostgreSQL → instrumentation → Sentry. - New `closeHttpListenerForShutdown()`: stops accepting requests, closes idle keep-alive sockets, waits up to 5 s for open connections, then closes whatever is left. Requests still in flight are drained while every service is available, and none can reach a route after `sql.end()`, on the signal path and the programmatic path alike (the programmatic path's later `server.close` finds the listener closed and skips). - `server/src/app.ts`: the app shutdown hook (`shutdownAppServices`) now stops the plugin job scheduler, whose tick queries the database, so a programmatic `shutdown()` leaves no timer running against the ended pool. - `server/src/index.ts` - `startServer()` is now a thin wrapper around the boot sequence. When the boot sequence throws after the pool exists, the wrapper ends the pool (and the separate migration pool, when configured) before it rethrows. This covers the `process.exit(1)` path in the main module and the CLI `paperclip run` path alike. - The orderly shutdown passes the same `closeDatabaseClients` to `finalizeServerShutdown`. - `endDatabaseClient` tolerates a client without `$client` (test doubles) and uses a 5 second end timeout. - Docs: `docs/deploy/database.md` gets a "Connection Pool Settings" table with every `DATABASE_*` pool variable, its default and its effect; `doc/DATABASE.md` lists the two new variables. - Tests - `packages/db/src/client-options.test.ts`: parsing of the new variables, `0` for the idle timeout, rejection of malformed values, driver option mapping, and the `resolveDatabaseClientOptions` defaults. - `packages/db/src/client.test.ts` (embedded PostgreSQL): `createDb(url)` reports `application_name = paperclip` for its own backend, and a pool with `idleTimeoutSeconds: 1` has zero backends in `pg_stat_activity` after the timeout. - `server/src/shutdown.test.ts`: the listener closes before the application services, and the database close runs between the application services and the embedded PostgreSQL stop; a failing database close is logged while the teardown still finishes; `closeHttpListenerForShutdown` closes idle sockets and resolves on close, force-closes after the grace period, and is a no-op when the listener was never bound. ## Verification - `pnpm --filter @paperclipai/db typecheck` — passes (`check:migrations` + `tsc --noEmit`). - `cd server && pnpm typecheck` — passes. - `cd packages/db && pnpm exec vitest run src/client-options.test.ts src/client.test.ts src/client-teardown-registry.test.ts` — 9 + 18 + 3 tests pass (the `client.test.ts` cases need embedded PostgreSQL; the new one waits up to 10 s for the idle reap and passed in about 3 s). - `cd server && pnpm exec vitest run src/shutdown.test.ts src/__tests__/server-startup-feedback-export.test.ts src/__tests__/bootstrap-claim-routes.test.ts` — 34 + 11 tests pass. The startup-feedback suite exercises `startServer()` with a mocked `createDb`, which is why `endDatabaseClient` tolerates a client without `$client`. - Manual check for a reviewer: start the server against any PostgreSQL, then run `SELECT application_name, state, count(*) FROM pg_stat_activity GROUP BY 1, 2;`. Paperclip's backends now show `paperclip`. Leave the server idle for more than 60 s and the idle backends disappear. Send SIGTERM and the backends close before the process exits. ## Risks - Behavior change with no environment set: idle pooled connections now close after 60 s. The next query after an idle period pays a reconnect (single-digit milliseconds on a local socket). postgres.js reconnects transparently. Set `DATABASE_IDLE_TIMEOUT_SECONDS=0` to keep the previous behavior. - `application_name` changes from `postgres.js` to `paperclip`. Anything that filtered `pg_stat_activity` on the old name would need an update; nothing in this repo does. - The HTTP listener now closes at the start of the final teardown (after the heartbeat run drain, which still needs the API for running agents). The pool close runs after the application services. A late query from a timer that survived the service shutdown would fail with a driver "connection ended" error instead of running; the known database-backed timer (the plugin job scheduler) is now stopped in the service shutdown. - The listener drain adds at most 5 s to a shutdown while long-lived connections (for example WebSocket clients) are open; after that they are closed forcibly. - `startServer()` is split into a wrapper and the boot sequence. The exported signature and return type are unchanged. - No migration, no schema change. ## Model Used - Claude Fable 5.1 (`claude-fable-5-1`) via Claude Code, extended thinking, tool use (file edits, shell, test runs). The change was produced with the model and reviewed by the submitting human. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_014t3bi2beVNVVHAxK36dmXm --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
c723bb4dfc |
fix(skills): reuse validated runtime revisions during preparation (#13042)
## Thinking Path > - Paperclip manages AI agents and prepares their runtime inputs before each turn. > - Shared company skills are part of those inputs for native and legacy adapters. > - Runtime materialization refreshed the full inventory again for every declared file. > - Remote skill directories were also downloaded and rebuilt on every turn. > - Measured preparation took 42–73 seconds while runner execution took 7–9 seconds. > - This change reads the inventory once and reuses validated installed revisions. > - Agents retain their selected skills while repeated preparation avoids upstream work. ## Linked Issues or Issue Description **What happened?** One 114-skill preparation performed 407 inventory refreshes, 48 directory rebuilds, and 388 GitHub file fetches. Reusing existing local copies took 151 ms. **Expected behavior** Each listing refreshes inventory once. Unchanged installed remote revisions reuse complete, validated local copies. Local edits remain visible. Explicit updates select new revisions. **Steps to reproduce** 1. Import GitHub skills with supporting files. 2. Run an agent turn, then run another with the same installed revisions. 3. Observe repeated inventory scans, downloads, and runtime directory replacement before execution. Related prior attempts: #2330 and #9268 (still open; #9268 last updated July 9). Those use a marker compared with `updatedAt`. This patch follows the required content validation, immutable revision, company isolation, atomic publication, and read-only semantics, and removes refresh-per-file multiplication. ## What Changed - Split public file reading from reading an already loaded skill. Runtime listing refreshes inventory once. - Add a company-scoped revision cache with file manifests outside the delivered skill directory. Fingerprints omit cosmetic metadata. - Validate exact file inventory, sizes, and hashes before warm reuse. Reject traversal and symlinks. Stage complete builds and serialize atomic publication across processes. - Preserve local/catalog direct sources, stored Markdown fallback, explicit version snapshots, and legacy mutable-ref compatibility. Report missing supporting files and keep older valid revisions readable. - Clean both runtime layouts on rename/removal and record `skills.prepare` under preparation timing. - Add service/cache regressions and an isolated 114-skill benchmark, including a new-process warm run. ## Verification - Final targeted skill-service/cache/trace validation: 86 tests pass (61 embedded-PostgreSQL service tests, 19 cache tests, 6 trace tests). Database tests executed rather than skipped. Focused skill routes, adapter selection, and native runtime context also pass. - `pnpm -r typecheck` and `pnpm build` pass locally at `22caa1fe4`. - The full `pnpm test:run` matrix passes on supported Linux CI at the final head: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34236097762). Local full-suite execution encountered PostgreSQL startup contention, a random allocated-port boundary, and a socket hang-up; every affected suite passed on an isolated rerun. The interrupted local serialized run is not claimed as a complete local pass. - Repeatable benchmark: `pnpm --filter @paperclipai/server exec tsx ../scripts/benchmark-skill-preparation.ts`. Mixed 114-skill inventory with 429 remote files on Linux: cold 286 ms, warm median 96 ms / maximum 153 ms including a new process. Every warm sample performs one refresh, zero upstream fetches/rebuilds, and reports no missing entries; content assertions pass. - Controlled deployment against the previously deployed revision completed with zero lost runs. Real inventory: 114 skills, 670 declared files; 402 cached files match the prior installed copies byte-for-byte. Ten post-deployment warm preparations: median 129 ms / maximum 208 ms; new-process warm 194 ms, zero downloads/rebuilds/missing entries. - Five sequential real browser questions persisted in 10.6–20.7 s (median 12.2 s), versus 50–83 s before. Skill preparation median 240 ms, with one 2.37 s outlier. Total preparation median 3.337 s / maximum 8.728 s **does not fully meet** the <3 s / <5 s target. The excluded historical-run redaction query takes about 1.36 s per scan at two preparation call sites; wider application latency coincided with the outlier, without a cache rebuild. These residuals are reported rather than discarded. - Disposable skill reimport verified through actual selected-skill runs: the next run read the changed code. Fixture removed and agent configuration verified unchanged. - Greptile 5/5, zero unresolved review threads, all final-head CI checks green. ## Risks - Cold preparation still requires upstream availability for supporting files. An unavailable revision is reported missing and never falls back to an older revision. - Valid older revisions and quarantined invalid entries consume additive disk space until skill cleanup. An abruptly killed publisher can leave a lock that requires operator cleanup after confirming its PID is dead. - Warm validation reads all cached file bytes. Very large inventories still have proportional local I/O cost. - No HTTP API, schema, agent configuration, or first-party Telemetry changes. OpenTelemetry retains its operator endpoint gate. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository inspection, code editing, and test execution. The exact serving snapshot and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted and isolated reruns; full Linux CI matrix passes, local full-run caveats above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b97101893f |
feat(projects): select multiple GitHub source repositories (#13010)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Projects give tasks a common source repository and execution context. > - The current project form asks for a raw URL and unrelated metadata. > - Teams need to select several repos from GitHub connections they can use. > - This pull request implements the reviewed project form and repository editor. > - The server checks credential ownership and shared audiences before discovery. > - Existing workspace URLs and runtime identity rules remain compatible. ## Linked Issues or Issue Description **Problem or motivation** Project creation accepts one raw repository URL. It does not help users select repos from their usable GitHub connections or attach several repos together. **Proposed solution** Add a shared GitHub repository picker to project creation and Configuration. Support multiple selections, transactional persistence, and the existing GitHub setup flow. Simplify the project form and Configuration tab as reviewed. **Alternatives considered** Keep a raw URL field or add a separate repository table. The existing workspace collection already supports several repositories and keeps legacy URLs compatible. **Roadmap alignment** This builds on the shipped MCP Tool Gateway and Apps capability. It does not change runtime credential delegation. Related work: #11662 addresses the existing dialog's viewport limits. #4552 addresses generic Git URLs; this change preserves those URLs in existing workspaces. ## What Changed - Add company-scoped repository discovery from usable personal and shared GitHub grants, with provider-ID deduplication, PAT pagination, and partial failure handling. - Document the repository endpoints and board access requirements in OpenAPI. - Validate new selections and save projects with multiple repository workspaces in one transaction. Preserve legacy URLs and existing selections whose access was lost. - Implement the reviewed Create project dialog, shared repository editor, scrolling, and mobile layout. - Move repositories above environment variables, remove Status and Goals controls and env help paragraphs, move Created to the bottom, and redirect Overview to Configuration. - Reuse GitHub setup in dialogs, preserve project drafts, and verify popup completion through the API. - Replace the configuration story's DOM adapter with explicit production composition. Keep the reviewed mobile and short-viewport stories. ## Verification - Passed: `pnpm build`, `pnpm -r typecheck`, `pnpm build-storybook`, and `pnpm check:token-gates`. - Passed: focused repository access, database persistence, configuration, and connection setup tests. - Passed: `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/project-repositories.spec.ts`. - The browser tests use a real temporary server/database. They cover create, forty persisted repos, mobile scrolling, save/reload, legacy URL editing, and rejection without a partial project. - GitHub responses and popup completion use deterministic fixtures. No real GitHub account was authorized by the test suite. - All CI general, serialized server, and browser test shards pass on the final commit. - The local full-suite run overlapped review edits and was stopped; fresh repository, OpenAPI, UI/CLI, and connection tests pass. Unrelated local worker, built-in-agent, and routine timing/socket failures passed isolated reruns. - Final commit `1b3308dca`: all CI gates pass, including build, runner verification, typecheck, canary dry run, and security checks. Greptile is 5/5 with no unresolved review threads. - Storybook visual regression is opt-in and was skipped by CI; the Storybook build passed locally. ## Risks - Repository discovery depends on provider availability. Failed connections are reported while successful results stay usable. - Selections identify source workspaces; they do not grant agents new credentials. The existing primary-workspace and responsible-user identity rules still apply. - No database migration is needed. Existing API status, goals, dates, and manual workspace URLs remain supported. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and browser tools. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
297d8741f5 |
fix: resolve duplicate connections to the same GitHub account (#13022)
## Thinking Path > - Paperclip lets people share agents while keeping GitHub access personal. > - Each managed Git or GitHub operation selects an eligible connection grant. > - Connecting the same GitHub account twice creates two grants. > - The old resolver counted grants and rejected them as competing identities. > - Managed commands then ran anonymously and reported a misleading login failure. > - This change compares GitHub account IDs and selects one eligible grant for the same account. > - The benefit is reliable access after reconnecting, with clear diagnostics for real failures. ## Linked Issues or Issue Description Refs #13005. **What happened?** Two active connections owned by one Paperclip user pointed to the same GitHub account. Managed Git refused both as ambiguous. The agent could not push, although the account was connected and had repository access. **Expected behavior** Multiple grants for the same GitHub account resolve to one eligible authorization. Different accounts remain ambiguous. Unavailable access explains its cause without blocking unrelated work. **Steps to reproduce** 1. Connect the same GitHub account twice for one Paperclip user and allow the shared agent through both connection audiences. 2. Start an instruction as that user. 3. Run managed gh or git push. Before this fix, no credential is provided. ## What Changed - Compare stable GitHub account IDs when more than one eligible grant exists. Never deduplicate by login alone. - Prefer an available grant, then the newest authorization with a stable ID tie-breaker. Refresh and webhook timestamps do not change the selection. - Keep the selected credential and connection policy together. Do not combine permissions or fall back from a dedicated account to a personal account. - Print the redacted unavailable reason in managed command output. Unrelated local operations still work anonymously. - Add database and executable launcher regressions, and document selection behavior. ## Verification - Final `pnpm -r typecheck` and `pnpm build` passed. - Fourteen operation credential integration tests passed, covering duplicate personal/dedicated grants, incomplete credentials, distinct accounts with the same login, missing identity metadata, revocation, membership, connection audiences, and A → B → A steering. Existing Git credential and gateway suites and both executable launcher tests also passed. - The local broad test run encountered three embedded-Postgres lifecycle timeouts and stale modules from edits made during that run. A fresh process rerun of all four affected suites passed all 35 tests. The full Node 24 CI test matrix passed on the final commit. - CI passed all 31 checks on `797973b30beb16ba5fa69ed281835e1ab812b449` (Storybook visual regression was correctly skipped). An unrelated Company Settings UI test failed once; the focused local reproduction and rerun of its CI shard both passed without code changes. - Fresh Greptile review of the final commit: 5/5, with no open findings. Security checks passed. - Live acceptance passed with both duplicate connections enabled: managed `gh api user` returned the expected account, managed `git push` succeeded, and the agent created #13023 and pushed its review fixes. No host login or credential changes were used. - Applied the final source/compiled patch to the affected instance with backups, after confirming no runs were active. Restarted service health and the final resolver selection were verified. The patch is an overlay on the existing deployment; this PR supplies the upstream fix. ## Risks The resolver selects one authorization for an already permitted GitHub account. It does not combine repository permissions across connections. If the selected authorization has narrower access, that operation can still be denied by GitHub. Different provider account IDs and unknown duplicate identities continue to fail closed. No schema, host credential, or connection permission changes are included. ## Model Used OpenAI GPT-6 through Codex assisted implementation and verification with shell, database, and browser tools. The exact model variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1cc45086d3 |
feat: use the responsible person's GitHub for shared agent operations (#13005)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Several people can send instructions to the same agent and task. > - A fixed GitHub token in the provider process can keep the first person's access after another person's message is accepted. > - Task ownership cannot select credentials for each accepted instruction or preserve the identity of an operation already in progress. > - This pull request records ordered execution identity contexts and resolves credentials when managed Git, gh, or GitHub tools start. > - The benefit is automatic personal GitHub access for shared agents, with durable continuation rules and no teammate credential fallback. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: orchestration, connection grants, database, runtime adapters, native runners, and run details. **Problem or motivation** A shared agent must use the person whose instructions it has accepted. A queued message must retain its author. A retry or approval without new instructions must retain the originating identity. GitHub must remain optional for ordinary work. **Proposed solution** Persist execution identity separately from task ownership. Give new processes a run-scoped broker capability and token-free managed launchers. Capture identity at operation start. Keep an explicit dedicated-agent grant as an override. Show redacted diagnostics in run details. **Alternatives considered** Per-task ownership, fixed provider tokens, and mutable repository author configuration do not handle accepted steering or concurrent operations. A manual account-selection action would add unnecessary setup to each turn. **Roadmap alignment** This completes the existing Multiple Human Users, MCP Tool Gateway & Apps, Secrets Manager, and Self-healing Runs capabilities. The implementation follows the maintainer-approved plan. Related work: Refs #12843, Refs #12907. Existing proposals #4618 and #8945 cover per-agent or per-worktree author configuration. This change instead follows the accepted human instruction across runtime types. Refs #11831 for governed personal connection delegation; this change preserves connection audience checks and does not use standing delegation as a personal credential fallback. ## What Changed - Add durable, ordered identity contexts and active run references. Preserve message authors through consolidation, steering, retries, delegation, approvals, routines, and restart. - Add an authenticated operation-time GitHub credential broker and local/remote managed git and gh launchers. Keep personal tokens out of the long-lived provider process. - Resolve GitHub gateway and server-side Git operations through the same responsible-person or dedicated-grant selection rules. - Make absent and unavailable GitHub credentials non-blocking at generic startup. Clear host and prior-person credentials. Keep anonymous Git access where supported. - Add run-detail identity history and the dedicated-account warning. Keep task ownership and queue-versus-steer decisions unchanged. - Preserve personal OAuth declarations through connection edits. Retain exact selected grants in the gateway. - Fix continuation races found during real acceptance: verify a warm owner before credential rotation, and wait for bounded durable runner suspension before the next run starts. - Make migrations replay-safe. Retain identity through agent/run deletion, remove it with its company, and clean terminal launcher directories before releasing execution environments. Document coordinated release and rollback. ## Verification - Full workspace typecheck, build, and token gates passed. The complete local suite passed in its normal test groups: 17,120 passing tests, including all 143 serialized server suites. After integrating the newly merged runner API work, full local typecheck and build passed again, along with 890 focused integration tests. All 31 checks on the integrated revision passed, including build, browser E2E, release registry, canary dry run, typecheck, security and all test suites. Greptile is 5/5 with all review threads resolved. - Current focused checks passed: 142 native executor tests, 67 runtime lifecycle tests, 9 durable identity tests, 75 credential/routine tests, 19 low-trust/resumption tests, and the executable migration replay test. - Authenticated browser acceptance with two Paperclip users and two GitHub accounts on one shared native agent passed. Real commits and pushes followed A → B accepted steering → queued A continuation in the same saved conversation. GitHub commit author and committer identities matched all three operations. Both runs succeeded and task ownership stayed unchanged. - Real GitHub MCP calls switched from A to B after accepted steering. A delegated subtask retained its originating identity across a server restart. - Disabling B's GitHub connection left ordinary work successful. Managed gh was unauthenticated and the provider had no inherited GH_TOKEN or GITHUB_TOKEN. - The browser displayed run-detail diagnostics and the exact dedicated-account warning. A final controller-restart check followed by another-person continuation retained the conversation, selected the correct GitHub login and Git author, and removed each terminal launcher directory. - Company-lifetime migration and all five previously failing CI suites passed locally (167 tests). Same-token gateway A → B → A and six broker/launcher boundary tests passed. - Remote callback, launcher, sandbox, and runtime contract tests passed. Both native and legacy Codex completed actual Daytona executions on the integrated revision ([campaign results](https://github.com/paperclipai/paperclip/actions/runs/34155056509)). The remote package-manager shim staging regression also passed locally. ## Risks - Deploy the migrations, server broker, launchers, and runner artifacts together. Existing processes finish with their original contract. New managed processes need the broker endpoint for GitHub operations. - Finish or stop new managed executions before rolling application code back. Keep the additive schema and identity history during rollback. - Scripts that require a persistent raw GH_TOKEN must use managed git, gh, or GitHub gateway tools. Run capabilities authorize code executing within that run to acquire its current identity; this is not hostile-code isolation within one execution principal. Managed commands prevent automatic credential carryover; arbitrary code deliberately copying a credential is outside that boundary. - Uncertain steering acknowledgement deliberately holds new credential acquisition until reconciliation. Already-started operations retain their captured identity. - GitHub private access and provider outages can still fail the specific operation that needs them. Dedicated grant failure does not fall back to personal access. ## Model Used OpenAI GPT-6 through Codex assisted implementation, review, shell execution, and browser acceptance. The exact model variant and context-window size are not exposed in this session. Tool use included TypeScript and Rust tests, database integration tests, GitHub CLI, and authenticated browser control. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5bddff0920 |
feat(runner): add guarded API search and call fallback (#13003)
## Thinking Path > - Paperclip manages AI agents and their work. > - The new runner gives agents dedicated tools for common tasks. > - Some API operations and parameters have no dedicated tool. > - Agents need a controlled way to find and use those operations. > - This pull request adds API search and calls through the real server routes. > - Existing tools remain the preferred path. The new tools are disabled by default. > - Paired tests measure correctness, tool choice, cost and time. ## Linked Issues or Issue Description **Subsystem affected** Paperclip Runner contracts, production tool authority and the server API catalog. **Problem or motivation** The runner cannot use much of the API described by the old Paperclip skill. A generic HTTP client would also let agents bypass runner control rules. **Proposed solution** Add `search_api` and `call_api`. Resolve calls from the mounted API catalog. Use server-held, run-bound credentials. Preserve route checks and runner lifecycle rules. Keep the tools disabled until an operator enables selected companies. **Alternatives considered** A dedicated tool for every endpoint would add a large initial prompt. An unrestricted HTTP tool would weaken authorization and replay controls. **Roadmap alignment** This extends the native runner tooling. The repository owner requested this design and implementation. The roadmap and related open PRs were checked. No duplicate API escape-hatch PR was found. ## What Changed - Register two compact fallback tools in canonical contracts and provider projections. - Build deterministic API discovery from OpenAPI, mounted experimental routes and the old skill reference. - Execute bounded JSON, text, file and download requests through authenticated HTTP routes. - Recheck active runs, company access and work modes. Block runner lifecycle, scheduling, credential and approval bypasses. Keep routine annotation collaboration available. - Retain mutation receipts. Report uncertain outcomes without blindly repeating writes. - Add a company rollout gate and a durable eval worker with complete cost accounting checks. - Record child-task creation in the activity log with the agent and run. - Add contract, authorization, file, replay and real runnerd/PRP/HTTP tests. - Document rollout gates and paid coverage limits. The companion eval repository retains immutable attempts and reports. ## Verification - Final app commit `da58370524c3626a744eec20164397c5fb6ba9ef`: all 32 checks passed; the unrelated Storybook visual check was skipped. Greptile 5/5; no unresolved review threads. - Full Linux build and recursive typecheck passed. Repository tests were run by project and serialized shard; all 143 serialized server suites passed. - Runner TypeScript: 1,599 passed, two skipped. Rust release: 451 passing test reports. Conformance and replay parity passed. The required API check passed 837 tests, including runnerd → PRP → authority → real HTTP. - Bindings cannot enable API tools without the explicit deployment flag. Unit and real-authority tests prove the default-off boundary. - The standalone API check builds and stages its own binary. It passed after existing staged and debug binaries were removed from the test container. - UI and CLI tests passed. Initial environment failures (missing jq, Docker overlay file identity, and parallel linker memory pressure) and focused passing reruns are retained. The macOS full runner suite has platform-specific failures; Linux is the qualified full-check platform. - Eval harness: 27 tests passed; existing CI discovery ran 86 tests with two unrelated skips. Credential export rejection is tested against the actual report command. - Luna and OpenRouter Sonnet each passed 60 common-workflow runs: ten workflows, three repetitions per arm, zero unnecessary API fallback. - Sonnet passed 11 selected capability/contract cases after fixes. Gemini passed three smoke cases. DeepSeek exceeded the 120-second limit and remains unqualified. - Luna's two cost flags received focused follow-up. The original flags and a later n=1 latency flag remain visible. Sonnet had no cost or latency increase above 20%. - The catalog contains 785 entries; 58 were exercised across all stages. Most operation probes remain unrun and some need additional fixtures. Authored probes do not establish successful coverage. - Total conservative accounted cost: $9.875960. Active paid-campaign time: 88.16/90 minutes. No missing accounting. Later security and harness fixes have provider-free verification; no paid validation is claimed for those revisions. - Inspect the [qualification report](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/READINESS.md) and [verification record](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/verification.json). ## Risks - This is a broad authenticated API surface. Keep the default-off gate until an operator selects initial rollout companies. - Paid coverage is incomplete. Small regression samples do not prove all workflows are unchanged. - A timeout or server failure can follow a committed mutation. The result reports an unknown outcome and requires state inspection. - The new definitions add prompt tokens. The report retains cost flags and cache variation. - No database migration is required. - Repository rules require code-owner approval before merge. Technical CI and automated review are complete. ## Model Used OpenAI Codex based on GPT-6 assisted with code, tests and review. The exact serving model ID and context window are not exposed in this session. It used reasoning, tool calls and code execution. Eval models: `gpt-5.6-luna` with low reasoning, `openrouter/anthropic/claude-sonnet-5`, `openrouter/google/gemini-3.8-flash`, and `openrouter/deepseek/deepseek-v4-flash-0731`. Attempts retain runtime versions, model identity, usage and source provenance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f6a211479f |
fix: share current CLI runtimes across sandbox adapters (#12994)
## Thinking Path - Paperclip Runner needs its runtime preinstalled for fast sandbox startup. - Native and local adapters should launch one current CLI installation per provider. - An older global copy can shadow that installation, and exact native compatibility pins must match it. - Update the qualified releases and binary digests, expose shared CLI entrypoints from the provider pack, and prefer the image-owned bin directory. - Keep dependency installation in the image build; task startup only discovers, links, and verifies artifacts. ## Linked Issues or Issue Description **What happened?** Remote native startup rejected a stale global Codex, while CLI-only images lacked runnerd entirely. **Expected behavior** An image-baked runtime starts without uploading binaries or installing packages. All adapters share the same current provider CLI. **Steps to reproduce** Start a native remote task with the old global Codex and the updated runtime available only under `/opt/paperclip-runner/bin`. **Paperclip version or commit** Discovery behavior at `54a99d884`. **Deployment mode** Docker with a remote sandbox. ## What Changed - Prefer `/opt/paperclip-runner/bin`, then the user's local bin directory, then PATH. Existing metadata and version validation remains in force. - Qualify Codex 0.153.4, OpenCode 1.18.29, and Claude SDK 0.3.263 / CLI 2.1.263. Update binary digests, TypeScript/Rust checks, registry defaults, and the displayed OpenCode version together. - Share Codex and Claude's native executable with the ACP bridges through exact dependency overrides. Preserve the separately qualified ACP bridge implementations and their security patches. - Expose shared provider-pack CLI launchers; fail the pack build if Codex ACP resolves a separate Codex installation. Update the eval image's other agent CLIs to current stable releases and remove duplicate global provider installs. - Document the single-current-CLI policy in source comments and development guidance. Latest stable releases are resolved at review/build preparation and pinned; task startup never auto-updates. ## Verification - Native-session and adapter-registry suites: 158 tests passed. - Provider suites: 88 tests passed, 7 Linux-only checks skipped on macOS. One existing macOS temporary-path alias assertion passed when rerun with canonical `TMPDIR=/private/tmp`. - Package-contract and OpenCode materialization tests: 11 passed. - Full typecheck, build, and token gates passed. Rust native-provider/recovery tests: 19 passed. - Broad local suite: 5,974 passed, 23 failed, 41 skipped. Failures are in unchanged macOS workspace/path/port and connection suites; focused runtime tests pass. All latest-head Linux PR checks passed, including the full test shards, typecheck, build, runner verification, browser suites, and canary dry run. - The standalone fleet image built with one current provider CLI each and passed native Codex/Claude binary-integrity checks. A disposable Daytona sandbox reported ready in 798 ms; its baked runner completed an API-key `gpt-5.6-luna` turn in 2,430 ms and returned the expected marker with a usage receipt. No runtime artifacts were uploaded or installed. - The normal shared `codex exec` entrypoint also completed an API-key `gpt-5.6-luna` turn in 2,321 ms. - Both image builds verify the complete generated lockfile against a reviewed SHA-256 before package installation or lifecycle execution. Root lockfile changes remain CI-owned. Merge and rollout remain on hold for operator review. ## Risks - Updating provider CLIs changes their behavior for all adapters; version probes and live native smoke testing are required before image promotion. - The image-owned directory takes precedence. Its entries must launch the same shared CLI as the global PATH, not a private older/newer copy. - Application qualification pins and the deployed image must move together. No startup fallback installation is added. - No schema or authentication-policy changes. ## Model Used OpenAI GPT-6 (Codex). The session does not expose a more specific model ID or context-window size. Used reasoning, repository inspection, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
932ddb7b37 |
feat: browse GitHub repository access across organizations (#12998)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub connections give agents access to approved repositories. > - One GitHub identity can use installations across several organizations. > - The permissions page linked to one installation and showed an unfiltered list. > - Users could not easily find another organization or inspect a large selection. > - This change adds account filtering, search, and access configuration links. > - Users can inspect repository access in one compact view. ## Linked Issues or Issue Description Refs #12993. Related repository-catalog work in #11228 and #11234 was checked. This change only improves the existing GitHub connection permissions page. **What existing behavior does this improve?** The GitHub connection permissions page and its repository display metadata. **Current behavior** The page links directly to an existing installation. The repository list has no account filter, search, height limit, or private-repository marker. Refresh access occupies a separate section. **Proposed behavior** Show all authorized repositories by default. Filter by account or organization and search by name. Open GitHub's account chooser to configure access across organizations. Show GitHub icons and private-repository locks. Keep refresh beside configuration and limit the visible list to about ten rows. **Reason and benefit** Users can find repositories across organizations and configure missing access without creating another GitHub identity. Large repository lists no longer fill the page. **Breaking changes** None. Repository display metadata gains an optional private flag. Older snapshots remain valid and gain the flag after access refresh. No SQL migration is required. ## What Changed - Add an All accounts view, account filter, search, and empty states. - Link both configuration controls to GitHub's app account chooser. - Place an accessible refresh icon beside the configuration button. - Keep the repository heading and list in one section. - Add GitHub icons and private-repository locks. - Cap the scrollable list at ten rows using a design token. - Persist GitHub's private flag only when the provider returns a boolean. - Recover missing legacy app configuration from GitHub installation metadata. - Update tests and the GitHub connection runbook. ## Verification - Focused tests passed: 54 permissions-page tests and four GitHub metadata tests. - UI and server typechecks passed before submission. Token gates passed. - Browser checks verified account filtering, search, empty results, and the configuration destination. - The live list contained 40 repositories. Its final height was 272 pixels, which fits ten single-line rows with gaps. Scrolling retained all rows. - A live access refresh populated 30 private-repository lock icons from GitHub metadata. - Full workspace typecheck and build passed. The broad local suite stopped in the general-server group with 18 failed files. Failures include macOS temporary-path handling and embedded PostgreSQL startup. That run also overlapped the legacy fix and retained a stale GitHub module; the final focused run passed all 58 tests. Clean-runner CI is tracked separately. - Latest-head review is 5/5 with the legacy chooser finding resolved. All CI checks passed on commit `0ff2b63f348f5c87d8b7df6e43388f60f5d872d9`, including build, typecheck, all test shards, browser tests, and canary dry run. ## Risks - Older repository snapshots lack visibility metadata until refreshed. Unknown visibility does not display a lock. - The account filter lists authorized installation owners. Users add other organizations through GitHub's chooser. - Filtering changes only the displayed list. GitHub remains authoritative for repository access. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) via Codex. Reasoning, code execution, and browser tools were used. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bac60d9d31 |
fix: preserve GitHub sign-in and show connected repository access (#12993)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub connections give agents an account with selected repository access. > - Fresh local instances enroll with production Paperclip Cloud. > - Enrollment could finish while the GitHub OAuth profile remained disabled. > - Setup then switched to a personal access token form without explanation. > - This change preserves sign-in intent and shows the connected account and repositories. ## Linked Issues or Issue Description Related: #12907, #12943, #12947. Existing open GitHub connection work was checked. No duplicate was found. **What happened?** After Cloud enrollment, a fresh test-drive asked for a GitHub key. Production did not advertise the managed GitHub profile. Staging did. The permissions page also omitted the authenticated username and repository names. **Expected behavior** Continue with GitHub OAuth when available. Explain unavailable sign-in and allow retry otherwise. Show the GitHub username and complete accessible repository list. **Steps to reproduce** Start a fresh test-drive. Choose GitHub and complete instance enrollment while the Cloud GitHub profile is disabled. Open an existing GitHub connection's permissions page. ## What Changed - Preserve managed sign-in intent when the gallery omits its profile. - Refresh the selected gallery entry on retry without resetting the audience. - Fetch all pages of GitHub installations and repositories. - Store only repository IDs, full names, and installation IDs in grant metadata. - Show the GitHub username, repository list, management link, and refresh action. - Discard the repository snapshot after newer installation lifecycle events. Preserve snapshots verified after delayed events. - Lock and re-read grant metadata when applying installation events or saving refreshed access. Patch only webhook fields for other events. Reject snapshots if access changed during the external fetch, using unique access revisions even when timestamps collide. - Show repository installation recovery for managed OAuth even when the app also offers an advanced PAT method. - Update tests and the GitHub connection runbook. No SQL migration is required. ## Verification - Local typecheck, build, and token gates passed. All latest-head CI gates passed, including the complete test matrix and browser suites. Greptile is 5/5 with no unresolved findings. - All 382 focused setup, permissions, metadata, service, and webhook tests passed across final runs. One socket-hang-up test passed on rerun with the full service suite. Final service, metadata, and webhook checks passed all 230 tests. - The broad local suite was stopped after failures. Seven workspace-runtime exposure and control-conflict failures reproduce on base commit `54a99d884`. The broad run also overlapped local iteration; final focused tests and clean-checkout CI are tracked separately. - Browser: a fresh production-backed instance completed enrollment, retried after profile enablement, reached GitHub consent, recovered from a missing installation, and completed OAuth. - Browser: the permissions page showed the authenticated username and the selected private test repository. A real `get_me` call returned the same account. Reading the selected repository passed; reading an unselected private repository failed with 404. - Browser: a second fresh instance completed enrollment and OAuth without a PAT form or unavailable state. Its username and repository list survived reload and refresh. A real get_me call on the final code returned the displayed account. ## Risks - Repository names are now stored in company-scoped grant metadata and shown with that credential. They are display data, not authorization data. - Large selections require more GitHub API calls. A failed later page rejects the refresh rather than reporting a partial list. - Older grants and webhook-invalidated snapshots require Refresh access to load the list. - Cloud profile enablement is separate deployment configuration. This PR does not change OAuth scopes or GitHub App permissions. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) via Codex. Reasoning, code execution, and browser tools were used. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fee8d8dc39 |
fix(runner): repair direct live provider bootstrap (#12932)
## Thinking Path > - Paperclip runs AI agents through qualified provider backends. > - The direct live eval workflow builds one immutable Runner runtime for every matrix cell. > - The workflow reinstalled the packed Runner with npm. > - That install discarded pnpm patches and selected provider dependencies outside the qualified lock. > - The first pnpm deployment model also placed its virtual-store marker at the wrong level; a real deployment keeps `.pnpm` beside the scoped Runner package. > - AgentCore enforced the current context-aware harness but the direct eval CLI did not supply the production v3 runtime context that harness requires. > - This pull request preserves the qualified dependency graph, resolves the real deployment layout, and makes direct evals exercise the production runtime-context contract. > - The benefit is that live eval cells reach their provider turn with the same artifacts and context contract that Paperclip qualified. ## Linked Issues or Issue Description Refs: #12931 **What happened?** The full direct live eval campaign failed every ACPX cell during `session.open`. The portable runtime had an incorrect dependency root. Its npm install also discarded the qualified ACP server patches. AgentCore cells first failed because Runner enforced `aws-agentcore-harness-v1` while the provisioned stack and eval profile use `aws-agentcore-harness-context-v2`; after aligning that revision, the direct eval CLI still omitted the required v3 runtime context. **Expected behavior** The direct eval runtime must preserve the frozen pnpm dependency graph and patched provider bytes. Runner, server validation, OpenAPI, and the deployed AgentCore stack must use one qualification revision. Direct eval attempts must supply the same immutable native runtime-context contract as production. **Steps to reproduce** 1. Dispatch `Runner Direct Live Protocol Evals` from `master`. 2. Select an ACPX Claude, ACPX Codex, or AgentCore roster. 3. Observe a pre-turn provider bootstrap failure. **Paperclip version or commit** `d96452db059338b329b458ba8fe359fef72f1363` **Deployment mode** GitHub Actions on the RunsOn Linux x64 fleet. ## What Changed - Build the reusable direct-eval runtime with `pnpm deploy --prod`. - Resolve ACPX dependencies from the actual scoped-package layout of a self-contained pnpm deployment. - Align AgentCore configuration and qualification checks on `aws-agentcore-harness-context-v2`. - Materialize a minimal immutable v3 runtime context for each isolated direct eval attempt. - Add workflow, package-authority, runtime-context, Rust, and server regression coverage. - Document the qualified packaging, runtime-context, and AgentCore revision contracts. ## Verification - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/runnerd-codex-transport.test.ts` (70 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/cli/eval-session-contract.test.ts` (14 tests) - Focused Runner contract tests (36 tests) - Focused server profile tests (47 tests) - Focused Rust managed-provider and native-selector tests (19 tests) - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs` - `actionlint .github/workflows/runner-protocol-live-evals.yml` - A local `pnpm deploy --prod` produced both qualified ACP server digests. - A Linux reproduction of the first follow-up smoke identified the real deployment root and the missing AgentCore runtime context. ## Risks The AgentCore revision change rejects profiles that still use the obsolete v1 value. This is intentional because the provisioned context-aware harness and current eval profile use v2. Direct eval prompts now receive the same fixed runtime-context preamble as production, so behavior scores may move; that is the intended qualification surface. The workflow package layout changes, but tests assert the new entrypoint and dependency root. This change does not modify the browser full-stack E2E workflow. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The context-window size is not exposed in this session. The model used extended reasoning, repository tools, code execution, Docker-based Linux reproduction, and GitHub Actions diagnostics. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (for example, `docs/...` or `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0c1e7504c0 |
fix(runner): persist warm Daytona workspaces across turns (#12904)
## Thinking Path > - Daytona preserves a stopped sandbox filesystem, but deleting or replacing a sandbox removes its only remote copy. > - Warm reuse therefore improves latency but cannot be Paperclip's durability boundary. > - The host execution workspace must remain authoritative after every successful turn, while same-run recovery must avoid overwriting unexported remote work. > - Result proposal, workspace export/merge, and terminal completion need a durable, replayable ordering so a crash never starts a duplicate provider turn. > - A paid browser acceptance suite must exercise both legacy Codex and Runner Codex for three real turns on one continuously warm Daytona sandbox. ## Linked Issues or Issue Description Refs #12901. Runner Codex did not previously export successful Daytona workspace changes back to the authoritative host workspace. That made warm reuse depend on Daytona's remote filesystem and left deleted/replacement sandboxes without a reliable reconstruction path. The existing paid fixture also lacked a focused three-turn continuity case for both Codex adapters. ## What Changed - Persist versioned, atomic native workspace-sync descriptors and durable seeds in `PAPERCLIP_HOME`, without credentials or a database migration. - Classify fresh, warm, replacement, and same-run-recovery workspace preparation explicitly; ambiguous lease/root/digest evidence fails closed. - Finalize native workspace export/merge after semantic result proposal and before run completion, with idempotent replay that never submits a second provider turn. - Surface legacy Codex workspace restoration failures instead of masking them, while preserving an earlier provider error when both fail. - Keep healthy reusable Daytona leases warm for legacy and native adapters, stamp finalized workspace generations, and retain existing cleanup behavior for per-turn or unhealthy leases. - Preserve Runner Codex's provider process/session across warm turns, including bounded post-terminal tail draining and exact authority rotation. - Add the exact paid `daytona-warm-continuity` matrix: - `legacy-codex × daytona × warm-three-turn` - `runner-codex × daytona × warm-three-turn` - Drive all three turns through the browser, verify ordered file continuity and stable lease/workspace/runtime identities, capture per-turn timings, and delete the sandbox immediately after assertions. - Document `pnpm test:e2e:runner -- --suite daytona-warm-continuity`; no package script was added. ## Verification - `pnpm typecheck` — passed, including migration safety (no migration added) - Focused server/runner Vitest coverage — 144 passed - `pnpm test:e2e:runner:unit` — 114 passed - `pnpm test:e2e:runner:typecheck` — passed - `pnpm --filter @paperclipai/paperclip-runner test:codex` — 66 passed, 1 helper ignored - `native-session-executor.test.ts` — 139 passed, including safe fail-closed cleanup after remote runner identity capture failure - Paid local browser acceptance, exact post-rebase Linux/amd64 runner binary: - Runner Codex — passed in 1.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - Legacy Codex — passed in 2.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - [Protected paid GitHub Actions campaign](https://github.com/paperclipai/paperclip/actions/runs/34026735033) against `7da42a91b95fa7fb2df126668ef7e37afb3b2b9d` — passed 2/2: - Runner Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Legacy Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Merge/enforcement, S3 history, and Pages publication jobs passed - Paid result artifacts were scanned for both provider credentials; neither secret was present. - Current PR checks — 31 passed, 1 expected Storybook skip; Greptile 5/5; Superagent security scan passed - `git diff --check origin/master...HEAD` — passed - Confirmed no `package.json`, lockfile, migration, or SQL changes. ## Risks - Workspace synchronization now sits on the terminal-success path, so a remote export failure deliberately prevents false success. Retryable state retains its lease/seed; loss of the only unexported remote copy fails closed. - Warm provider reuse has strict identity and quiescence checks. Mismatched or ambiguous evidence blocks reuse rather than risking concurrent provider work. - The paid suite incurs Daytona and Codex cost only in the existing protected scheduled/manual workflow and explicitly destroys its sandbox after each cell. ## Model Used OpenAI Codex with GPT-5 agentic reasoning, repository inspection, real browser E2E execution, Rust/TypeScript test execution, and GitHub Actions diagnostics. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing issue or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the dedicated suite invocation without adding a package script - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green on the current revision - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the current revision - [x] I will address all reviewer comments before requesting merge |
||
|
|
3796c6f259 |
fix(connections): project GitHub identity into sandbox runners (#12907)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed GitHub connections resolve a responsible user's or dedicated agent's identity into an audited, run-scoped credential projection. > - An enrolled instance could retain a hidden managed setup method after Cloud stopped advertising it, producing a blank, disabled setup step. > - Native runner processes also dropped the resolved GitHub projection before the provider shell, so `gh` and Git could not use the selected identity in Daytona. > - Daytona already provides the outer isolation boundary. Applying Codex's inner Linux sandbox there both duplicated containment and failed because nested user namespaces are unavailable. > - This change repairs setup fallback, carries only the bounded GitHub projection across each runner boundary, and allows only a controller-selected managed sandbox transport to act as the outer sandbox. ## Linked Issues or Issue Description **What happened?** An enrolled self-hosted instance could show a blank GitHub setup step when its managed profile was unavailable. Separately, a native Codex run in Daytona could resolve a managed GitHub connection on the Paperclip host but lose it before the provider shell. Once projected, Codex's nested sandbox failed before commands could run because Daytona does not expose the user-namespace operation used by the inner sandbox. **Expected behavior** Setup must select an advertised customer method when the managed method is unavailable. A Daytona run must receive the exact managed GitHub identity selected for that run, support `gh` and HTTPS Git, and rely on Daytona as its outer sandbox without weakening local or SSH execution. **Steps to reproduce** 1. Enroll a self-hosted instance while Cloud does not advertise the managed GitHub profile and open GitHub setup. 2. Observe the blank second step and disabled action. 3. Configure a native Codex agent with a Daytona environment and a responsible-user GitHub grant. 4. Run `gh api user` or HTTPS Git from the agent shell. 5. Observe missing GitHub environment projection or nested-sandbox startup failure. **Paperclip version or commit** The setup bug reproduces on `1dceee9a4`; the runner proof was developed from the same branch and verified at the latest head below. **Deployment mode** Self-hosted Paperclip enrolled with Paperclip Cloud, using the Daytona sandbox-provider plugin and native Paperclip runner. ## What Changed - Wait for connector enrollment hydration, retain a hidden managed method only while enrollment is needed, and otherwise select an advertised customer fallback. - Add a single bounded GitHub credential-environment projection for `GH_TOKEN`, `GITHUB_TOKEN`, the process-only Git helper token, GitHub commit identity, and at most 32 controller-generated Git config entries. - Forward that projection through the durable controller, Codex app-server transport, and Rust provider child without placing token values in arguments or config. - Allow Codex shell inheritance only for the exact projected GitHub keys and enable provider network access only when the managed credential exists. - Derive outer-sandbox authority exclusively from a managed `sandbox` transport; strip the same flag from configured, host, local, and SSH environments. - Define a named external-sandbox permission profile that Codex resolves to `dangerFullAccess` for default-mode Daytona turns while plan mode remains read-only. - Add regression tests for setup fallback, credential projection, local/SSH/sandbox authority separation, provider forwarding, and permission-profile selection. ## Verification - `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx` — 96 passed. - `pnpm exec vitest run src/drivers/codex/codex-security-config.test.ts src/drivers/codex/app-server-transport.test.ts src/control-plane/durable-prp-control-plane.test.ts` from `packages/paperclip-runner` — 31 passed. - Focused native-session executor tests — 3 passed. - `pnpm --filter @paperclipai/paperclip-runner typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --lib` — 194 passed. - Real Codex app-server configuration probe accepted `paperclip-runner-external-sandbox` and reported `sandbox.type=\"dangerFullAccess\"` while using the named profile. - [Signed Daytona image workflow](https://github.com/paperclipai/paperclip/actions/runs/33995270328) built commit `86571f7997e7100e47bd131aac1f1e773112a0ce`; the isolated environment was pinned to `sha256:ecef21105f8de382d75787e59439d936be239b77ae74a31c8ed3a17cde39b023`. - Live isolated Daytona proof passed: the three projected token variables were non-empty and equal; the host-scoped Git credential helper returned the same token without printing it; `gh api user` resolved `cryppadotta`; authenticated `git ls-remote https://github.com/paperclipai/paperclip.git HEAD` returned `1dceee9a4e75b13456760bb54c752deb2dba1d79`; no repository mutation occurred. - The persisted 28,476-byte run log contains no GitHub token shape, bearer header, credential-bearing URL, or private-key marker. - Latest-head pull-request CI and reviews provide the remaining full-suite gate. ## Risks - This deliberately gives shell Git and `gh` access to the run's resolved GitHub identity. It is the audited class-3 behavior required by the GitHub connection design and is outside per-tool Ask-first controls. - The credential source is the trusted broker projection, which overwrites configured environment values. The helper is scoped to HTTPS `github.com`, revalidates protocol and host, and never places its token in command arguments, URLs, or files. - Managed Daytona sandboxes become the containment boundary for default-mode provider commands. Local and SSH targets retain the inner Codex workspace sandbox, and plan mode remains read-only everywhere. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, with reasoning, browser control, shell access, and code execution. The product did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
60469a08e0 |
feat(agent-login): resume an active login session and permit concurrent login terminals (#12861)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent authentication uses server sessions, plugin workers, and browser login panels. > - A page reload loses an active login session, and one worker permits only one login terminal. > - These limits cause lost work and prevent two owners from logging in through one worker. > - This pull request lets the browser resume active sessions and lets workers serve concurrent login terminals. > - The benefit is reliable login recovery with a bounded process-wide route limit. ## Linked Issues or Issue Description **What existing behavior does this improve?** It improves agent credential login recovery and concurrent login terminal handling. **Subsystem affected** Cross-cutting (multiple of the above). **Current behavior** A page reload loses the active login session. A shared plugin worker rejects a second login terminal. **Proposed behavior** The browser reads and resumes the owner's active session. A worker supports multiple login terminal routes under a process-wide ceiling. **Reason and benefit** Owners keep login progress after a reload. Two owners can log in through one worker without removing the route limit. **Breaking changes** None. The change adds owner-scoped read routes and changes login terminal concurrency. ## What Changed - Replace the single worker login route with maps keyed by host route and worker session identifiers. - Add a process-wide login route ceiling and release each reserved slot on every exit path. - Add owner-scoped active-session reads with consistent negative responses and private cache control. - Keep the device-login prompt while the session has an active public status. - Add a durable setup-token cancel fallback for a lost in-memory session. - Resume active sessions when the agent configuration or onboarding panel mounts. - Remove routine unmount cancellation and keep explicit Cancel behavior. ## Verification - `pnpm --filter @paperclip/server test` — server route, service, and plugin-worker-manager suites. - `pnpm --filter @paperclip/plugin-sdk test` — worker RPC host suite. - `cd ui && npx vitest run src/components/AgentConfigForm.render.test.tsx src/components/OnboardingWizard.test.tsx`. - `cd ui && npx tsc -b`. - `tests/e2e/onboarding.spec.ts` — reload during login. - CI must pass on this pull request. ## Risks The change affects agent authentication and the sandbox-to-host boundary. Route cleanup must release every reserved slot. Owner checks must prevent cross-owner session access. Tests cover route cleanup, owner scope, reload recovery, and concurrent worker routes. ## Model Used Codex, OpenAI GPT-5, tool use and code review support. The implementation author owns the exact model details for the code changes. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR with the relevant issue-template fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ed5b7c5c8 |
refactor(server): extract the active-run output watchdog into a feature module (#12853)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server recovery service monitors active runs and applies watchdog decisions > - The watchdog rules and database operations lived in one large recovery service > - This structure made the rules harder to test and made company scoping harder to inspect > - This pull request moves the watchdog into domain, application, and adapter layers > - The benefit is a smaller recovery service, pure policy tests, and clear company-scoped ports ## Linked Issues or Issue Description **What existing behavior does this improve?** The active-run output watchdog that detects silence, suppression, and terminal evidence. **Current behavior** The recovery service contains the watchdog policy, use cases, database operations, and process control in one file. **Proposed behavior** A feature module separates pure policy, use cases and ports, and Postgres and process adapters. The recovery service delegates its public watchdog methods to this module. **Reason and benefit** The separation makes policy decisions easy to test. Company identifiers on every reader and writer port make tenant scope clear. Smaller service methods reduce change risk. **Breaking changes** None. The recovery service keeps its public methods and call sites. Related public watchdog work includes [#7043](https://github.com/paperclipai/paperclip/pull/7043) and [#7770](https://github.com/paperclipai/paperclip/pull/7770). ## What Changed - Add the `server/src/modules/active-run-watchdog/` feature module with domain, application, and adapter layers. - Move watchdog policy, use cases, Postgres access, and local process control into the module. - Keep the recovery service public methods and delegate them to the module. - Add 53 pure module test cases and retain 8 Postgres integration cases. - Add company scoping and transaction rollback coverage. ## Verification - Run `vitest run --config vitest.config.ts src/modules` and confirm 3 files and 53 cases pass. - Run `vitest run --config vitest.config.ts src/__tests__/heartbeat-active-run-output-watchdog.test.ts` and confirm 1 file and 8 cases pass. - Run the full server suite in pull request CI. - Compare the type-check result with a fresh baseline on the same checkout. ## Risks The main risk is a behavior change in recovery decisions during the move across layers. The pure policy tests cover the moved rules. The integration tests cover database behavior, company scope, and transaction rollback. Pull request CI runs the full server suite. ## Model Used OpenAI Codex, GPT-5, runtime-managed context window, tool use, code execution, and repository review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d2d647c34b |
fix(connections): honor identity after reconnect (#12897)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - Connections give agents access to external services with a selected credential identity. > - A removed connection keeps its database row so Paperclip can retain its history. > - A fresh GitHub setup can select a different identity from the removed connection. > - The retained row incorrectly kept its old credential policy after that new selection. > - The GitHub callback then could not save the new grant and returned the user to setup. > - This pull request applies the explicit identity selection when Paperclip revives an archived row. > - The benefit is a successful GitHub reconnect after the user changes from a dedicated agent account to a personal account. ## Linked Issues or Issue Description **What happened?** After a user removed a dedicated-agent GitHub connection, a fresh setup with “My GitHub account” returned to the setup page with `oauth=failed`. The Cloud claim succeeded, but the local connection still used the old `per_agent` policy. **Expected behavior** A fresh setup must apply the explicit identity choice. An interrupted draft or an explicit reconnect must keep its existing identity. **Steps to reproduce** 1. Connect GitHub with a dedicated agent identity. 2. Remove the connection. 3. Start a fresh GitHub connection with “My GitHub account.” 4. Complete GitHub OAuth. 5. Observe that Paperclip returns to the setup page instead of the permissions page. **Paperclip version or commit** `342c01fee` **Deployment mode** Local dev (`pnpm dev`) with embedded Postgres and the staging managed connector. **Additional context** This follows the GitHub access UI change in #12893. ## What Changed - Apply an explicit Access identity when a fresh gallery setup revives an archived connection row. - Preserve the identity for interrupted drafts and explicit reconnects. - Do not carry credential material across an identity change. - Apply the omitted organization default during a fresh archived-row recovery. - Restore the prior grants and credential policy transactionally if a revived setup rolls back. - Disable the connection and surface a specific failure if that restoration cannot complete. - Preserve newer concurrent grant changes with a row lock and optimistic version check. - Preserve newer concurrent connection identity/configuration changes with a locked state fingerprint. - Add regressions for dedicated-to-personal OAuth, organization-default recovery, rollback, rollback failure, and concurrent grant/connection changes. ## Verification - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts` — 222 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - Browser proof on an isolated local instance: dedicated connection removed, personal setup selected, staging GitHub OAuth completed, permissions page opened, connection reported active and healthy, personal grant active, old agent grant revoked. ## Risks - Low migration risk. This change has no schema migration. - The behavior changes only when a fresh setup explicitly selects an identity for an archived connection row. - Existing draft resume and explicit reconnect behavior stays unchanged. - Connection-manager checks still protect changes to a retained credential identity. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, with reasoning, browser control, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
70c9ca7410 |
fix(server): stop mock leakage between interaction-route tests (#12807)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server test suite checks issue-thread interaction routes > - Shared Vitest mocks can keep queued one-shot values between tests > - A leftover value can change the issue returned to a route and cause a false authorization failure > - This pull request resets all mocks and restores the plain run-attribution value before each test > - The benefit is stable interaction-route tests that do not depend on test order ## Linked Issues or Issue Description This pull request has no public issue link. The bug details follow. **What happened?** The interaction-route server test suite failed intermittently in continuous integration. The test named `lets a watchdog-scoped assignee withdraw through ordinary containment` sometimes failed because `withdrawInteraction` received no call. The test suite used `vi.clearAllMocks()`, which clears call history but does not clear queued one-shot mock values. A queued value from `mockIssueService.getById` could change a later test's issue and make the route return `403`. **Expected behavior** Each test must start with empty mock queues and the default run-attribution value. Test results must not depend on test order. **Steps to reproduce** 1. Run the interaction-route test file many times in sequence. 2. Run the same file with shuffled test seeds. 3. Observe the intermittent containment failure before this change. **Paperclip version or commit** Current `master` plus commit `8cd56b38a77f1feecac495f57a48d3f0a1b3b01c`. **Deployment mode** Built from source. The failure occurs in the server test suite. ## What Changed - Replace the partial mock reset with `vi.resetAllMocks()`. - Reset `mockRunAttribution.value` before each test. - Keep the change within the interaction-route test file. ## Verification - The target suite passes 77 of 77 runs. - The test count remains 64 `it(...)` sites, including five `it.each(...)` blocks that expand to 77 runs. - The file contains no skipped or focused tests. - The diff changes no production code. - A defect-detection round trip reproduces the failure when the containment guard is broken and passes after the guard is restored. - Ten sequential runs and five shuffled-seed runs pass 77 of 77. ## Risks Low risk. The change affects test setup only. It does not change production code or route behavior. ## Model Used OpenAI GPT-5, exact runtime model `gpt-5`, tool use and code execution, context window not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
45725ad820 |
fix(server): hoist the comment-cancel route test suite's module graph (#12877)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server test suite checks issue comment and cancellation routes.
> - The comment-cancel route suite reloaded about 40 modules before
every test.
> - Repeated module reloads created a race between service mocks and
real services.
> - The race caused an HTTP 500 when the test expected HTTP 200.
> - This pull request loads the mocked module graph once for the suite.
> - The benefit is stable route tests with clearer diagnostics for
future failures.
## Linked Issues or Issue Description
**What happened?**
The comment-cancel route test suite failed intermittently in continuous
integration with an HTTP 500 where the test expected HTTP 200. The suite
reset modules and re-imported the route graph before every test. A
re-import could bind the real service module to the test's minimal fake
database and cause a `TypeError`.
**Expected behavior**
The suite must run all seven route tests without intermittent HTTP 500
responses. A future server error must show its underlying cause in the
test output.
**Steps to reproduce**
1. Run `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/issue-comment-cancel-routes.test.ts`.
2. Repeat the test command under continuous integration load.
3. Compare the result with a run that reloads the route module graph
before every test.
**Paperclip version or commit**
`0a6a7087ed6c3bb1cadf59fcfec362d6fc9a6d14`
**Deployment mode**
Built from source with the server test runner.
**Installation method**
Built from source.
**Agent adapter(s) involved**
Not adapter-specific (core test issue).
**Database mode**
Not database-related.
**Access context**
Unclear / not applicable.
**Additional context**
The route module and error-handler middleware remain real. The service
layer remains mocked. Authorization and non-leakage assertions remain
unchanged.
## What Changed
- Register mocks once and load the route module graph once through
`hoistModuleGraph`.
- Remove the per-test module reset and re-import.
- Add a `res.on("finish")` diagnostic listener for server error context.
- Keep all seven test titles and the existing authorization assertions.
## Verification
- Run `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/issue-comment-cancel-routes.test.ts`.
- Confirm that all seven tests pass.
- Confirm that the full continuous integration suite passes on this pull
request.
## Risks
Low risk. This change modifies one test file and does not change
production code or test coverage. The local worktree cannot start this
suite because it lacks `packages/adapters/droid-local`; continuous
integration must verify the complete repository dependency set.
## Model Used
OpenAI Codex, GPT-5. The model used repository tools, GitHub tools, and
code review reasoning. The execution context window is not exposed by
the runtime.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
58ed2b64ea |
test(server): remove a concurrent-import race in the approval routes suite (#12876)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Approval routes use server tests to protect access and idempotency behavior > - The approval routes test suite loaded mocked modules at the same time > - Concurrent module loading could lose a service mock and produce a false test failure > - This pull request loads the shared module graph once and reuses it across the suite > - The benefit is stable approval route tests with unchanged coverage ## Linked Issues or Issue Description **What happened?** The approval routes test suite loaded two mocked modules in one concurrent import. A module interleaving could remove the approval service mock. The route then returned HTTP 404 instead of the expected HTTP 403. **Expected behavior** The suite must keep the approval service mock when it loads the route modules. The access test must return HTTP 403 on every run. **Steps to reproduce** 1. Run `npx vitest run server/src/__tests__/approval-routes-idempotency.test.ts`. 2. Repeat the test command while the test runner loads the module graph. 3. Observe a false HTTP 404 result when the module mock interleaves. **Paperclip version or commit** Current `master` at the base commit of this pull request. **Deployment mode** Built from source with the server test runner. ## What Changed - Reused the existing `hoistModuleGraph` helper for the approval route modules. - Loaded the route modules once in sequence instead of in one concurrent import. - Kept per-test mock behavior, Express app setup, database doubles, test names, and assertions unchanged. ## Verification - `npx vitest run server/src/__tests__/approval-routes-idempotency.test.ts` — 11 of 11 tests passed. - Ten repeat runs passed. - `npx tsc --noEmit -p server` produced no new errors against the base branch. - All Paperclip CI checks passed. - Greptile reported 5/5 with no open findings. ## Risks This change affects test module setup only. It does not change production code or test coverage. Risk is low. ## Model Used OpenAI Codex with the `gpt-5` model family. The serving snapshot and context-window size are not exposed. The agent used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and found no duplicate for this test race - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f0c1d4548 |
feat(cli): add isolated test-drive command (#12894)
Add a foreground-only test-drive workflow with isolated data, provider-backed CEO bootstrap, OpenCode/OpenRouter support, worktree execution setup, reuse safeguards, and delayed browser opening. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5da6499860 |
fix(connections): reuse one-time cloud enrollment (#12891)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed connections let agents use provider credentials without exposing those credentials to the control plane UI > - A self-hosted instance must first establish a trusted credential destination with Paperclip Cloud > - The GitHub connection flow repeated that trust decision before provider consent > - The local setup route also lost step 2 after enrollment and could display the PAT identity defaults before enrollment > - This pull request makes enrollment a one-time instance decision and sends later provider starts directly to provider consent > - The benefit is a shorter flow with one clear Paperclip approval and no required service restart ## Linked Issues or Issue Description Refs #12843. Companion Cloud change: [paperclipai/paperclip-cloud#391](https://github.com/paperclipai/paperclip-cloud/pull/391). ## What Changed - Made `stage=setup` authoritative during initial route hydration and enrollment return. - Added a contained one-time enrollment screen with provider-specific copy. - Accepted a provider `authorizationUrl` from Paperclip Cloud only when it matches the exact GitHub or Google OAuth endpoint. - Preferred the direct provider URL while retaining the legacy confirmation URL fallback. - Preserved the company-bound identity and agent-access draft across the full-page enrollment callback, including cold company-context hydration. - Kept GitHub defaulted to “My GitHub account” and “Any agent,” including before Cloud advertises the managed method. - Updated GitHub identity and agent-access copy for responsible-person and dedicated-agent behavior. - Labeled the provider action “Continue to GitHub.” - Added parser, routing, cold-hydration, access-restoration, visibility, fallback, defaults, and copy tests. ## Verification - `pnpm exec vitest run server/src/services/paperclip-cloud-connector.test.ts ui/src/pages/apps/AppsConnect.test.tsx` (114 tests passed) - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` - Live browser proof used a new data directory on `127.0.0.1:3117` and the exact Cloud PR revision on staging. - The fresh flow selected “My GitHub account” and “Any agent,” showed one enrollment approval, returned to local step 2, and connected GitHub without a second Paperclip confirmation or login. - The connected screen showed one selected repository, a long-lived token, installation metadata, a successful access refresh, and healthy webhook delivery. - Gmail on the same instance went directly to Google consent without another Paperclip approval. - Restarting the same data directory preserved enrollment. A second new data directory required exactly one new approval. - A final fresh-data-dir rerun selected a dedicated GitHub identity for Ada before enrollment, approved the instance once, returned to step 2, retained Ada after a Back check, connected directly through GitHub, and finished with “Used only by Ada,” one selected repository, and a long-lived token. - Port 3100 remained untouched throughout the proof. ## Risks - The new Cloud field is additive and restricted to the exact GitHub and Google OAuth origins and paths, with no embedded credentials or URL fragment. - An older Cloud response still works through `confirmationUrl`. - A self-hosted instance still requires one signed Cloud enrollment. Managed Cloud instances do not render the enrollment screen. - Provider authentication and consent remain mandatory after instance enrollment. - No schema migration is included in this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, browser control, and multi-file repository editing. The context window size was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bcc6fe7a44 |
fix(runner): restore multi-turn remote sessions (#12840)
## Thinking Path > - Paperclip manages AI agents and their work. > - The runner executes agent turns on local and remote providers. > - A remote per-turn session must save its state before Paperclip releases its sandbox. > - The session runtime returned after 100 milliseconds while the remote checkpoint still ran. > - The next turn also checked the local state path instead of the verified remote backup. > - This pull request waits for the bounded remote close and accepts only a verified suspended backup. > - The benefit is reliable multi-turn execution without weaker identity checks. ## Linked Issues or Issue Description **What happened?** A successful remote agent turn released its sandbox before the runner saved the verified continuation backup. The next turn failed with `runner_state_identity_mismatch`. **Expected behavior** Paperclip must finish the bounded remote checkpoint before it releases the sandbox. A later turn must validate and restore the digest-matched suspended backup. **Steps to reproduce** 1. Run a native ACPX Claude Plan test in a non-reusable Daytona sandbox. 2. Reject the first plan to start a second turn. 3. Observe that the second turn fails before provider execution. **Paperclip version or commit** The failure reproduced at `13775a90b078ff64872f50961ea1b83d575e7bc6`. **Deployment mode** GitHub Actions with a Daytona sandbox. ## What Changed - Wait for the internally bounded remote runner close and checkpoint before the host returns. - Preserve the existing short cleanup bound for other providers. - Validate remote continuation lifecycle from a complete digest-verified backup when local runner state is absent. - Keep corrupt, non-suspended, mismatched, and unverified state fail-closed. - Make native Plan completion and accepted-Plan wake prompts deterministic. ## Verification - A prior 45-cell local campaign passed 44 cells. The only failure was the OpenCode Plan prompt variance fixed here. - A focused OpenCode local Plan rerun passed. - ACPX Claude Daytona message and question cells passed. - Focused regressions cover delayed checkpoint close and verified remote backup lifecycle. - GitHub Build and the focused ACPX Claude Daytona Plan cell will validate this exact head. ## Risks Remote runnerd sessions now wait for their internally bounded close/checkpoint path before returning; generic provider cleanup retains the existing 100 millisecond bound. Durable run success still cannot be reversed. The environment release guard still blocks sandbox destruction when no verified backup stamp exists. ## Model Used OpenAI Codex, GPT-5.6, extended reasoning, with code execution and GitHub Actions inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal task id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open findings - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0ffc091473 |
feat(connections): add durable GitHub identities and webhooks (#12843)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents need source control access for repository work > - A shared token cannot preserve the responsible person's identity or an agent's dedicated identity > - GitHub App tokens also need durable refresh, repository access checks, and webhook delivery > - Paperclip already has managed connections, encrypted grants, run secret leases, and merge-confirmation behavior > - This pull request extends those systems with GitHub identities instead of adding a parallel credential system > - The benefit is durable GitHub access with explicit identity, repository, runtime, and webhook boundaries ## Linked Issues or Issue Description No public GitHub issue describes this connection change. This description follows the feature request template. **Subsystem affected** Connected Apps, connection grants, secret resolution, native Git runtime setup, webhook processing, and the Apps UI. **Problem or motivation** Users need to connect GitHub once and let agents use the correct GitHub identity. A run should use a dedicated agent account when one exists. Otherwise, it should use the responsible person's account. The connection must survive token expiry, repository access changes, and temporary instance downtime. **Proposed solution** Add user-owned and agent-owned GitHub grants to the existing connection model. Resolve one identity for MCP, Git, `gh`, health checks, and webhook bindings. Store provider tokens in the existing encrypted secret system. Refresh expiring token pairs under the existing lease and compare-and-swap path. Register signed Cloud webhook bindings and process normalized pull request and installation events through a durable local inbox. **Alternatives considered** An organization-wide GitHub token would lose person and agent attribution. Environment variables alone would bypass the managed connection and grant model. A new GitHub-only credential store would duplicate the existing secret and access systems. GitHub App installation tokens and private-key custody remain outside this first version. **Roadmap alignment** This change implements the Connected Apps direction. It also extends the shipped MCP Tool Gateway, per-agent secret access, and action-attribution systems. It does not add a repository catalog. The open repository catalog work in [#11234](https://github.com/paperclipai/paperclip/pull/11234) is related and complementary. ## What Changed - Added agent-owned connection grants and a per-agent credential policy with company and subject constraints. - Added a managed GitHub App method while keeping the personal access token method as an advanced fallback. - Added durable access-token and refresh-token handling with proactive rotation and one automatic recovery after a provider `401`. - Added GitHub identity and installation summaries without storing repository-name lists. - Added signed Cloud webhook binding, event lease, acknowledgement, local idempotency, pull request merge processing, and installation access handling. - Added one identity resolver for MCP, native Git, `gh`, checkout, health checks, and webhook bindings. - Added a class-3 run projection for `GH_TOKEN`, `GITHUB_TOKEN`, a `github.com`-only credential helper, SSH-to-HTTPS rewrite, and GitHub noreply commit attribution. - Added personal and dedicated-agent setup choices plus identity, repository, continuity, and webhook status in the Apps UI. - Added schema migrations, tests, and connection documentation. ## Verification - The current head is fully green in GitHub CI, including build, typecheck, all serialized/general server shards, all browser shards, policy, canary dry run, review, and security checks. - Live staging proof completed with a non-expiring GitHub App user token, selected-repository installation, repository add/remove refresh, managed MCP, native `gh`, HTTPS clone/push/delete, GitHub noreply commit attribution, signed merged-PR webhook acceptance, durable Cloud-to-instance delivery, and installation-access event processing. Temporary branches and temporary repository access were removed afterward. - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed before and after the rebase onto `origin/master`. - `pnpm build` passed. - The focused connector suite passed 285 tests after the rebase. - The full stable suite passed 5,790 tests and failed 22 tests across 8 general server files. The failures reproduced as shared-runner environment issues. They included `/tmp` versus `/private/tmp`, closed database connections, and invalid high ephemeral ports. The focused connection tests pass in isolation. ## Risks - Migrations add agent grant subjects and a durable connection-event inbox. Migration numbering and safety checks pass. - A raw GitHub user token enters the agent process for Git and `gh`. Per-tool Ask-first controls cannot limit those shell operations. The UI warns users about this boundary. - GitHub App user tokens can be non-expiring. Paperclip performs a continuity check every 30 days, but provider revocation still requires a reconnect. - The webhook path accepts only signed and bounded payloads. It stores a minimal normalized record and no raw provider payload. - GitHub repository permissions remain authoritative. Removed access can make a cached repository count temporarily stale, but runtime access fails immediately. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, browser control, and multi-file repository editing. The context window size was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
263f181fed |
fix(runner): complete live hot restart adoption (#12852)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner owns durable provider sessions and streams their work to the control plane. > - Pull request #12845 added native restart recovery for live and dead local runners. > - A real browser test found three live-adoption gaps after that pull request merged. > - Lazy runner process ownership was not always stored before restart. > - The old controller did not release its PRP authority without closing the provider turn. > - Reconnect events could arrive before the active provider turn was restored. > - This pull request closes those gaps and proves the same turn completes after a UI hot restart. ## Linked Issues or Issue Description Refs #12845 Related search results: #12646 covers indeterminate command results after a runner restart. It does not cover controller adoption or active-turn rebinding. No open duplicate pull request was found. ## What Changed - Store lazy runnerd process ownership after provider session creation, read, and resume. - Detach native PRP controller authority during coordinated hot shutdown. Keep the live provider turn running. - Restore the exact checkpointed provider session when bounded PRP identity events have been compacted. - Restore the active provider turn before reconnect events are replayed. This prevents `turn_binding_mismatch`. - Keep exact live ownership by the current controller out of generic orphan recovery. - Add driver, transport, and server regression tests for these paths. ## Verification - Ran 12 Codex driver lifecycle tests. - Ran 53 runnerd transport tests. - Ran 143 recovery and orphan-reaper server tests. - Ran all 8 real-process restart recovery scenarios. - Ran all 96 existing runner E2E unit tests. - Ran runner TypeScript typecheck. - Ran server TypeScript typecheck. - Ran the migration replay test and migration safety checks. - Tested the board UI on an isolated local instance. A real local Codex-backed turn entered a 120-second terminal wait. The UI `Restart now` action replaced the server and kept the same runner PID, process start time, run ID, native session ID, runner ID, provider session ID, and active turn. The original turn then completed. - Confirmed one heartbeat run, no retry row, one result, one proposed-result event, one terminal event, no protocol errors, no active recovery state, and no surviving runner or provider process. ## Risks - A live runner can continue provider work while no server owns the control route. Recovery fails closed when the process fingerprint or durable identity is ambiguous. - Provider identity can be restored from the database only for an exact verified adoption claim. An authenticated live `session.snapshot` validates that identity before the driver can resume. - The new detach path applies only to native sessions that expose restart detachment. Other adapters keep their existing shutdown behavior. - This follow-up does not change the database migration or `package.json`. The migration in #12845 remains replay-safe through `ADD COLUMN IF NOT EXISTS` and its embedded-Postgres idempotence test. The dedicated real-process command remains in `doc/DEVELOPING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact serving build and context-window size are not exposed. The run used extended reasoning, repository tools, shell execution, and in-app browser automation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
7b094724e6 |
fix(runner): recover native sessions across restarts (#12845)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner keeps durable run and provider state outside one server process. > - A server restart can leave that runner alive or can interrupt it after a provider checkpoint. > - The old startup path used handoff intent and PID evidence, but it did not reconstruct native ownership. > - That gap could block the issue, create a replacement run, or start duplicate provider work. > - This pull request adds durable same-run recovery for coordinated and uncoordinated restarts. > - The benefit is exact recovery of the run, runner, session, provider, steering, and finalization state. ## Linked Issues or Issue Description Refs #9628. That pull request added earlier local-adapter hot-restart work. This change adds native PRP authority reconstruction and same-run provider resume. Refs #10935. That pull request handles missing hot-restart snapshots. This change also supports hard restarts with no snapshot. Refs #11624. That pull request prevents unsafe retry after an adopted legacy process exits. This change reconciles native terminal evidence before provider recovery. Refs #12070. That pull request improves process liveness checks. This change also binds recovery to a process-start fingerprint and fails closed on ambiguity. **What happened?** The server could record hot-restart intent, but startup did not rebuild native runner ownership. A live runner could not re-register its PRP authority. A dead runner could not resume the exact native and provider session on the same heartbeat run. Generic recovery could then block the issue or create replacement work. **Expected behavior** A live native runner must reconnect with the same PID and logical identities. A dead runner must resume the same durable session and heartbeat run with only a new operating-system PID. A proposed or terminal result must finalize once before any provider turn starts. Ambiguous process or session evidence must stay blocked without a signal or duplicate spawn. **Steps to reproduce** 1. Start a Paperclip Runner heartbeat and wait for an active provider turn. 2. Restart only the Paperclip server, with or without a hot-restart marker. 3. Observe that the old startup path does not reconstruct the native control-plane authority. 4. Kill both the server and runner after a provider checkpoint. 5. Observe that the old path cannot resume the exact native session on the original heartbeat run. **Paperclip version or commit** The defect was reproduced from commit `1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto the current `master`. **Deployment mode** Local development and self-hosted server deployments that use the local Paperclip Runner. ## What Changed - Added correlated hot-restart requests and version-compatible native handoff fields. - Added controller boot identity, process-start identity, controller generation, recovery state, request id, and bounded history to the native finalization ledger. - Added transactional recovery claims for live-runner reattach, dead-runner resume, and incomplete bootstrap. - Added fail-closed ownership takeover rules and process identity validation. - Added live runner adoption to the local runner transport without a duplicate spawn. - Added same-run provider checkpoint resume and legacy retry-row compatibility. - Reconciled proposed and terminal results before runner or provider recovery. - Bound the HTTP and PRP listener before startup recovery and delayed scheduling and generic reapers until classification completes. - Added restart-aware health diagnostics, run-log recovery transitions, durable runner diagnostics, and bounded shutdown finalizer draining. - Moved restart-survivable diagnostics into runner-owned, pre-redacted bounded writes; raw stdout and stderr are never persisted. - Added process-start fencing for controller, runner, and provider PIDs; startup classifies every candidate without an implicit cap. - Added crash-recoverable, contention-safe development restart-request coordination and failed-startup listener cleanup. - Added a credential-free real-process restart suite for eight restart, scale, and identity scenarios. - Documented native restart operation, persistence, diagnostics, and verification. ## Verification - The documented native restart commands passed. They ran eight real-process/database recovery scenarios and the live runner adoption transport test. - Native executor tests passed: 111 tests. - Heartbeat recovery tests passed: 124 tests. - Hot restart, health, and shutdown tests passed: 52 tests. - The broader affected server suite passed: 350 tests. - Focused native recovery and startup tests passed: 49 tests. - Runner transport and control-plane tests passed: 63 tests. - Runner-owned diagnostic tests passed for write-time bounding, credential redaction, private file modes, and raw stream non-persistence. - Development restart coordination tests passed: 11 tests. - Database migration checks and the partial-application/replay regression test passed. - Server, database, and Paperclip Runner typechecks passed. - `git diff --check` passed. - Full Paperclip PR CI passed, including build, canary, all five general server shards, all five serialized server shards, all three browser E2E shards, workspace suites, and release-registry verification. - Greptile completed at 5/5 with no outstanding findings, recommendations, follow-ups, or open review threads. ## Risks - Moderate risk. This changes startup ordering and ownership transfer for active native runs. - The migration adds nullable columns and does not rewrite existing rows. - Recovery fails closed when process or durable session identity is incomplete or contradictory. - The first implementation supports the local Paperclip Runner. Remote targets keep their existing behavior. - The real-process suite covers cleanup and asserts that no runner or provider process survives each test. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime did not expose a more specific model revision or context-window size. Repository editing, shell execution, database tests, and real-process test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |