mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
0cedb45df3ed830baf5293cabe59ff03bb6091b2
501
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0cedb45df3 |
build(deps-dev): bump typescript from 5.9.3 to 7.0.2 (#11880)
Bumps [typescript](https://github.com/microsoft/TypeScript) from 5.9.3 to 7.0.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/microsoft/TypeScript/releases">typescript's releases</a>.</em></p> <blockquote> <h2>TypeScript 7.0.2</h2> <p><a href="https://devblogs.microsoft.com/typescript/announcing-typescript-7-0/">https://devblogs.microsoft.com/typescript/announcing-typescript-7-0/</a></p> <p>This tag was originally released at: <a href="https://github.com/microsoft/typescript-go/releases/tag/typescript%2Fv7.0.2">https://github.com/microsoft/typescript-go/releases/tag/typescript%2Fv7.0.2</a></p> <h2>TypeScript 6.0.3</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.2%22">fixed issues query for TypeScript 6.0.2 (Stable)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.3%22">fixed issues query for TypeScript 6.0.3 (Stable)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.2%22">fixed issues query for TypeScript 6.0.2 (Stable)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0.1 RC</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0-rc/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0 Beta</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0-beta/">release announcement</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22+is%3Aclosed+">fixed issues query for Typescript 6.0.0 (Beta)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/microsoft/TypeScript/commit/1e4744d68260a7cb91b62b12edc3f6a2187faaf1"><code>1e4744d</code></a> Merge branch 'main' into ts7-release</li> <li><a href="https://github.com/microsoft/TypeScript/commit/a5a219c3b5da0db4fa0ecf6c0b1f588c9af9c669"><code>a5a219c</code></a><code>microsoft/typescript-go#4558</code></li> <li><a href="https://github.com/microsoft/TypeScript/commit/ecfe30dce91368d52c9a49b6095bb0b673a238f8"><code>ecfe30d</code></a> Update status localization</li> <li><a href="https://github.com/microsoft/TypeScript/commit/5de25b5f8fec2ca35eadaed041f1f06d2e214895"><code>5de25b5</code></a> Hide executable name in TypeScript status</li> <li><a href="https://github.com/microsoft/TypeScript/commit/d7ce74a75da2b80e8201506a1599c06549432b93"><code>d7ce74a</code></a> Show bundled TypeScript version for packaged servers</li> <li><a href="https://github.com/microsoft/TypeScript/commit/29be66a607707f90d7a53103a4469bb3015a4d54"><code>29be66a</code></a> Correct TS 7 release version to 7.0.2</li> <li><a href="https://github.com/microsoft/TypeScript/commit/ed2bd1bfa4aac5211ce4bc58fcd1313c7eddc8ff"><code>ed2bd1b</code></a> Merge branch 'main' into ts7-release</li> <li><a href="https://github.com/microsoft/TypeScript/commit/887307575c58ea640dbeba3b4e8fdb6347cd3044"><code>8873075</code></a> Bump the github-actions group across 1 directory with 3 updates (microsoft/ty...</li> <li><a href="https://github.com/microsoft/TypeScript/commit/9427131ae2d4e230a90ee8a09daac4e75da3e311"><code>9427131</code></a> Set up stable / nightly extension split, other prep (microsoft/typescript-go#...</li> <li><a href="https://github.com/microsoft/TypeScript/commit/d4eaca5460a1f5f02a829e62706794b0a6fb903e"><code>d4eaca5</code></a><code>microsoft/typescript-go#4549</code></li> <li>Additional commits viewable in <a href="https://github.com/microsoft/TypeScript/compare/v5.9.3...v7.0.2">compare view</a></li> </ul> </details> <details> <summary>Maintainer changes</summary> <p>This version was pushed to npm by <a href="https://www.npmjs.com/~microsoft1es">microsoft1es</a>, a new releaser for typescript since your current version.</p> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
397de98193 |
feat(runner): add flagged Codex execution adapter (#12188)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner now has protocol, provider, tool, package, persistence, and hidden server boundaries. > - The server still cannot select that path for a real agent heartbeat. > - A new runtime must not change any existing direct adapter. > - An experimental runtime must fail closed when its rollout flag is off. > - This pull request adds one guarded Codex vertical slice through runnerd. > - The benefit is a production-built runner path that users cannot start by default. ## Linked Issues or Issue Description Refs #11962 Refs #12111 Refs #12169 Refs #12176 **Subsystem affected** Cross-cutting. The change affects the runner package, server orchestration, shared settings, and adapter configuration UI. **Problem or motivation** The hidden PRP coordinator cannot execute a real heartbeat. The application also needs an explicit rollout boundary before it can expose the experimental runner. Existing direct adapters must keep their current execution and finalization behavior. **Proposed solution** Add `paperclip_runner` as a Codex-only adapter behind the default-off `enableNativeRunner` instance flag. Select the native runtime only for that adapter. Persist the run binding before runnerd starts. Wait for the durable PRP result and terminal event. Resume the real Codex provider thread on later heartbeats. Keep persisted native runs readable and recoverable after the flag changes. **Alternatives considered** The server could route `codex_local` through runnerd. That option would change an existing adapter and weaken rollback safety. The server could expose all providers now. That option would add unreviewed provider behavior. The build could depend on a prebuilt runner binary. That option would make source builds architecture-dependent and difficult to verify. **Roadmap alignment** This work supports the shipped enforced-outcomes, governed-tool, and self-healing-run milestones. It does not add a new roadmap surface. It is the guarded execution step after the merged hidden runner boundaries. **Additional context** This is the next replacement for the closed large runner pull request. Task-thread presentation remains a separate follow-up so this change can preserve the current direct-adapter UI. ## What Changed - Add `paperclip_runner` as an explicit Codex-only adapter. - Add the default-off `enableNativeRunner` instance flag. - Reject fresh create, hire, import, switch, and execution requests while the flag is off. - Allow edits to persisted runner agents while the flag is off. - Recover an already persisted native run even after the flag is disabled. - Keep every built-in direct adapter on its existing runtime path. - Persist an immutable native run binding and revisioned completion contract before runnerd starts. - Execute server to PRP to runnerd to Codex to server through the hidden coordinator. - Validate the durable result against the terminal event and exact completion criteria before finalization. - Preserve the Codex provider thread ID and use `thread/resume` on the next heartbeat. - Strip unsupported Codex configuration fields from the experimental adapter. - Build a target-native release runner binary from source and vendor it into the server distribution. - Install Rust only in the Docker build stage. Do not add a workflow or lockfile change. - Stop the runner process group on completion, cancellation, and forced shutdown. ## Verification - Run `pnpm --filter @paperclipai/paperclip-runner check:all`. All 69 TypeScript tests and 58 Rust tests pass. Protocol, conformance, replay, formatting, and generated-file checks pass. - Run the 12 focused adapter, settings, runtime-selection, coordinator, direct-isolation, and real Codex integration test files. All 186 tests pass. - The real integration test uses PostgreSQL, HTTP, WebSocket, runnerd, and a fake Codex app server. It proves one `thread/start` followed by one `thread/resume`. - Run `pnpm -r typecheck`. - Run `pnpm build`. - Run `pnpm check:token-gates`. - Build the Docker `build` target from a clean context. Confirm that the server distribution contains an executable `paperclip-runnerd` built with Debian Rust 1.85. - Start the server through the source-mode tsx entry point with the package `dist` directory absent. Confirm the vendor shim resolves source exports and the server boots. - Run `pnpm test:run` twice. On this macOS host, 405 files pass and 1 file skips. Eight untouched workspace and loopback tests fail because macOS resolves `/tmp` and `/var` through `/private` and because PID-derived test ports exceed 65535. Linux CI must pass the full suite. - Confirm that the diff contains 52 files. Confirm that it contains no `.github` or `pnpm-lock.yaml` change. ## Risks - The feature flag is off by default. A fresh native start fails with a stable error while the flag is off. - A persisted native run remains recoverable after the flag changes. This prevents rollout changes from corrupting recorded work. - Only local Codex execution is accepted. Other providers and remote work modes fail closed. - Existing direct adapters do not start runnerd, create native rows, use native status arbitration, or enter native finalization. - The runner receives its one-use bootstrap ticket through the child environment. The server does not put the ticket in command arguments or logs. - The server validates the company, task, agent, run, runner, session, completion contract, result, and terminal binding before it accepts completion. - The build compiles a target-native Rust binary. Cross-platform release packaging remains a later concern. Source builds and Docker builds compile for their current target. - Docker needs enough build memory for the existing server TypeScript compile. The Docker build stage sets a 4 GB V8 heap limit. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment ID and context-window size are not exposed. The model used agentic reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and applicable tests pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fcb84d472d |
feat: already-imported transfer error names the landed company (#12144)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Chunked company-import transfers are deduplicated by content: a
byte-identical zip that already finished an apply is rejected
> - The rejection said only "this exact package was already imported by
a completed transfer" without saying where that import went
> - Users who could not find the earlier import read the rejection as
data loss and kept retrying, or exported again and created duplicate
companies
> - This pull request makes the declaration response carry the company
the completed apply created, and both clients name it in the error
> - The benefit is that the dedupe rejection now points at the existing
import instead of implying it vanished
## Linked Issues or Issue Description
**What existing behavior does this improve?**
The `alreadyCompleted` rejection when re-declaring a chunked
company-import transfer.
**Subsystem affected**
Shared transfer contract
(`packages/shared/src/company-import-transfer.ts`), transfer declaration
route (`server/src/routes/companies.ts`), web import page, CLI import
command.
**Current behavior**
`POST /api/companies/import/transfers` returns `alreadyCompleted: true`
with no pointer to the earlier import. Web and CLI raise "This exact
package was already imported by a completed transfer. Re-export the
package to import it again."
**Proposed behavior**
The response includes an optional `company` field (`{id, name,
issuePrefix} | null`) resolved from the completed run's company link.
Web and CLI raise a shared message: `… It created the company
"Paperclip" (PAPA) — open it from the company switcher. Re-export the
package to import it again.` A company that was deleted since (or a link
that was never written) degrades to `null` and the original message.
**Breaking changes**
None. The new response field is optional; old clients ignore it.
## What Changed
- `CompanyImportTransferCreated` gains optional `company`, plus a shared
`buildAlreadyImportedMessage` used by both clients.
- The declaration route's `alreadyCompleted` branch resolves the landed
company null-safely via `companyService.getById`.
- Web (`ui/src/pages/CompanyImport.tsx`) and CLI
(`cli/src/commands/client/company.ts`) raise the shared message.
## Verification
- `cd packages/shared && npx vitest run
src/company-import-transfer.test.ts` — 3 tests (named company, id
fallback, no-company original message).
- `cd server && npx vitest run
src/__tests__/company-import-transfer-routes.test.ts` — 24 tests; the
re-declaration test now asserts the company payload and the
deleted-company null path.
- `cd cli && npx vitest run
src/__tests__/company-import-transfer.test.ts` — 17 tests; new test pins
the named-company message.
- `cd ui && npx vitest run src/pages/CompanyImport.test.tsx` — 23 tests.
- `pnpm run typecheck` clean in shared, server, ui, cli.
## Risks
- Low risk. The lookup runs only on the `alreadyCompleted` branch and is
null-safe; the transfer run is already scoped to the requesting actor
(user + instance context in the actor key), so the response never names
a company the caller did not import.
## Model Used
- Claude Fable 5 (`claude-fable-5`, Anthropic) with extended thinking
and tool use, via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
11f6c754c9 |
feat: dedicated import pause reason with visible paused-assignee notices (#12140)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import parks every imported agent as a safety default, and issue assignment wakes are dropped for paused agents > - The pause was recorded as the generic reason "system" and was almost invisible: the chat-style task thread showed nothing, the legacy notice had no action, and the new-task dialog gave no hint > - Users assigned tasks in an imported company, nothing ran, and there was no explanation — the imported company looked broken > - This pull request records a dedicated "import" pause reason and makes the paused state visible and fixable where the user is looking > - The benefit is that a silent no-op becomes an explained state with a one-click resume ## Linked Issues or Issue Description **What existing behavior does this improve?** Working with a company whose agents arrived paused from a company import. **Subsystem affected** Shared constants (`PAUSE_REASONS`), company import service (`server/src/services/company-portability.ts`), task thread and new-task dialog UI. **Current behavior** Imported agents get `pauseReason: "system"`, the same value plugin-managed and built-in agent pauses use. Assigning an issue to a paused agent silently drops the wake. The chat-style task thread renders no paused notice; the legacy thread's notice says "It was paused by the system." with no action and only renders when the composer is shown. **Proposed behavior** Import writes `pauseReason: "import"`. The paused-assignee notice explains the import pause, offers an inline "Resume agent" button (suppressed for budget pauses, which clear on their own), and renders for read-only viewers. The chat-style task thread shows the same notice above the composer. The new-task dialog warns when the selected assignee is paused. **Breaking changes** None. `PAUSE_REASONS` is widened, not changed; the column already stores free-text values in other paths, and every consumer is an equality check with a manual fallback, so an older client shows the generic fallback copy for the new value. ## What Changed - `packages/shared/src/constants.ts`: `"import"` added to `PAUSE_REASONS`. - `server/src/services/company-portability.ts`: the import pause patch writes `pauseReason: "import"`. - `ui/src/components/IssueChatThread.tsx`: `IssueAssigneePausedNotice` gains import copy, a Resume button, test ids, and is exported; it now renders even when the composer is hidden. New `onResumeAssignee` / `resumeAssigneePending` props. - `ui/src/components/TaskChatThread.tsx`: renders the paused-assignee notice above the composer dock (the chat-style thread previously had no paused surface at all). - `ui/src/pages/IssueDetail.tsx`: wires a resume mutation (`agentsApi.resume`) through both thread variants and invalidates the company agent list. - `ui/src/components/NewIssueDialog.tsx`: inline note when the chosen assignee is paused, with import-specific copy. ## Verification - `cd server && npx vitest run src/__tests__/company-portability.test.ts` — 82 tests pass (pause pin updated to `"import"`). - `cd ui && npx vitest run src/components/IssueChatThread.test.tsx src/components/NewIssueDialog.test.tsx src/components/TaskChatThread.test.tsx` — 120 tests pass (new: notice copy per reason, resume click, budget suppression, active-agent null render, dialog note). - `pnpm run typecheck` in `packages/shared`, `server`, and `ui` — clean. ## Risks - Low risk. The resume action calls the existing `POST /agents/:id/resume` route with its existing guards. Existing rows keep `"system"` and fall back to the current generic copy; only new imports write `"import"`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) with extended thinking and tool use, via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4d2af732ae |
feat(runner): add native persistence contracts (#12169)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need durable records so Paperclip can explain results and final status changes. > - The current heartbeat tables support direct adapters, but they do not model native runner evidence. > - The runner transport and server coordinator must share a strict finalization contract before they write production data. > - This pull request adds that contract and its additive database boundary. > - It does not select the Paperclip Runner or change any existing adapter execution path. > - The benefit is a reviewable persistence layer that preserves all current behavior and supports later guarded integration. ## Linked Issues or Issue Description Refs #11962 Refs #12129 ## What Changed - Add native run result, finalization, completion, assessment, status decision, and status effect tables. - Add inert native metadata to heartbeat runs and events. Keep `legacy` as the default runtime mode. - Bind each evidence relationship to one company, issue, run, contract, result, assessment, and decision with composite constraints. - Add a strict `paperclip.native_finalization.v1` shared type and validator. - Preserve database functions, triggers, and the unique indexes required by foreign keys in JavaScript backups. - Add migration, backup, mixed-owner denial, validator, and direct-adapter compatibility tests. - Document the new records and their ownership rules. ## Verification - Run `pnpm -r typecheck`. - Run `pnpm build`. - Run `pnpm db:generate`. The schema output and migration safety checks remain current. - Run `PAPERCLIP_PSQL_PATH=/Applications/Postgres.app/Contents/Versions/latest/bin/psql pnpm exec vitest run packages/shared/src/validators/native-finalization.test.ts packages/db/src/client.test.ts packages/db/src/backup-lib.test.ts server/src/__tests__/heartbeat-workspace-busy.test.ts server/src/__tests__/heartbeat-comment-wake-batching.test.ts`. All 52 tests pass. - The full local `pnpm test:run` run completed 4,688 tests. It found 30 existing macOS test-environment failures. A serial rerun with the canonical `/private/tmp` path reduced those failures to six existing listener-diagnostics and skill-browser cases. None of those suites use files in this change. - The full Linux GitHub Actions matrix passes. This includes all general-server, serialized-server, workspace, browser, build, typecheck, canary, and aggregate verification jobs. - Greptile passes at 5/5. Contributor trust, Superagent, Socket, and Snyk pass with no finding from this change. - Storybook visual regression skips by path because this pull request has no UI or Storybook change. - Confirm that the diff contains 25 files. Confirm that it contains no workflow or `pnpm-lock.yaml` changes. ## Risks - The migration adds tables, columns, indexes, a function, a trigger, and ownership constraints. It does not remove or rename existing data. - Composite foreign keys reject mixed-company, mixed-issue, and mixed-run evidence even when each ID exists. - The status-version trigger runs only when an issue status changes. Backup tests confirm that restore retains this trigger and its dependencies. - Native source identifiers are unique when present. Legacy event rows remain unchanged. - This change does not add a unique run sequence constraint. The later native writer must allocate its sequence atomically before that invariant can be safe. - Existing adapters keep their current execution and finalization paths. New heartbeat runs default to `legacy` mode. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment ID and context-window size are not exposed. The model used agentic reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
445547c989 |
feat(duplex): run the Daytona sandbox callback bridge over Node HTTP/2 (#12120)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers carry agent work through controlled execution channels > - The Daytona callback bridge uses a bespoke line-framed protocol over its duplex channel > - The bespoke protocol adds framing work and does not use the Node transport that already supports multiplexed streams > - This pull request carries raw bytes across the channel, adds a Node HTTP/2 bridge, and selects it for Daytona > - The benefit is one authenticated, multiplexed callback session with queue_v1 as the bounded fallback ## Linked Issues or Issue Description **Subsystem affected** The packages/plugins Daytona provider and the shared duplex execution path. **Problem or motivation** The Daytona callback bridge uses a bespoke line-framed protocol over the provider duplex channel. This adds protocol work and limits stream handling. **Proposed solution** Carry raw bytes through the cross-layer channel. Add an authenticated Node HTTP/2 host server and sandbox client gateway. Select http2_v1 for Daytona and retain queue_v1 as the fallback. **Alternatives considered** Keep the current duplex_v1 protocol. This keeps the bespoke framing path and does not provide one HTTP/2 session for callback streams. **Roadmap alignment** ROADMAP.md lists Daytona under cloud and sandbox agents. This change improves the shipped Daytona provider path. **Additional context** The branch adds no dependency. Node 24 provides the http2 module. The host token check and canonical path parser remain the single dispatch path. ## What Changed - Carry raw Uint8Array chunks through the adapter, plugin, worker, runtime, and Daytona layers. - Encode bytes as base64 only across the JSON-RPC hop, because JSON has no binary type. - Add the bounded host HTTP/2 server and the in-sandbox HTTP/2 client gateway. - Authenticate every stream with the per-run bridge token before route work. - Parse the path once and reuse the canonical result for route and forwarding work. - Select http2_v1 for Daytona and fall back once to queue_v1 when the client preface is absent. - Add transport, session, stream, and fallback telemetry. - Mark HTTP/2 as the preferred transport and queue_v1 as the soft-deprecated fallback. ## Verification - `npx vitest run packages/adapter-utils/src` — 990 passed and 4 skipped. - `npx vitest run server/src/__tests__/plugin-worker-manager-duplex.test.ts` — 32 passed. - `npx vitest run --config packages/plugins/sandbox-providers/daytona/vitest.config.ts` — 220 passed and 6 skipped. - `npx tsc --noEmit` in `packages/adapter-utils`, `packages/shared`, `packages/plugins/sdk`, and `server` — clean. - No `package.json` or `pnpm-lock.yaml` file changed. - The live Daytona test skips when `DAYTONA_API_KEY` is absent. - The root `npx tsc --noEmit` command has a pre-existing missing `packages/adapters/droid-local` reference on this branch and on `master`. ## Risks - The transport change affects several duplex layers and could expose byte-boundary errors. - A missing HTTP/2 client preface falls back once to queue_v1 and records `preface_missing`. - The host token check and canonical path parser must remain on the shared dispatch path. - The live Daytona test needs `DAYTONA_API_KEY` and does not run in this agent sandbox. ## Model Used OpenAI GPT-5, tool-enabled coding agent with repository inspection, GitHub CLI, and shell execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d1573244b5 |
refactor: disambiguate the Telemetry and Observability data paths (#12128)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip records first-party events, OpenTelemetry data, and local run-log events > - The code and documents used one term for these three data paths > - This naming made the required review level unclear > - This pull request names each data path in the module names, documents, and code comments > - The benefit is a clear review rule without a runtime change ## Linked Issues or Issue Description **Issue type** Unclear or confusing. **Where is the issue?** `packages/shared/src/telemetry/README.md`, `doc/observability.md`, `doc/run-log-events.md`, and the duplex instrumentation modules. **What's wrong?** The repository used Telemetry for first-party events, OpenTelemetry data, and local run-log events. This usage made the data path and review level unclear. **Suggested fix** Use Telemetry only for Paperclip first-party events. Use Observability for OpenTelemetry data. Use the run log for rows in `heartbeat_run_events`. Related public pull requests: #8476 and #9672. ## What Changed - Rename the duplex instrumentation modules and identifiers from `Telemetry` to `Observability`. - Move the Observability and run-log contracts out of the Telemetry README. - Add `doc/observability.md` and `doc/run-log-events.md` as the canonical documents. - Add a file-path review rule to `AGENTS.md`. - Correct the remaining code comments that name the wrong data path. - Keep all event names, payloads, database records, spans, configuration keys, environment variables, and runtime paths unchanged. ## Verification - `npx vitest run packages/shared/src/telemetry/readme-contract.test.ts` passes. - `npx vitest run packages/adapter-utils/src/published-exports.test.ts` passes. - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-timing.test.ts` passes with 42 tests. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - `pnpm --filter server typecheck` passes. - The old module name does not remain in TypeScript or JSON files, except for the intentional publication guard. - CI and Greptile checks remain pending after PR creation. ## Risks - The old duplex module subpath no longer has a compatibility shim. The board accepted this intentional hard break. - The new duplex module subpath stays blocked from package publication. - The change has no runtime effect. The main risk is an incorrect document or module reference. ## Model Used OpenAI GPT-5 Codex, exact model ID `gpt-5`, with tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR with the documentation issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a14e51d592 |
refactor(environment): classify environment capabilities from static driver definitions (#12045)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environment runtime drivers provide workspace, lease, and custom image behavior > - Runtime code used driver identity checks and several capability-specific members > - These checks spread capability rules across the runtime and made new drivers harder to verify > - This pull request adds one general capability classifier and one static driver support table > - The benefit is one fail-closed capability model that keeps current behavior and supports future drivers ## Linked Issues or Issue Description **What existing behavior does this improve?** Environment runtime capability checks for workspace realization, custom images, lease capabilities, and duplex authorization. **Subsystem affected** Cross-cutting (multiple of the above) **Current behavior** The runtime selects several capability paths from driver identity and separate capability members. Custom image gates also trust provider declarations without checking every matching live worker method. **Proposed behavior** The runtime uses one general capability classifier and one static support table. Custom image gates require both the provider declaration and every matching live worker method. The public capability names remain unchanged. **Reason and benefit** The change keeps capability rules in one place. It removes identity conditions from runtime consumers and makes unsupported drivers fail closed. **Breaking changes** None. The public names sandboxCapabilities, sandboxProviders, and EffectiveSandboxCapabilities remain available. ## What Changed - Add classifyEnvironmentCapabilities and static support definitions for all four driver families. - Add resolveCapabilities to every environment runtime driver. - Move driver traits into environment-driver-traits.ts and migrate runtime consumers. - Require provider declarations and matching live worker methods for all custom image gates. - Migrate duplex authorization to the general resolver and remove the dead sandbox-only member. - Delete the unused resolveEffectiveSandboxCapabilities wrapper and update its test. ## Verification - pnpm --filter @paperclipai/server typecheck - pnpm exec vitest run server/src/__tests__/environment-capability-contract.test.ts server/src/__tests__/environment-runtime.test.ts — 92 tests pass - pnpm exec vitest run server/src/__tests__/environment-driver-traits.test.ts server/src/__tests__/general-capability-classifier.test.ts — 12 tests pass - pnpm exec vitest run server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/environment-execution-target-capabilities.test.ts server/src/__tests__/environment-execution-target-duplex.test.ts server/src/__tests__/environment-execution-target-duplex-kill-switch.test.ts — 66 tests pass ## Risks The main risk is a capability gate that denies a valid driver or permits an invalid driver. The static support matrix, live worker method checks, and regression tests reduce this risk. No database, public API, or published type name changes. ## Model Used OpenAI Codex, GPT-5, with tool use and code execution. The deployment does not provide a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (for example, docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c62bb4b16b |
feat: environment delete with agent reassignment and consented sandbox destroy (#12053)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments define where agent runs execute: local, SSH, or provider sandboxes > - Operators can create and edit environments, but the UI has no way to delete one > - The server already exposes `DELETE /environments/:id` and a delete-blast-radius preflight, but no UI consumes them, and a delete blocked by reusable sandbox leases gives the operator no path forward > - This pull request adds the delete flow to the environment configuration page: a preflight-driven modal that reassigns dependent agents, names the workspaces that hold blocking sandbox leases, and can destroy those sandboxes with explicit consent > - The benefit is that operators can retire stale environments from the UI without database surgery, and dependent agents move to a chosen replacement instead of silently falling back ## Linked Issues or Issue Description Refs #8554 Refs #11124 **Subsystem affected** Environments (server routes, environment runtime service, and the environment settings UI). **Problem or motivation** The environment configuration page has no delete control. The server delete endpoint exists, but nothing in the UI calls it. When reusable sandbox leases block a delete, the 409 error names no owner, so the operator cannot find the blocking workspace. Agents that use the environment as their default lose it silently through the FK `on delete set null`. **Proposed solution** Add a delete button with a confirmation modal on the environment edit page. The modal reads the delete-blast-radius preflight. It offers a dropdown to reassign dependent agents to another environment before the delete. It lists each workspace that holds a blocking reusable sandbox lease, with a link. When those leases are the only blocker, the confirm button destroys the sandboxes inline (`?destroyReusableSandboxLeases=true`) and then deletes. A failed teardown falls back to `pending_cleanup` for the sweep, so no sandbox is orphaned. ## What Changed - `ui/src/pages/CompanyEnvironments.tsx`: delete button on the edit page header, confirmation modal with agent reassignment select, lease-holder list, impact notes, and a consent-labeled destroy-and-delete action - `ui/src/api/environments.ts`: `deleteBlastRadius` and `remove` client methods; `remove` takes an optional `destroyReusableSandboxLeases` flag - `server/src/routes/environments.ts`: `DELETE /environments/:id` accepts `?destroyReusableSandboxLeases=true`; it destroys the environment's reusable sandbox leases first, but only when those leases are the sole delete blocker, then re-checks the blast radius before it deletes - `server/src/services/environment-runtime.ts`: new `destroyReusableSandboxLeasesForEnvironment` — destroys every reusable sandbox lease an environment still owns while the environment config (provider credentials) is still available - `server/src/services/environments.ts`: the delete blast radius now returns `reusableSandboxLeaseHolders` (lease id, workspace, issue) so clients can name what blocks a delete - `packages/shared/src/types/environment.ts`: `EnvironmentDeleteReusableLeaseHolder` type on the blast radius - Tests: route gating for the consent flag (destroy runs, mixed-blocker rejection, surviving-lease rejection), runtime destroy scoped to an environment, blast-radius holder join, and UI tests for the reassignment flow, holder links, and the consent button ## Verification - `npx vitest run server/src/__tests__/environment-routes.test.ts server/src/__tests__/environment-service.test.ts server/src/__tests__/environment-runtime.test.ts ui/src/pages/CompanyEnvironments.test.tsx` - Manual: open Settings → Environments → edit an environment. The trash icon opens the modal. With agents on the environment, pick a reassignment target and confirm; agents move and the environment deletes. With reusable sandbox leases, the modal names the holding workspaces and the confirm button reads "Destroy N sandboxes and delete". ## Risks - The consented path destroys provider sandboxes. It runs only when reusable leases are the sole blocker, so a delete that would still be rejected never destroys anything. A failed teardown routes to `pending_cleanup` and the delete stays blocked until the sweep resolves it. - Agent reassignment issues one PATCH per agent from the client. A mid-sequence failure leaves some agents reassigned; the reassignments are valid on their own and the UI refreshes to the actual state. - Hard blockers (managed local, instance default, pending cleanup) keep the existing 409 behavior and disable the confirm button. ## Model Used - Claude (Anthropic) — Fable 5, model id `claude-fable-5`, extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
05b35d4669 |
feat(duplex): bound aggregate duplex route resource consumption with a process-owned byte ledger (#12003)
## Thinking Path > - Paperclip runs AI agents through adapters and sandboxed execution targets. > - Duplex routes retain bytes across route data, broker messages, decoder buffers, and readiness replay. > - Per-route limits bound each route but do not bound the total retained bytes across many routes. > - A process-owned ledger must charge each retained buffer before allocation and release the charge during cleanup. > - This pull request adds the aggregate ledger, connects it to host and sandbox duplex paths, and adds route coverage. > - The benefit is a fail-closed process-wide byte limit that keeps concurrent duplex work within a safe resource budget. ## Linked Issues or Issue Description **Subsystem affected** This change affects packages/adapter-utils and server duplex orchestration. **Problem or motivation** Many routes can each stay below their per-route limits while their combined retained bytes exceed a safe process budget. **Proposed solution** Add a process-owned aggregate byte ledger. Charge route data, broker bytes, decoder buffers, and readiness replay bytes before allocation. Release each charge during cleanup. Use a separate sandbox_process decoder cap for the in-sandbox path. **Alternatives considered** Keep only per-route limits. This does not bound the combined process use. Set a fixed limit at one call site. This misses retained bytes in other duplex paths. **Roadmap alignment** This is a tightly scoped reliability and resource-safety improvement. It does not duplicate a roadmap feature. **Additional context** The aggregate ceiling uses a safe 256 MiB default. An invalid override falls back to that default and reports the rejected value. ## What Changed - Add a process-owned aggregate byte ledger for duplex route resource use. - Charge and release route data, broker forward and response bytes, decoder buffers, and readiness replay bytes. - Bound host-to-worker pending writes and standard input transport bytes. - Add a separate decoder cap for the sandbox_process path. - Make invalid aggregate-ceiling overrides fall back to the safe default without host startup failure. - Add adapter-utils and server tests for charging, release, rejection, cleanup, and many-route aggregate limits. ## Verification - pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit - pnpm --filter @paperclipai/server exec tsc --noEmit - Run the focused adapter-utils duplex ledger and execution-target tests. - Run the server aggregate-ledger route test. - Confirm all required pull request checks pass on this branch. ## Risks The ledger touches several duplex buffer paths. A missed release could reduce later capacity until process restart. The tests cover charge, release, rejection, cleanup, and route aggregation. The change uses a safe default when configuration input is invalid. ## Model Used OpenAI GPT-5 Codex. The runtime model ID and context window are not exposed to this task. The model used tool calls, shell commands, and code review workflow support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes/Closes/Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub issue references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c5050396c7 |
fix(daytona-duplex): chunk host-to-sandbox writes and make a transport close legible (#11986)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agent work through adapters and sandbox providers > - The Daytona duplex path sends host input through a provider pseudo-terminal WebSocket > - Large messages exceed the provider limit, and a transport close can look like a process exit > - This pull request chunks UTF-8 input and carries transport-close state through the duplex path > - The benefit is reliable large input and accurate loss reporting ## Linked Issues or Issue Description **What happened?** The Daytona duplex path sent a full input payload as one WebSocket message. A payload above the provider limit closed the channel. The wait path also mapped a non-numeric exit result to a process exit without exit data. **Expected behavior** The provider must receive large input as ordered UTF-8 chunks. A transport close without exit data must record `transport_closed`, while a numeric exit must record `provider_exit`. **Steps to reproduce** 1. Start a Daytona duplex session. 2. Send an input payload larger than 65536 bytes. 3. Observe that one message closes the provider channel. 4. End a session without a numeric exit code. 5. Observe that the loss reason reports a process exit. **Paperclip version or commit** Commit `1761e79ec9097c65d94f90a8ba20416f8ab718a6`. **Deployment mode** Built from source with the Daytona sandbox provider. ## What Changed - Add a shared UTF-8 byte chunker with a 32768-byte cap. - Route both Daytona pseudo-terminal write paths through the chunker. - Preserve multi-byte UTF-8 sequences across read-side chunks. - Carry an explicit `transportClosed` state through the worker and host wait paths. - Record `transport_closed` for a reason-less transport close and `provider_exit` for a numeric exit. - Keep orderly completion suppression for both exit paths. ## Verification - The Daytona plugin suite passes 194 tests. - The adapter-utils broker, codec, and telemetry suites pass 73 tests. - The plugin SDK duplex and worker RPC host suites pass 37 tests. - The server plugin worker manager duplex suite passes 78 tests. - The execution target sandbox and ACPX execute suites pass 257 tests. - TypeScript checks pass for adapter-utils, plugin SDK, server, and the standalone Daytona plugin. ## Risks The chunk size adds a loop for large input payloads. The 32768-byte cap stays below the provider limit. The optional loss field preserves compatibility for other providers. ## Model Used OpenAI Codex, GPT-5, extended reasoning, tool use, and code execution. The runtime does not expose a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
10d2781a29 |
feat(sandbox): add the duplex bridge broker, gated transport selection, and fixed observability (#11769)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox adapters provide controlled execution for untrusted provider environments. > - The sandbox channel needs one persistent duplex transport with strict host control. > - The transport must remain off unless the instance setting and provider capability both allow it. > - The host must detect loss, bound resource use, and expose only safe telemetry. > - This pull request adds the broker, gated selection, kill-switch wiring, fixed observability, and real-process proof. > - The benefit is safer sandbox execution with bounded failure behavior and inspectable transport results. ## Linked Issues or Issue Description No public issue exists for this change. The related pull requests are #11738 and #11750. **Problem or motivation** The sandbox duplex channel needs a host-controlled broker, strict transport gates, bounded provider input, and safe loss telemetry. Without these controls, a provider can cause replay, resource growth, unsafe endpoint selection, or data exposure through telemetry. **Proposed solution** Add a host broker with nested time limits, request limits, one-shot loss, and per-id deduplication. Select duplex transport only when the instance setting and provider capability both equal true. Assign the endpoint and nonce on the host. Reject invalid readiness data and use the file bridge on failure. Add fixed redacted telemetry and a real-process end-to-end test harness. **Alternatives considered** Keep the file bridge as the only transport. This avoids new channel behavior but does not provide persistent duplex operation for supported sandbox providers. **Roadmap alignment** This change supports the Cloud / Sandbox agents section in ROADMAP.md. ## What Changed - Add the duplex bridge broker with bounded forward, response, and gateway wait budgets. - Bound concurrent requests, lifetime requests, and request-id bytes before retention or forwarding. - Select duplex transport only when both required gates are true. - Assign the loopback port and nonce on the host and enforce a liveness-only READY frame. - Fall back to the file bridge after invalid readiness, contamination, bind failure, or timeout. - Carry the kill switch through the server, acpx engine, and six local adapters. - Add fixed, redacted duplex telemetry with a provider allowlist. - Add a real-process end-to-end harness for readiness, round trips, loss, and teardown. - Add regression coverage for limits, loss, UTF-8 splits, concurrency, and telemetry dimensions. ## Verification - Adapter-utils, server, and Daytona typechecks pass locally. - Adapter-utils tests pass, including the codec, broker, execution-target sandbox, and real-process harness. - Server kill-switch tests pass. - Live Daytona tests pass with the required provider key and skip without that key. - The root pnpm-lock.yaml file has no diff. - The branch contains ten commits after origin/master. ## Risks - Duplex transport remains disabled unless both gates equal true. - A provider remains an untrusted boundary and needs least-privilege credentials and quotas. - The server telemetry recorder stays deferred; the default recorder does nothing. - A provider that pre-binds the host port causes a fail-closed fallback to the file bridge. - The change adds no database migration and changes no root lockfile. ## Model Used OpenAI GPT-5, exact model family GPT-5, large context window, reasoning, and tool use. The model assisted with Git handoff validation and PR preparation. The implementation commits came from the engineering worktree. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (e.g. docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
417336f8be |
fix(workspaces): attach PR preparation to existing branches (#11703)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Execution workspaces isolate an agent task from the primary checkout. > - Pull request preparation can need a branch that already contains completed work. > - The workspace policy could not require an exact existing branch. > - Workspace cleanup also treated worktree creation as branch ownership. > - This pull request adds an exact existing-branch policy and separate branch ownership metadata. > - The benefit is safe pull request preparation that preserves every existing commit and operator-owned branch. ## Linked Issues or Issue Description **What happened?** A pull request preparation run could not pin its execution workspace to an exact existing branch. Workspace reuse and cleanup could also confuse worktree creation with branch ownership. **Expected behavior** The run must attach only to the requested branch in an isolated Git worktree. It must fail if the branch is missing, busy, or inconsistent. Cleanup must not delete a branch that Paperclip does not own. **Steps to reproduce** 1. Create a branch that contains completed work. 2. Configure a pull request preparation task to use that branch. 3. Start the task and observe that the prior policy cannot require the exact branch. **Paperclip version or commit** This behavior reproduces on the base revision before this pull request. **Deployment mode** Local development with isolated Git worktrees. ## What Changed - Add `existingBranch` to the execution workspace policy and shared validation contracts. - Require `existingBranch` to use an isolated Git worktree and reject conflicting branch templates. - Attach to the exact branch without creating, renaming, resetting, or deleting it. - Track branch ownership separately from worktree creation and use that ownership during cleanup. - Return HTTP 422 for invalid existing-branch settings on every issue-producing route. - Add a bounded repair script for existing pull request preparation tasks. - Add focused policy, route, heartbeat, runtime, and ready-comment tests. - Document the exact-branch behavior and safety rules. ## Verification - `pnpm exec vitest run server/src/__tests__/execution-workspace-policy.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/issue-existing-branch-validation-status.test.ts server/src/__tests__/workspace-runtime.test.ts server/src/services/workspace-runtime-exposure.test.ts server/src/services/workspace-runtime-ready-comment.test.ts` passed 335 tests. - `pnpm -r typecheck` passed for all workspace projects. - `pnpm test:run` passed 4,431 tests. Two unrelated embedded-Postgres setup hooks timed out under aggregate load. Their isolated rerun passed 74 tests. - `pnpm build` passed for all workspace projects. - The two review regressions passed with 139 unrelated tests skipped. - All latest-head CI gates passed after one unrelated timing-sensitive test passed on rerun. - Greptile scored the latest head 5/5 with no unresolved review threads. ## Risks - Invalid workspace settings now return HTTP 422 instead of a generic validation response. - The exact branch must already exist and must not be checked out by another worktree. - The new policy fails closed when it cannot prove branch identity or ownership. - This change has no database migration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex from the GPT-5 family assisted with this change. The runtime did not expose its exact deployment ID or context window. The agent used high-reasoning mode, repository tools, shell execution, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
adfbe2d4b9 |
feat(environments): refer to the managed default environment by name, not the sandbox driver key (#11838)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed deployments provision a platform-managed default environment
for agent runs; the UI shows this environment in selectors, the agent
form, run details, and the environments page
> - Those surfaces append the raw driver key to the environment name, so
users see labels like "Paperclip Computer (sandbox)", "Paperclip
Computer · sandbox", and fallback copy such as "Managed sandbox" and
"The sandbox has no ready authentication"
> - "sandbox" is infrastructure vocabulary, not the product name of the
environment; showing it next to the managed environment's name is
confusing and off-brand
> - This pull request renders platform-managed environments by name
alone and rewords the sandbox-phrased copy, while user-created
environments keep the driver suffix so mixed lists stay distinguishable
> - The benefit is that the default environment reads as one clear
product name everywhere, and self-hosted users lose nothing: their own
environments still show the driver
## Linked Issues or Issue Description
**What existing behavior does this improve?**
Display of the platform-managed default environment across the UI.
**Subsystem affected**
UI (environment selectors, agent config form, environments page, agents
page, run details) and the claude-local/codex-local adapter auth checks.
**Current behavior**
The agent form labels the inherited default environment as "Name
(sandbox)". Environment selectors and the environments list render "Name
· sandbox". The agents page describes the environment as "<provider>
sandbox provider". The agent form's fallback label is "Managed sandbox".
Adapter auth checks say "The sandbox has no ready authentication for
this adapter."
**Proposed behavior**
Platform-managed environment rows (`metadata.managedByPaperclip`) render
their name alone. The fallback label is "Paperclip Computer". The agents
page describes managed environments as "Managed by Paperclip". Run
details omit the driver suffix for sandbox-driver environments (the
adjacent Provider entry already identifies the mechanism). Adapter auth
checks say "This environment has no ready authentication for this
adapter."
**Reason and benefit**
The managed environment carries a product name. Appending the raw driver
key ("sandbox") to it is noise and contradicts the product naming.
User-created environments keep the driver suffix, so mixed lists stay
distinguishable.
**Breaking changes**
None. Message text of the auth check is not read programmatically; the
UI keys off `ADAPTER_AUTH_MISSING_CHECK_CODE`. Rows without the managed
marker render exactly as before.
## What Changed
- New `environmentDisplayLabel` helper in
`ui/src/lib/managed-sandbox-environment.ts`: managed rows → name alone;
other rows → "Name · driver".
- `AgentConfigForm`: inherited-default label uses the helper; fallback
copy "Managed sandbox" → "Paperclip Computer"; environment options use
the helper.
- `ProjectProperties`, `CompanyEnvironments`: environment selector
options use the helper; the environments-list row hides the driver
suffix on managed rows; the managed detail page's fallback description
no longer says "sandbox".
- `Agents` page: managed environments are described as "Managed by
Paperclip" instead of "<provider> sandbox provider".
- `CommentThread` run details: the driver suffix is omitted for
sandbox-driver environments.
- claude-local and codex-local adapters: auth-missing check message/hint
reworded from "sandbox" to "environment" (ACP and environment-test
paths); claude-local probe/effort/login hints reworded the same way.
- Run status lines: "Syncing workspace to sandbox", "Exporting git
changes from sandbox", "Starting adapter in sandbox", and friends now
say "environment"; "Finalizing sandbox workspace" → "Finalizing
workspace". Templated transfer-progress lines map the `sandbox`
transport key to "environment" for display (`runtime-progress.ts`).
- Agent form sign-in panel: "Sign in to the sandbox" → "Sign in to the
environment"; "Authenticated. The sandbox has credentials now." → "…The
environment has credentials now."
- Feature catalog + instance settings card: "Managed Sandbox Only" →
"Managed Environment Only" (setting key unchanged; the card keeps its
alphabetical slot).
- Server agents routes: execution-target failure and test-identity copy
no longer say "sandbox"; workspace-mode label "Cloud sandbox" → "Cloud
environment".
- Tests: new `environmentDisplayLabel` unit cases; new `AgentConfigForm`
render case asserting the managed default renders without "(sandbox)" or
"· sandbox"; status-line assertions updated across adapter-utils, server
heartbeat/live-run, and UI chat suites.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `pnpm --filter @paperclipai/adapter-claude-local typecheck` and
`--filter @paperclipai/adapter-codex-local typecheck` — clean.
- `vitest run` for `managed-sandbox-environment.test.ts`,
`AgentConfigForm.render.test.tsx`, `CompanyEnvironments.test.tsx`,
`Agents.test.tsx`, `CommentThread.test.tsx`, `NewAgent.test.tsx` — all
green (118 tests across the two runs).
## Risks
Low risk. Cosmetic label changes only; no data or API changes. Rows
without `metadata.managedByPaperclip` render exactly as before, so
self-hosted deployments with their own environments see no change. The
only self-hosted-visible wording changes are the adapter auth-check
message and the driver suffix omission on sandbox-driver rows in run
details.
## Model Used
- Claude (Anthropic) — claude-fable-5 (Claude Fable 5), Claude Code CLI,
extended thinking, tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
docs reference these labels)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
38d8f37172 |
fix(build): enforce Node 24 across Paperclip (#11792)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs across the CLI, server, adapters, plugins, CI, and container images. > - These surfaces declared different Node.js versions from 20 through 24. > - A newer `@types/node` major can expose APIs that the supported runtime does not provide. > - Node.js 20 is no longer a suitable project baseline, and Node.js 24 is the current LTS line. > - This pull request sets Node.js 24.11.0 as one repository-wide baseline, adds a drift check, and gives users actionable startup guidance when their runtime is too old. > - The benefit is one clear runtime contract for development, release, installation, and published packages. ## Linked Issues or Issue Description Refs #2734 Refs #11727 Refs #739 ## What Changed - Require Node.js 24.11.0 or newer in all 42 package manifests and runtime checks. - Use Node.js 24 in GitHub Actions, Docker images, smoke images, sandbox setup, portable installs, and esbuild targets. - Align every direct `@types/node` declaration on `^24.0.0`. - Prevent Dependabot from opening major `@types/node` upgrades without a matching runtime decision. - Add `.nvmrc` and a CI policy check for Node version drift. - Update ACP version gates, tests, and user documentation for the new minimum. - Print a non-blocking warning on CLI and server startup when Node is unsupported, with remediation through a version manager or the documented downloaded `install.sh` workflow. - Deduplicate that warning when `paperclipai run` boots the CLI and server in the same process. ## Verification - `node scripts/check-node-version-policy.mjs` - `node --check scripts/check-node-version-policy.mjs` - `node --check cli/esbuild.config.mjs` - `node --check scripts/generate-npm-package-json.mjs` - `bash -n scripts/install.sh scripts/test-install-sh-docker.sh scripts/e2e-install-lifecycle.sh` - Parsed all 42 package manifests and confirmed `engines.node` is `>=24.11.0`. - `git diff --check` - `vitest run packages/adapter-utils/src/sandbox-install-command.test.ts` passed with 3 tests. - `vitest run cli/src/node-version.test.ts` passed with 4 tests. - Directly exercised the shared warning helper for unsupported-version messaging and same-process deduplication. - The focused exe.dev suite could not resolve the locally unbuilt plugin SDK from this isolated worktree. A full offline workspace install was also blocked because the package-manager signature verifier requires registry access. The full suite was not run locally; draft CI performs a clean install and evaluates the wider impact. ## Risks - This is a breaking runtime change for users, plugins, and deployments that still use Node.js 20 or 22. - Published workspace packages will now produce an engine warning or failure in strict package managers on older Node.js releases. - Node.js 24 can reveal dependency, native module, Playwright, or agent CLI compatibility issues in CI. - The bootstrap installer now installs Node.js 24 when the current runtime is older than 24.11.0. - The portable sandbox fallback is pinned to Node.js 24.11.0 and depends on that upstream tarball remaining available. - Unsupported runtimes continue booting after a warning, so a later incompatibility can still fail at its point of use. - The CLI and server share the warning policy through the published `@paperclipai/shared` package; packaging checks must keep that subpath export available. - This PR does not commit `pnpm-lock.yaml` because repository policy assigns lockfile generation to CI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact deployment ID and context window are not exposed in this session. Reasoning, repository tools, shell execution, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3abe9e2134 |
build(deps): bump zod from 3.25.76 to 4.4.3 (#11719)
Bumps [zod](https://github.com/colinhacks/zod) from 3.25.76 to 4.4.3. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/colinhacks/zod/releases">zod's releases</a>.</em></p> <blockquote> <h2>v4.4.3</h2> <h2>Commits:</h2> <ul> <li>4c2fa95ce3f3390fbc522324e406b4e9e89b88f9 docs: use Zernio primary wordmark for gold sponsor logo</li> <li>2aeec83eb135e3a83756e973ef44845fc5a455d2 docs: prune lapsed gold sponsors and rebalance logo sizing</li> <li>7391be88ac1ee5cd02057f5ccc012a1f5df4efd0 docs: prune lapsed silver/bronze sponsors and add active ones</li> <li>2c703322a21b4e2b12f33f49ea8430c451a68b4f docs: normalize bronze sponsor logos to github avatar pattern</li> <li>9195250cab0e7950efe39c3926d6c203b4b0a170 docs: remove Mintlify from bronze sponsors (churned)</li> <li>b8dffe9e62f17e6571e6249d05cc5102b54d94e4 docs: remove Numeric and Speakeasy (2+ missed monthly cycles)</li> <li>1cab69383fcdeae2a366d5e2a2fc4d8fc765d168 fix(v4): restore catch handling for absent object keys (<a href="https://redirect.github.com/colinhacks/zod/issues/5937">#5937</a>) (<a href="https://redirect.github.com/colinhacks/zod/issues/5939">#5939</a>)</li> <li>c2be4f819064eed62c7c350a2d399b5faecd15f8 fix(v4): generalize optin/fallback to transform; restore preprocess on absent keys (<a href="https://redirect.github.com/colinhacks/zod/issues/5941">#5941</a>)</li> <li>f3c9ec03ba7a28ae72d25cc295f38674bee0f559 4.4.3</li> <li>1fb56a5c18c27102dbc92260a4007c7732a0ccca docs: document release procedure in AGENTS.md</li> </ul> <h2>v4.4.2</h2> <h2>Commits:</h2> <ul> <li>0c62df0ea19fd05abdf90473e9eef7eea530fab2 Clean up docs navigation and stale labels (<a href="https://redirect.github.com/colinhacks/zod/issues/5901">#5901</a>)</li> <li>20cc794895cc8604fe0c87d83a5d1c3f89fad0ac chore: add security policy and refresh tooling deps</li> <li>6fbe07b0177efdd1bf1c0b05160e70d7a0702337 fix(docs): heading anchor links now include the hash so it doesnt scoll all the way up, follows navbar logic (<a href="https://redirect.github.com/colinhacks/zod/issues/5791">#5791</a>)</li> <li>4bbed1b1c73eca4ce9e59b1189ed236aa6c8b5bd Tighten discriminated union option typing</li> <li>bbac3e567e7fccfaaf7cdc97f1ce30c295e2c908 Update PR guidance for agents</li> <li>cf0dc942a32805c292fff59ade20a7ace980735a Merge remote-tracking branch 'origin/main' into fix-discriminated-union-key-constraint</li> <li>292c894a5fd2aa42e527900b83d8d7a3009a709c docs: add Zernio gold sponsor</li> <li>1fc9f311c28dcf80d0bb5a36b177086cbc3d8eca docs: document codec inversion</li> <li>1373c85da9aeff704a9762d27bc58699618aefb7 docs: remove AI disclosure guidance</li> <li>e20d02b473c08e3a4e557bc610b1b5fac079b649 chore: ignore triage notes</li> <li>e58ea4d91b1dfe8194b73508203213cbc7e9c936 docs: test Zod Mini tab code heights</li> <li>905761a5d127e8d5dd2ebb3bc88c75cb0b8149ff docs: document preprocess input type narrowing</li> <li>bf64bac850d4dee2b7dde7e64909d5d796d32043 chore: tighten test guidance in AGENTS.md</li> <li>8ec4e73f4c4693b6361ad591be40fb41eb8a9f95 chore: update play.ts scratch</li> <li>02c2baf7d0d615872fa4528a8020603b71211702 Make z.preprocess defer optionality to inner schema (<a href="https://redirect.github.com/colinhacks/zod/issues/5929">#5929</a>)</li> <li>88015df8e25c44fb5385eb3ef28935119cd5edea fix(docs): drop deprecated <code>baseUrl</code> from tsconfig</li> <li>c59d4474e3b4cad1b323462186cf607178ce8267 4.4.2</li> </ul> <h2>v4.4.1</h2> <h2>Commits:</h2> <ul> <li>481f7be4238c83ed58183f921b2646f340a91c6a ci: gate release publishing on full test workflow</li> <li>95ccab423aec720b2523c3a64cdc7e3204537cc7 test(v3): restore optional undefined expectations</li> <li>cede2c63739a5823d6aa5093d291e9a111da943d fix(v4): reject tuple holes before required defaults (<a href="https://redirect.github.com/colinhacks/zod/issues/5900">#5900</a>)</li> <li>edd0bf0f5ada4a8dc581c259407d7bbad0a71ea7 release: 4.4.1</li> <li>180d83d1dbe6a59260710cc8637a3dea2281ee56 docs: remove Jazz featured sponsor</li> </ul> <h2>v4.4.0</h2> <h2>4.4.0</h2> <p>This is a minor release with a wide set of correctness and soundness fixes. Some fixes intentionally make Zod stricter, so code that depended on previously accepted invalid or ambiguous inputs may need small updates.</p> <h2>Potentially breaking bug fixes</h2> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/colinhacks/zod/commit/1fb56a5c18c27102dbc92260a4007c7732a0ccca"><code>1fb56a5</code></a> docs: document release procedure in AGENTS.md</li> <li><a href="https://github.com/colinhacks/zod/commit/f3c9ec03ba7a28ae72d25cc295f38674bee0f559"><code>f3c9ec0</code></a> 4.4.3</li> <li><a href="https://github.com/colinhacks/zod/commit/c2be4f819064eed62c7c350a2d399b5faecd15f8"><code>c2be4f8</code></a> fix(v4): generalize optin/fallback to transform; restore preprocess on absent...</li> <li><a href="https://github.com/colinhacks/zod/commit/1cab69383fcdeae2a366d5e2a2fc4d8fc765d168"><code>1cab693</code></a> fix(v4): restore catch handling for absent object keys (<a href="https://redirect.github.com/colinhacks/zod/issues/5937">#5937</a>) (<a href="https://redirect.github.com/colinhacks/zod/issues/5939">#5939</a>)</li> <li><a href="https://github.com/colinhacks/zod/commit/b8dffe9e62f17e6571e6249d05cc5102b54d94e4"><code>b8dffe9</code></a> docs: remove Numeric and Speakeasy (2+ missed monthly cycles)</li> <li><a href="https://github.com/colinhacks/zod/commit/9195250cab0e7950efe39c3926d6c203b4b0a170"><code>9195250</code></a> docs: remove Mintlify from bronze sponsors (churned)</li> <li><a href="https://github.com/colinhacks/zod/commit/2c703322a21b4e2b12f33f49ea8430c451a68b4f"><code>2c70332</code></a> docs: normalize bronze sponsor logos to github avatar pattern</li> <li><a href="https://github.com/colinhacks/zod/commit/7391be88ac1ee5cd02057f5ccc012a1f5df4efd0"><code>7391be8</code></a> docs: prune lapsed silver/bronze sponsors and add active ones</li> <li><a href="https://github.com/colinhacks/zod/commit/2aeec83eb135e3a83756e973ef44845fc5a455d2"><code>2aeec83</code></a> docs: prune lapsed gold sponsors and rebalance logo sizing</li> <li><a href="https://github.com/colinhacks/zod/commit/4c2fa95ce3f3390fbc522324e406b4e9e89b88f9"><code>4c2fa95</code></a> docs: use Zernio primary wordmark for gold sponsor logo</li> <li>Additional commits viewable in <a href="https://github.com/colinhacks/zod/compare/v3.25.76...v4.4.3">compare view</a></li> </ul> </details> <details> <summary>Maintainer changes</summary> <p>This version was pushed to npm by <a href="https://www.npmjs.com/~GitHub%20Actions">GitHub Actions</a>, a new releaser for zod since your current version.</p> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db4defdfbf |
feat: operator-configurable settings visibility via PAPERCLIP_HIDDEN_SETTINGS (#11823)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The instance settings surface (Access, Plugins, Adapters, General,
Experimental) assumes the person at the keyboard operates the whole
instance
> - Operators who host Paperclip for others — a managed cloud or an
internal shared server — expose settings pages and toggles that do not
apply to their deployment, and the related mutation APIs stay open
> - A hosted tenant can open Plugins or Adapters, try an action, and hit
a confusing failure, because only a few hardcoded platform floors exist
> - This pull request adds a generic, operator-configured visibility
mechanism: one env var hides declared settings surfaces in the UI and
floors their mutation routes with a stable 403 code
> - The benefit is a clean hosted-tenant settings surface for any
operator, with zero behavior change for normal self-hosted instances
## Linked Issues or Issue Description
**Subsystem affected**
Instance settings (server routes and UI), the shared settings registry
in `packages/shared`, and the `/api/health` bootstrap payload.
**Problem or motivation**
An operator who hosts Paperclip for other people cannot hide settings
surfaces that the platform manages. Tenants see Access, Plugins, and
Adapters pages, backup retention, and host-level experimental toggles
that do nothing useful for them. The mutation APIs behind these surfaces
also stay open, so a tenant admin can attempt actions the platform must
control. ROADMAP.md names a cleaner shared deployment story as a goal
("Teams should be able to run the same product in hosted or semi-hosted
environments without changing the mental model").
**Proposed solution**
Add a declarative registry of hideable settings surfaces and one env
var, `PAPERCLIP_HIDDEN_SETTINGS`. The server parses the list at boot,
reports it on `/api/health`, and rejects value-changing writes to hidden
surfaces with a stable `settings_operator_managed` 403 code. The UI
reads the list from the health payload and removes the hidden pages,
sections, and toggles from navigation, routes, and page content. Unknown
keys warn and are ignored, so one list can roll across a fleet with
mixed app versions. With the variable unset, behavior is byte-identical
to today.
**Alternatives considered**
- Hardcode the hidden set for cloud instances in this repo: rejected,
because each hosting operator needs a different policy, and policy does
not belong in shared code.
- Deliver the hidden set through the managed-config document: rejected,
because that channel is cloud-specific and fail-closed on unknown
fields; a plain env var works for any operator, including self-hosted
shared servers.
- Lock the controls with a badge instead of hiding them: rejected for
these surfaces, because they are meaningless to tenants, not merely
platform-controlled; the existing managed-overlay lock stays the right
tool for controlled flags.
**Roadmap alignment**
Supports the "shared deployment story" item in ROADMAP.md: hosted and
semi-hosted deployments keep the same product with a settings surface
that matches what the tenant can actually do.
## What Changed
- New `packages/shared/src/settings-visibility.ts`: registry of hideable
surfaces (every instance settings page — profile, environments, access,
heartbeats, experimental, plugins, adapters; every Instance → General
section; every experimental flag as `instance.experimental.<key>`), the
`PAPERCLIP_HIDDEN_SETTINGS` parser, and the `settings_operator_managed`
error code. The General page stays visible as the settings root and
redirect target.
- New `server/src/services/settings-visibility.ts`: parse-once accessor;
unknown keys log one warning and are ignored.
- `/api/health` reports `hiddenSettings` on every response shape; the
field is omitted when nothing is hidden.
- Server floors on hidden surfaces, with same-value echo tolerance (the
`executionMode` precedent): field-backed general sections and
experimental keys reject value-changing PATCHes, and hiding the whole
Experimental page floors every toggle; plugin lifecycle and config
writes, adapter management writes, and the Access admin routes (reads
included) return 403 `settings_operator_managed`. Reads the app itself
needs (plugin `ui-contributions`, adapter metadata, plugin job trigger)
stay open. Pages without instance-scoped mutation routes are hidden in
the UI only.
- UI: new `useHiddenSettings` hook and `HiddenSettingsPageGate` route
gate (hidden pages redirect to the settings root); the settings sidebar
and tab bar drop hidden entries; remembered settings paths remap to the
default page; `InstanceGeneralSettings` skips hidden sections; every
`ExperimentalToggleCard` now carries its flag key and renders nothing
when hidden.
- Removed the dead `InstanceSidebar` component (referenced only by its
own test).
- Docs: `docs/deploy/environment-variables.md` documents the variable
and the key registry.
## Verification
- `pnpm vitest run` over the new and extended suites: shared registry
and parser, representative floor tests per route class (changed-value
403, same-value echo 200, unset env 200, page-level Experimental
hiding), the health field, the route gate, nav filtering, and
section/card hiding with one hidden example per surface kind — 168 tests
pass.
- Full root `pnpm typecheck` passes.
- Manual: booted a server with the variable set. `/api/health` lists the
keys; an unknown key logs one warning and the server boots; hidden pages
redirect; hidden sections and cards do not render; hidden-field PATCH
returns 403 with `details.code = "settings_operator_managed"`; a
same-value echo returns 200. Unset the variable: the full settings
surface returns and responses are byte-identical to master.
## Risks
- Low risk for self-hosted instances: with the variable unset, the
hidden set is empty, the health field is omitted, and no floor
activates.
- Flooring plugin config writes assumes hosted deployments configure
plugins through the platform. If a future bundled plugin needs
tenant-entered config, the floor needs a narrow carve-out.
- Hidden-key floors tolerate same-value echoes, so API clients that
round-trip full GET responses keep working.
- Hiding a toggle does not change its value; operators pair hiding with
the desired default where the value matters.
## Model Used
Claude Fable 5 (Anthropic, `claude-fable-5`) with extended thinking and
agentic tool use, driven through the Claude Code CLI (file edits, test
execution, and live-server verification loops).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
0fa318b8da | feat(artifacts): bridge Markdown work products into the document review surface (#11822) | ||
|
|
cb0009b097 |
fix: preserve recovery retries across restarts (#11817)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane must keep each active issue on a clear execution or recovery path. > - A missing issue disposition can require more than one bounded repair attempt. > - A server restart could lose that repair path or move source ownership to the recovery owner. > - A parked or expired retry could also make the user interface show a false healthy state. > - Concurrent recovery loops must not schedule the same repair attempt twice. > - This pull request keeps retry state durable, makes scheduling atomic, and keeps source ownership stable. > - The benefit is that recovery continues after a restart and operators see the correct state. ## Linked Issues or Issue Description **What happened?** A run that ended without a valid issue disposition could lose its repair path after a server restart. Manager recovery could also change the source owner. In addition, a parked or expired retry could make the issue look healthy when no active work existed. Concurrent reconciliation could also schedule the same repair attempt twice. **Expected behavior** Paperclip must keep bounded source and manager repair attempts across restarts. Recovery ownership must stay separate from source issue ownership. The server and user interface must report only a live retry as active work. Each repair attempt must be scheduled at most once per company. **Steps to reproduce** 1. Start an agent run on an issue. 2. End the run without a valid issue disposition. 3. Let the first repair attempt schedule a retry. 4. Restart the server, let the retry time pass without a live run, or start two reconciliation loops together. 5. Observe that the repair path can stop, the issue can show a false healthy state, or duplicate retries can be created. **Paperclip version or commit** The problem existed on `master` before candidate head `d8e620fe86bade7df18decac332007f5821ae04f`. **Deployment mode** The problem affects self-hosted servers and local builds that use automatic recovery. ## What Changed - Persist bounded source-owner and manager repair lineages with stable fingerprints and retry limits. - Resume incomplete disposition repairs after a server restart. - Keep recovery ownership separate from source issue ownership and enforce source mutation authority. - Project live retry evidence into issue and blocker summaries. - Show recovery owner, return owner, attempt count, and retry state in the board user interface. - Treat expired or parked retries as attention states unless a queued or running attempt exists. - Atomically deduplicate disposition-repair wake requests with a company-scoped partial unique index. - Reuse the winning run when concurrent reconciliation loses the uniqueness race, without duplicate scheduling activity. - Honor disabled on-demand wake policy before recovery scheduling and again before delayed retry promotion. - Keep the new index migration safe for lagging seeded databases that already contain the index. - Add server and user interface tests for recovery, restart, ownership, retry, concurrency, and blocker states. - Update the implementation and execution semantics documents. ## Verification - Focused server recovery and ownership suites: 282 tests passed on the repaired base candidate. - Focused user interface recovery suites: 128 tests passed on the repaired base candidate. - Atomic-deduplication schema and recovery suites: 111 tests passed on the first Greptile repair. - Recovery and scheduled-retry wake-policy suites: 126 tests passed at `d8e620fe86bade7df18decac332007f5821ae04f`. - The exact lagging-source migration-order test passed after the index migration became idempotent: 1 test passed and 62 unrelated tests were skipped. - `@paperclipai/db` and `@paperclipai/server` typechecks passed at the current head. - Migration generation and migration safety checks passed for migration `0226_tan_colossus.sql`. - `pnpm check:token-gates` passed on the repaired base candidate. - `pnpm -r typecheck` passed on the repaired base candidate. - `pnpm build` passed on the repaired base candidate. - `pnpm test:run` passed 4,540 tests on the repaired base candidate. Four fixed-port cases met listeners that already existed on the host. - The two unchanged fixed-port files passed in an isolated network namespace: 129 tests passed and 27 tests were skipped. - Independent Security and QA reviews approved `63c0423aab54c66f2293a20b0fb3f3b013ee3ba8`; exact-head re-review is required after automated checks settle on `d8e620fe86bade7df18decac332007f5821ae04f`. ## Risks - Recovery orchestration affects issue liveness and ownership. The new paths use bounded attempts, stable fingerprints, row locks, authority checks, and database uniqueness. - A conservative attention state can show more warnings when a scheduled retry has no queued or running attempt. It does not hide stopped work. - Migration `0226_tan_colossus.sql` creates a partial unique index on a known-large table. Migrations run transactionally, so `CONCURRENTLY` is unavailable. The matching disposition-repair key namespace is introduced by this release, so deployed databases have no matching rows before the index is added. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex from the GPT-5 model family used agentic reasoning, tool use, and code execution. The runtime did not expose the exact model ID or context window. - Anthropic Claude Opus 5 used a 1M context window, tool use, and code execution for part of the user interface repair, as recorded in the commit history. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
233c12f029 |
feat: add kimi-local adapter for Kimi Code CLI (CLI + ACP engines) (#9967)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local agent adapters (`claude_local`, `gemini_local`, `grok_local`, …) are the integration surface that lets Paperclip run coding CLIs on the host machine > - The Kimi Code CLI (`kimi`, Moonshot AI) has a documented non-interactive mode, `kimi -p --output-format stream-json` with session resume via `kimi -r`, but Paperclip has no built-in adapter for it > - So Kimi users (especially Kimi membership / OAuth subscribers) cannot onboard their CLI to Paperclip agent teams > - This pull request adds a complete built-in `kimi_local` adapter (both execution engines, session management, instructions + skills delivery, thinking-effort control, environment test, UI and CLI modules, docs) following the established `gemini_local`/`grok_local` package pattern > - Kimi Code ships an ACP server (`kimi acp`), so the adapter runs on Paperclip's shared acpx engine by default (streaming transcript with live tool status, like `claude_local`/`gemini_local`) and falls back to a headless CLI lane (`kimi -p --output-format stream-json`) when ACP prerequisites are unavailable > - The benefit is that Kimi Code becomes a first-class Paperclip agent lane: selectable in the UI, resumable across heartbeats, with the same operating context (instruction bundle, skills, effort) and streaming transcript the other local adapters get ## Linked Issues or Issue Description - Supersedes #9880 (same branch; expanded from the CLI-only lane into a complete adapter with the default ACP engine lane, control-plane skill install, and live transcript wiring) - Refs #9879 (adapter request for Kimi Code CLI, filed with this PR) - Refs #163 (original Kimi support request) Duplicate/related prior PRs, per the dedup search (both appear stale: no updates or maintainer review since May 2026, and both target an older Kimi CLI interface; calling them out for reviewer context per CONTRIBUTING.md): - Refs #6276 (`feat: add kimi-local adapter`): targets an older array-based content format (`{type: think}`/`{type: text}` blocks), not the current documented stream-json schema - Refs #5202 (`feat(adapter): add Kimi CLI local adapter with Wire protocol support`): builds on a `--wire` JSON-RPC interface that current Kimi Code CLI (0.27.0) no longer documents; the current documented headless interface is `-p --output-format stream-json` This PR is a fresh implementation against current master and the currently documented/verified Kimi CLI behavior (see Verification). Happy to fold in anything useful from the earlier attempts if a reviewer prefers. ## What Changed - **New adapter package** `packages/adapters/kimi-local` (`@paperclipai/adapter-kimi-local`), modeled on `gemini-local`/`grok-local`: - `src/server/execute.ts`: spawns `kimi -p <prompt> --output-format stream-json` (argv array, no shell), `-m <model>` only when configured, `-r <sessionId>` when the stored session cwd matches the run cwd, automatic fresh-session retry on unrecoverable-session errors, headless-safe env (`CI=1`, `NO_COLOR=1`, `KIMI_CODE_NO_AUTO_UPDATE=1`, `TERM=dumb`; user-configured values win), full remote (ssh/sandbox) execution lane with runtime install via `@moonshot-ai/kimi-code` - **Instruction bundle delivery**: the prompt path directive now names the sibling instruction files (`./HEARTBEAT.md`, `./SOUL.md`, `./TOOLS.md`) alongside the prepended entry file, and local runs pass `--add-dir <instructions-dir>` so Kimi can actually open them (matching `claude_local`). Without this, only the entry file reached Kimi and agents improvised the operating workflow that `HEARTBEAT.md` documents - **Thinking effort**: a configured `effort` is forwarded as the `KIMI_MODEL_THINKING_EFFORT` operational override (Kimi has no per-invocation effort flag). It is only sent for models that advertise `support_efforts` (currently `kimi-code/k3`) to avoid provider rejections, and `medium` maps to `high` since Kimi has no medium tier (`low`/`high`/`max` pass through) - **Skills delivery**: desired Paperclip skills are delivered via Kimi's `--skills-dir` flag from a dedicated per-run directory (a local snapshot, or the synced snapshot on remote targets), so skills load reliably and in isolation. Paperclip never overwrites the shared `$KIMI_CODE_HOME/skills` home, so skills installed by the operator or other agents are left intact. `--skills-dir` is only passed when at least one skill is desired, so unconfigured agents keep Kimi's default skill discovery - **Live run status**: the adapter now forwards each streamed stream-json line to `onEvent` (assistant `content` as an assistant snippet, `tool_calls` as tool-name events), which drives the issue-thread activity indicator (`currentToolName` / `lastAssistantSnippet` / `lastEventAt`). Previously the adapter only wrote the raw run log, so the issue thread showed a stale "no output for N s" line with no tool or reasoning context while Kimi worked. Tool results are omitted so the last meaningful "Using X" / snippet is not overwritten by a generic label - `src/server/parse.ts`: parses the verified Kimi stream-json event shapes (`assistant` text, `assistant.tool_calls` with JSON-string arguments, `tool` results, trailing `meta.session.resume_hint` for session-id capture) plus failure classifiers (`kimi_auth_required`, transient network, unrecoverable session). A signaled exit (null exit code, not a timeout) is now reported as a failure rather than coalesced to success, and the error message names the terminating signal - `src/server/skills.ts`: lists/syncs Paperclip skills for the adapter's skill-management surface - `src/server/test.ts`: environment test covering CLI resolution + `kimi --version`, cwd check, auth detection (OAuth credential dirs, keyed `[providers.*]` in config.toml, or the `KIMI_MODEL_NAME` + `KIMI_MODEL_API_KEY` env pair), and a live hello probe - `src/ui/` (stdout-line parser for transcripts, config builder) and `src/cli/` (stream event formatter) modules - Root metadata: three managed model aliases (`kimi-code/kimi-for-coding`, `kimi-code/kimi-for-coding-highspeed`, `kimi-code/k3`), effort-capable-model metadata (`EFFORT_CAPABLE_MODELS`, effort mapping helpers), `agentConfigurationDoc` - Tests: 101 tests across parse, execute (args building, resume gating, retry, auth error code, timeout, signaled-exit failure, effort forwarding/gating/mapping, `--add-dir` instructions directive, `--skills-dir` gating, `onEvent` runtime-event forwarding), ACP engine (engine resolution, acpx config build, node-version gate), ACP transcript delegation, environment test, UI parse/build-config - **ACP engine lane (default)** (`src/server/acp.ts` + shared `adapter-utils/acpx-engine`): Kimi Code ships an ACP server (`kimi acp`), so `kimi_local` now runs on Paperclip's shared acpx engine by default, matching `claude_local`/`codex_local`/`gemini_local`. The issue-thread transcript streams live (assistant text deltas, tool calls with a `pending`->`completed` status lifecycle) instead of the CLI lane's bursty complete-message output. Registered `kimi_local -> "kimi"` in `ACPX_ADAPTER_AGENT_IDS` and resolved the built-in agent command to `kimi acp`; `execute.ts` dispatches to the ACP executor first with an automatic CLI fallback when ACP prerequisites fail (`engine=acp` requires ACP, `engine=cli` pins the headless lane); `index.ts` falls back to the shared acpx session codec; the UI/CLI delegate `acpx.*` events to the shared acpx transcript parser and event formatter. The headless CLI lane (above) remains as the fallback - **Registration** (one entry each, mirroring existing adapters): server adapter registry + `BUILTIN_ADAPTER_TYPES`, `AGENT_ADAPTER_TYPES` (shared), UI adapter registry + display registry (`Kimi Code`, Moon icon) + capabilities defaults, CLI adapter registry, `Dockerfile` (package copy + `npm install --global @moonshot-ai/kimi-code@latest`), `vitest.config.ts` workspace, `scripts/release-package-manifest.json` - **Behavioral sets** mirroring `gemini_local` (Kimi resumes sessions the same way): `GIT_SENSITIVE_LOCAL_ADAPTER_TYPES`, `SESSIONED_LOCAL_ADAPTERS` (heartbeat + recovery), `REMOTE_MANAGED_ADAPTERS`, ssh/sandbox execution-target allow-lists, `ADAPTER_DEFAULT_RULES_BY_TYPE` (`timeoutSec: 0`, `graceSec: 15`), and `LEGACY_SESSIONED_ADAPTER_TYPES` + `ADAPTER_SESSION_MANAGEMENT` in adapter-utils - **UI touch-points**: New Agent default-model branch, AgentConfigForm command map (`kimi_local: "kimi"`) + model defaults + a Kimi-specific thinking-effort option list (`Low`/`High`/`Max`, reflecting Kimi's tiers rather than borrowing Claude's), OnboardingWizard (command map, model default, `kimi login` / `KIMI_MODEL_NAME + KIMI_MODEL_API_KEY` auth hints, manual-debug command line), InviteLanding enabled adapters - **Control-plane skill install** (`cli/src/commands/client/agent.ts`): `paperclipai agent local-cli` seeded the Paperclip control-plane skills into `~/.codex/skills` and `~/.claude/skills` so Codex/Claude agents auto-discover the API reference every run. Kimi had no equivalent target, so `kimi_local` agents began each session without the control-plane skill and rediscovered routes (e.g. the company-scoped `POST /api/companies/{companyId}/issues`) by trial and error. Added `~/.kimi-code/skills` (honoring `KIMI_CODE_HOME`) as a third install target for parity. Independent of the per-run `--skills-dir` delivery, which only applies to explicitly configured skills. - **Docs**: `docs/adapters/kimi-local.md` (prerequisites, auth options, config fields including `effort`, session resume, instruction bundle, skills delivery, control-plane skill install) + a row in `docs/adapters/overview.md` Out of scope (deliberately): model profiles, built-in agent `allowedAdapterTypes` additions. ## Verification\n\nCurrent-master rebase verification (OpenAI Codex, 2026-08-03): 13 focused files / 231 tests pass; adapter-utils, server, UI, CLI, and Kimi adapter typechecks pass; full repository build and UI token gates pass. The branch is conflict-free against master at head `1249df117c5e12e5771b9a570a6340866450619e`.\n\nAutomated (all from repo root, pnpm 9.15.4, Node 22): - `vitest run packages/adapters/kimi-local`: 89/89 pass (includes coverage for the instruction `--add-dir` directive, effort forwarding/gating/mapping, `--skills-dir` gating, the signaled-exit failure path, and `onEvent` runtime-event forwarding with cross-chunk line buffering) - `vitest run server/src/__tests__/adapter-registry.test.ts server/src/__tests__/adapter-routes.test.ts server/src/services/heartbeat-stop-metadata.test.ts ui/src/adapters/adapter-display-registry.test.ts`: 37/37 pass - `vitest run cli/src/__tests__/skills.test.ts`: 13/13 pass (the control-plane skill install target follows the existing Codex/Claude install path, whose symlink logic is unchanged) - `vitest run packages/shared`: 307/307 pass; `vitest run packages/adapter-utils`: pass except one pre-existing, unrelated failure (`mcp-isolation.integration.test.ts` requires Claude CLI ≥ 2.1.207; host has 2.1.185, fails identically on unmodified master) - `pnpm --filter @paperclipai/adapter-kimi-local typecheck|build`, plus typecheck of `server`, `ui`, `cli`, `adapter-utils`: all clean - `pnpm install --frozen-lockfile`: passes (the PR diff itself contains no lockfile changes, per repo policy; verified against a locally regenerated lockfile) - `node scripts/check-no-git-push.mjs` and `node scripts/check-forbidden-tokens.mjs`: pass - CI note: the `policy` job's release-bootstrap step is expected to stay red until a maintainer bootstraps the first npm publish of `@paperclipai/adapter-kimi-local`; see the CI Note for Maintainers comment. All other contributor-actionable checks are green. Manual end-to-end (real Kimi CLI 0.27.0, OAuth login, dev server on an isolated instance): 1. Server `GET /api/adapters` lists `kimi_local` as builtin with correct capability flags; models endpoint returns the three Kimi models 2. `POST .../adapters/kimi_local/test-environment`: all checks pass, including a live `kimi -p` hello probe 3. Created a `kimi_local` agent and invoked two heartbeats: run 1 spawned `kimi -p ... --output-format stream-json`, Kimi used its `Read` tool, produced the expected answer, and the session id was captured from the `session.resume_hint` meta event; run 2 resumed the **same** Kimi session (`sessionIdBefore == sessionIdAfter`) via `-r` 4. UI: adapter appears in the New Agent dropdown; selecting it shows the Kimi command placeholder, the three models, and the Kimi config fields; the run transcript renders Kimi tool calls via the adapter's stdout parser The instruction-bundle, thinking-effort, and `--skills-dir` changes landed after the manual run above. They are covered by the unit tests listed under Automated, and the Kimi CLI flags they rely on (`--add-dir`, `--skills-dir`, `KIMI_MODEL_THINKING_EFFORT`) were confirmed against the installed Kimi Code CLI 0.27.0 (`kimi --help`, config-file thinking-effort docs). Screenshots (assets branch on the fork, not part of the diff):       ## Risks - Low risk to existing behavior: the change is additive, one new workspace package plus single-entry registrations alongside existing adapters; no existing adapter code paths are modified. - The adapter invokes the locally installed `kimi` CLI; like other local adapters, run behavior depends on the host's Kimi version. The parser is written against the documented/verified 0.27.0 stream-json schema and degrades gracefully (malformed lines are skipped, failures surface as run errors). - `--skills-dir` overrides Kimi's auto-discovery of user and project skills for the run. This is intentional (paperclip-managed agents get a reproducible, isolated skill set), and it is only passed when at least one Paperclip skill is desired, so unconfigured agents keep default discovery. - Thinking effort is only forwarded to models that advertise `support_efforts` (currently `kimi-code/k3`); `EFFORT_CAPABLE_MODELS` must be extended when more Kimi models gain support, otherwise a configured effort is silently ignored for them. - `Dockerfile` now installs `@moonshot-ai/kimi-code@latest` globally alongside the other agent CLIs, so image size increases slightly. - Maintainer action needed for the npm bootstrap gate: the `policy` job's release-bootstrap step fails until the first npm publish of `@paperclipai/adapter-kimi-local` (the gate from #5146 that every new adapter package has passed through). Enrollment with `publishFromCi: true` is required by the manifest validator (dropping the entry, `false`, or `private` are all rejected), so this is intentionally left to a maintainer. Remaining CI lanes are expected to run once it is done. ## Model Used\n\n- **Current-master rebase, conflict adaptation, and registry-parity coverage:** OpenAI, **GPT-5 Codex** (Codex agent; exact serving model ID and context-window size were not exposed to the runtime), with repository, shell, Git, and GitHub tooling. It preserved Hawik’s commit authorship, reconciled ACPX and environment-capability changes, added current registry tests, and ran the verification above.\n- **Adapter implementation and initial review:** Moonshot AI, **Kimi K3 Coding** (latest), via **Kimi Code CLI v0.27.0** (`kimi-code/k3` alias, 1M-token context window, thinking mode, agentic tool use). The CLI agent explored the repo, wrote the adapter implementation (delegated to a coder sub-agent of the same model), ran tests, and drafted the first version of this PR body. A second model-driven review pass (read-only, same model) audited the diff for security/correctness before submission; its findings (shell-quoting hardening, auth-detection false positive, session-compaction registration, test gaps) were fixed and are included. - **Harness-context fixes and review responses:** Anthropic, **Claude Opus 4.8** (`claude-opus-4-8`) via Claude Code. Diagnosed from run logs that Kimi received only the entry instructions file (not the `HEARTBEAT.md`/`SOUL.md`/`TOOLS.md` bundle) and that `effort` was never wired, then implemented the instruction `--add-dir` delivery, `KIMI_MODEL_THINKING_EFFORT` forwarding, and `--skills-dir` skill delivery, added the accompanying tests and docs, and addressed the automated review comments (preserving external skills on remote sync, treating a signaled exit as a failure). Also extended the `paperclipai agent local-cli` installer to seed the control-plane skills into `~/.kimi-code/skills` for Codex/Claude parity, wired `onEvent` runtime events so the issue-thread activity indicator reflects Kimi's tool and reasoning output live, and built the ACP engine lane (`kimi acp` via the shared acpx engine, default) so the transcript streams with live tool status like the other ACP adapters. The Kimi CLI flags, subcommand, and env var relied on here were verified against the installed Kimi Code CLI 0.27.0. - All CLI behaviors claimed here (`-p`, `--output-format stream-json`, `-r` resume, event shapes, `--add-dir`, `--skills-dir`, `KIMI_MODEL_THINKING_EFFORT`) were verified empirically against the installed Kimi CLI, not assumed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green *(only the release-bootstrap step remains red, pending the maintainer npm publish described in Risks)* - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(will address all Greptile comments as they arrive)* - [x] I will address all Greptile and reviewer comments before requesting merge --- ## Maintainer Addendum (2026-08-20) The shared acpx-engine and issue-chat changes (run-summary segmentation, placeholder tool-event coalescing, `ISSUE_CHAT_TRANSCRIPT_MAX_VISIBLE_ENTRIES` 30 → 400, live-reasoning UI) have been **extracted to #11761** so the cross-adapter behavior changes review and revert independently — both commits there preserve @hawikk's authorship. This PR is now the kimi-specific adapter only (60 files, +3,793/−8, essentially pure addition); the only shared-engine touch left is the `kimi acp` command resolution. `publishFromCi` is `true` — the package name is bootstrapped on npm. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Devin Foley <devin@paperclip.ing> |
||
|
|
e826188e82 |
refactor(settings): unify settings and speed up exports (#11789)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - Operators use the settings area to control a company and its Paperclip instance. > - The current navigation separates related settings and uses duplicate instance pages. > - Company exports also do independent reads in sequence and do extra work for previews. > - Hardened workspace commands can differ from their saved command after loopback binding. > - This pull request makes these related operator workflows consistent and faster. > - The benefit is one clear settings area, faster exports, and stable runtime command matching. ## Linked Issues or Issue Description Refs #338 Related: #9834 **What existing behavior does this improve?** This improves the company settings UI, company export preparation, and workspace runtime command matching. **Current behavior** Company and instance settings use separate navigation and duplicate pages. Export preparation reads many independent records in sequence. Preview generation can also build an unused organization image. A command with a forced loopback bind can fail to match its saved runtime command. **Proposed behavior** Use one settings navigation and put general instance controls on the company General page. Load independent export data with bounded concurrency, skip unused preview image work, and load the export page only when it is needed. Treat the loopback-bound form of a command as the same runtime command. **Reason and benefit** Operators get one clear settings area. Large company exports need fewer serialized reads. Export previews and initial UI loads do less work. Hardened runtime services remain linked to their saved command definitions. **Breaking changes** The obsolete instance General URL redirects to the unified settings page. Access and Heartbeats remain available, and legacy bookmarks keep their destinations. No API response shape or database schema changes. ## What Changed - Unified company and instance settings navigation and removed duplicate instance settings pages. - Embedded general instance controls in the company General page and kept access-sensitive navigation behavior. - Preserved instance Access and Heartbeats controls in the unified navigation and normalized old bookmarks to those destinations. - Improved environment and access-state handling when workspace seed requests overlap. - Added bounded export reads, a lighter preview path, deferred export preparation, and lazy export-page loading. - Matched loopback-bound runtime commands to their saved command definitions. - Added focused shared, server, and UI regression tests. ## Verification - `pnpm exec vitest run <18 changed test files>`: 18 files and 256 tests passed. - `pnpm check:token-gates`: passed all four token gates. - `pnpm -r typecheck`: passed for all workspace projects. - `pnpm build`: passed for all workspace projects. - `pnpm test:run`: tests ran without a reported failure, but the runner did not close after the server handoff tests. The process closed with status 0 after an interrupt. - Focused latest-head route tests: 2 files and 4 tests passed. - GitHub latest-head checks: all completed without failure. - Greptile: 5/5 with no unresolved review threads. ## Risks - Medium risk: settings routes and navigation changed across several operator roles. - Medium risk: bounded export concurrency increases simultaneous database reads. The limits stay below the normal pool size. - Low risk: runtime command matching accepts only the known Tailscale HTTPS loopback transformation. - No migrations are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with a GPT-5-family coding model. The runtime does not expose the exact deployed model ID or context-window size. Reasoning, tool use, and local code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details Exception: This task requires the existing execution branch. The harness does not permit a branch rename. - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a9d1f740f0 |
fix(workspaces): seed managed worktrees when the base checkout has no config (#11752)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents do that work in isolated git worktrees, and a managed
worktree runs its own Paperclip instance with a cloned database
> - That clone needs a seed source, and the source must come from
server-owned registration, never from state the workspace itself can
rewrite
> - The seed-source resolver requires the registered base project
workspace to hold its own `.paperclip/config.json`
> - A managed project workspace is a plain `git clone`, and no code
writes that file into it
> - Every isolated worktree provision, deferred seed, and workspace
repair therefore fails on a managed checkout
> - This pull request lets a named source supply the config when the
base checkout has none
> - The benefit is that managed worktrees provision again, and the seed
source stays server-owned
## Linked Issues or Issue Description
No public GitHub issue exists for this problem. It is described below.
**What happened?**
Agent runs that need an isolated worktree fail during provisioning. The
provision command exits with this error (paths redacted):
```
Execution workspace provision command "bash ./scripts/provision-worktree.sh" failed:
Registered base project workspace has no canonical Paperclip config:
<instance-home>/instances/default/projects/<company-id>/<project-id>/<repo>/.paperclip/config.json
```
`resolveRegisteredWorktreeSeedSource` sets `registeredConfigPath` to
`<baseCwd>/.paperclip/config.json` whenever the caller names a
registered base workspace. It then requires that file to exist.
`scripts/provision-worktree.sh` applies the same rule.
A managed project workspace never has that file.
`materializeManagedProjectWorkspace` creates it with `git clone` and a
rename, so the checkout holds repository content only. The control plane
keeps its config at `<home>/instances/<id>/config.json` instead.
The failure reaches three paths: worktree provisioning, deferred seeding
through `worktree ensure-seeded`, and workspace repair.
The behavior changed in #11671. That pull request replaced a fallback
chain with a single hard requirement. Fixture code in
`scripts/__tests__/provision-worktree-self-heal.test.mjs` writes a
config into the fake base workspace, so tests kept passing.
**Expected behavior**
A managed worktree provisions and seeds from the registered source. The
seed manifest still never selects that source.
**Steps to reproduce**
1. Register the Paperclip repository as a project with a `repoUrl`, so
the server materializes a managed checkout.
2. Assign an issue to an agent whose workspace strategy is
`git_worktree`.
3. Watch the workspace operation log for the provision command.
4. The command exits non-zero with the error above.
**Paperclip version or commit**
Reproduced on `master` at
|
||
|
|
faab2620ad |
feat(sandbox): add a duplex transport for Daytona behind a default-off kill switch (#11750)
## Thinking Path > - Paperclip is an open source app that manages AI agents for work > - Paperclip runs agents in local and remote sandbox environments > - A sandbox needs a bounded channel for commands and asynchronous input > - Daytona needs a real pseudo-terminal transport for this channel > - The sandbox gateway also needs a mode that handles channel loss safely > - This pull request adds the Daytona transport and gateway mode behind a default-off kill switch > - The benefit is a tested foundation for later transport selection ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above): sandbox providers, plugin SDK, server settings, and shared types. **Problem or motivation** The merged sandbox protocol has no runtime transport for Daytona. The generated sandbox gateway also has no duplex mode. A later transport-selection change needs both parts and a safe per-run gate. **Proposed solution** Add a Daytona `duplexCommandStream` transport over a raw pseudo-terminal. Add a generated gateway mode named `duplex_v1`. Add the `enableSandboxDuplexBridge` setting with a default value of `false`. Keep transport selection disabled until a later pull request. **Alternatives considered** Keep the protocol unused until the transport-selection change. This would delay provider tests and leave the gateway path without direct coverage. **Roadmap alignment** This change supports the completed Roadmap item for cloud and sandbox agents. It extends the merged sandbox channel foundation in pull request #11738. **Additional context** The Daytona provider remains an untrusted boundary. Deployments must use least-privilege provider credentials and provider-side quota controls. Operators must name an owner for duplex telemetry retention before rollout. ## What Changed - Add the Daytona `duplexCommandStream` capability over a raw pseudo-terminal. - Add a launch wrapper that disables echo and newline translation for NDJSON frames. - Close channels on lease release, destroy, resume of a stopped worker, and worker shutdown. - Declare the capability in the Daytona manifest and set `PLUGIN_VERSION` to `0.1.5`. - Add the worker-to-host notification sink at `ctx.duplexChannel.data` and `ctx.duplexChannel.exit`. - Add the generated sandbox gateway mode `PAPERCLIP_API_BRIDGE_MODE=duplex_v1`. - Add channel-loss results of `409 outcome_indeterminate` and `503 bridge_unavailable`. - Add the per-run setting `enableSandboxDuplexBridge`, with a default value of `false`. - Add unit tests, generated-source codec tests, lifecycle tests, and a credential-gated live Daytona test. ## Verification - Daytona suite: 185 tests pass. - Adapter utilities: 754 tests pass and 4 tests skip. - Plugin SDK: 62 tests pass. - Shared package: 28 tests pass. - Server duplex tests pass. - Shared, plugin SDK, server, and Daytona TypeScript checks pass. - The live Daytona test passes 3 cases when `DAYTONA_API_KEY` is set. - The live Daytona test skips 3 cases without `DAYTONA_API_KEY`. - CI must run the full workspace typecheck, test, and build gates after PR creation. ## Risks - The Daytona control plane and pseudo-terminal remain untrusted boundaries. - The duplex gateway changes behavior only when the mode and per-run setting enable it. - A lost channel fails requests without replay, so callers must handle indeterminate outcomes. - The transport-selection change must require both `duplexCommandStream === true` and `enableSandboxDuplexBridge === true`. - The provider credential and quota limits need operator control before rollout. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8161244284 |
feat(sandbox): add opt-in duplex command-stream foundation (capability, protocol, bounded host route, frame codec) (#11738)
## Thinking Path > - Paperclip provides a control plane for companies that run AI agents. > - Sandboxed agents need a safe execution path for persistent command streams. > - The existing callback transport does not provide a bounded, generic duplex route. > - The host must control capability access, route identity, protocol limits, and close behavior. > - This pull request adds an opt-in duplex command-stream foundation across the sandbox layers. > - The feature stays inert because no current provider declares the capability. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Sandbox command execution needs a persistent host-to-sandbox stream. The current callback bridge uses a file transport and does not provide this generic route. **Proposed solution** Add a fail-closed provider capability, generic worker protocol messages, a host-owned bounded route, cross-layer service mediation, and a versioned newline-delimited frame codec. **Alternatives considered** Keep the file transport and add feature-specific commands. This does not provide one reusable duplex contract or host-owned route bounds. **Roadmap alignment** This work supports the completed Cloud / Sandbox agents roadmap area and the safe autonomy goal in the product definition. **Additional context** The change passed a two-stage security review. The final code review verdict was approve after fixes for active-stream bounds and service-layer capability mediation. ## What Changed - Add the opt-in `duplexCommandStream` provider capability with fail-closed narrowing. - Add duplex open, write, stop, and close requests and data and exit notifications to the plugin worker protocol. - Add a host-owned route with bounds for chunk size, cumulative bytes, lifetime, protocol errors, pending requests, and pre-bind buffering. - Add close acknowledgement handling with worker retirement when the close remains unconfirmed. - Wire `openDuplexChannel` through the execution target, runtime service, and plugin worker. - Add a versioned frame codec with shared wire-compatibility vectors and split UTF-8 handling. ## Verification - `server/src/__tests__/plugin-worker-manager-duplex.test.ts` passes 18 tests. - `server/src/__tests__/environment-execution-target-duplex.test.ts` passes 11 tests. - `packages/adapter-utils/src/duplex-frame-codec.test.ts` passes 38 tests. - `server/src/__tests__/sandbox-capability-contract.test.ts` passes 15 tests. - Setup-token pseudo-terminal regression tests pass 47 tests. - Server TypeScript check passes. - Continuous integration will run the full required test, typecheck, build, and policy checks. ## Risks - Providers that opt into the capability must implement the complete worker protocol. - Route limit defaults can close a stream when a workload exceeds the configured bounds. - The capability remains disabled for current providers, so current production behavior does not change. ## Model Used OpenAI GPT-5 (`gpt-5`), with tool use and code execution. The model reviewed and prepared this pull request from the supplied implementation and verification record. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bd059a073d |
fix(workspaces): make managed runtimes reliable across restarts (#11740)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution workspaces need isolated databases, ports, and runtime services > - Concurrent workspaces could reuse ports or lose service ownership after a restart > - A markerless worktree also needed seed recovery, but normal markerless instances still needed to boot > - This pull request makes seed, port, and service ownership state explicit and recoverable > - It also checks live process and listener identity before it reclaims shared resources > - The benefit is reliable workspace startup, restart, adoption, and concurrent provisioning ## Linked Issues or Issue Description **What happened?** Managed workspaces could lose runtime service ownership after a control-plane restart. Concurrent worktrees could also reuse a port when their parent paths differed. A seed recovery change made every markerless instance resolve a worktree seed source, so normal instances without a source could not start. **Expected behavior** Paperclip must preserve healthy managed services across restarts. It must reserve unique ports across worktree parents. It must provision a registered markerless worktree, but it must skip seed work for a normal markerless instance. **Steps to reproduce** 1. Start two managed worktrees under different parent paths at the same time. 2. Restart the control plane while a managed service stays alive. 3. Start Paperclip with a config that has no seed markers and no registered worktree source. 4. Observe duplicate port selection, lost service adoption, or a seed-source startup error. **Paperclip version or commit** Current `master` plus the workspace runtime reliability changes in this pull request. **Deployment mode** Local development with managed execution workspaces and embedded Postgres. ## What Changed - Added a shared port registry with lease heartbeats, process identity checks, and live listener probes. - Reserved worktree ports across custom parent paths and repaired duplicate legacy assignments. - Preserved and adopted healthy managed services across control-plane restarts. - Reconciled guest bind modes and verified listener ownership before termination or reuse. - Provisioned registered markerless worktree databases and kept normal markerless instance startup as a no-op. - Added CLI, shared, server, and shell regression tests for seed, port, listener, restart, and adoption behavior. - Updated the worktree development documentation. ## Verification - `pnpm exec vitest run cli/src/__tests__/worktree.test.ts --reporter=verbose` — 63 tests passed. - `pnpm exec vitest run packages/shared/src/worktree-port-registry.test.ts --reporter=verbose` — 5 tests passed. - Focused runtime Vitest set — 199 tests passed across 37 suites. - `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs` — 10 tests passed. - `git diff --check` passed. ## Risks - Port reservation now depends on lease and process identity data. The fallback listener probe prevents early reclamation when process metadata is incomplete. - Runtime adoption is stricter about bind and owner identity. The tests cover healthy adoption, stale records, PID reuse, and unrelated listeners. - Markerless seed detection now separates registered worktrees from normal instances. The tests cover both paths. - There are no database schema migrations. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with the `gpt-5` model family. The serving snapshot and context-window size are not exposed. The agent used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
233be4b36c |
feat: parallelize sandbox file-sync behind a provider opt-in capability (#11736)
## Thinking Path > - Paperclip runs AI agents through local and remote execution adapters. > - Sandbox providers move workspace and asset files before and after agent runs. > - Serial file transfers delay startup and teardown when several operations do not depend on each other. > - Providers need an opt-in contract so existing providers keep their serial behavior. > - This pull request adds a bounded scheduler and routes inbound and outbound sync operations through it. > - The benefit is shorter sandbox setup and teardown with stable errors, clear telemetry, and a safe opt-in path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above): packages/shared, packages/adapter-utils, packages/plugins, and server. **Problem or motivation** Sandbox sync processes the workspace, assets, and referenced projects in series. This adds avoidable wait time to agent startup and teardown. **Proposed solution** Add a fail-closed provider capability named concurrentSyncOperations. Use a bounded scheduler with a limit of four operations. Preserve operation order for error reporting. Keep non-opted-in providers on the serial path. **Alternatives considered** Increase the serial transfer speed or add provider-specific schedulers. Those options do not provide one shared contract or stable behavior across providers. **Roadmap alignment** ROADMAP.md lists cloud and sandbox agents as a product area. This change improves sandbox execution without changing the control-plane contract. **Additional context** The Daytona provider opts in. Board trials on this commit showed overlap for inbound sync and outbound restore, with no referenced-project staging failures. ## What Changed - Add the concurrentSyncOperations sandbox capability and fail-closed parsing. - Add a bounded settle-all scheduler with stable input-order errors. - Parallelize inbound workspace, asset, and referenced-project sync operations when the provider opts in. - Parallelize outbound workspace and asset restore operations when the provider opts in. - Surface referenced-project failure text in run logs and server telemetry. - Add Daytona sync spans and the capability declaration. - Preserve in-flight upload scratch tarballs during workspace wipe. - Add unit and regression tests for the scheduler, coordinators, provider behavior, telemetry, and wipe race. ## Verification - Run the adapter-utils and server type checks. - Run the targeted adapter-utils, server, and Daytona test suites. - Run the full automated sweep. - Review six cold Daytona trials, with three serial and three parallel runs. - Confirm that parallel trials show inbound overlap and outbound restore overlap. - Confirm that providers without the capability keep serial behavior. ## Risks - Providers must opt in only when their file operations can run safely at the same time. - A provider that declares the capability incorrectly can expose transfer races. - The scheduler keeps a limit of four to bound resource use. - Providers without the capability keep the prior serial behavior. ## Model Used OpenAI GPT-5 in the Codex runtime. The model used tool calls, code inspection, and GitHub workflow support. The model did not author the implementation commits. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e0b64529b3 |
feat(auth): normalize agent login in the sandbox onto one session table and a capability contract (#11730)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox agents need a safe login path for each supported adapter > - Codex device login and Claude setup-token login used separate session stores and route logic > - Separate stores made session lookup, expiry, and login capability checks harder to keep consistent > - This pull request unifies both flows on one session table and one capability contract > - The benefit is one company-scoped login model with public session identifiers and shared lifecycle rules ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Codex and Claude sandbox login used separate session stores and different route paths. This split increased the risk of inconsistent company scoping, session lookup, and cleanup. **Proposed solution** Use `adapter_auth_sessions` for both login flows. Use public session identifiers for API access. Select login behavior from projected adapter capability data. Share the route spine, lease arguments, runner lifecycle, and reaper rules. **Alternatives considered** Keep two session tables and add matching fixes to both routes. This keeps duplicate logic and does not provide one capability contract, so this pull request uses shared infrastructure. **Roadmap alignment** This change supports the shipped Cloud / Sandbox agents milestone in `ROADMAP.md`. ## What Changed - Unify Codex device login and Claude setup-token login on `adapter_auth_sessions`. - Return and look up sessions with company-scoped public session identifiers. - Enforce one active session for each company, owner, and adapter. - Share the login route spine, sandbox lease arguments, runner lifecycle, and missing-auth check. - Add a standalone setup-token reaper with adapter-specific row selection. - Add optional login capability projection for adapters and drive route and UI selection from that data. - Rename the provider flag to `supportsLoginPty` and validate its deprecated alias. - Remove the old Claude setup-token session table and add the required migrations. ## Verification - Server typecheck passed with `tsc`. - Database typecheck passed. - UI typecheck passed with `tsc -b`. - Codex login service and route suites passed. - Setup-token session, route, and reaper suites passed. - Adapter session schema, plugin validator, capability projection, UI render, and Daytona suites passed. - GitHub Actions must confirm the complete CI gate after pull request creation. ## Risks - The migrations remove short-lived in-flight login rows during deployment. A login that spans the migration can continue until its provider lease expires. - The Codex credential store remains company-scoped. A cross-owner credential race remains a documented, board-accepted risk. - API clients that use internal session row identifiers no longer work. The API accepts only public session identifiers. ## Model Used Codex, GPT-5, exact runtime model ID not exposed in this handoff, large context window, reasoning, and repository tool use. The implementing engineer produced the code with AI assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e1df4c6068 |
fix(workspaces): keep deferred seed databases reliable (#11706)
## Thinking Path > - Paperclip manages agent work in isolated execution workspaces. > - A workspace depends on a valid database seed before it can run. > - Deferred seed failures were hidden behind a successful provision status. > - The seed restore also had two possible owners for the embedded PostgreSQL process. > - That allowed the target database to stop while the restore was still running. > - This pull request makes seed failures visible and gives the seed process sole lifecycle ownership. > - The benefit is that workspace provisioning reports the real result and does not stop its own target database. ## Linked Issues or Issue Description Related: #11684 **What happened?** Initial worktree provisioning could report success before its deferred database seed completed. The seed restore could also reuse a target embedded PostgreSQL process with another shutdown owner. This could stop the target database during the restore. **Expected behavior** Workspace status must show a failed deferred seed as a failure. The seed restore must own the target embedded PostgreSQL process until restore, migration, and validation finish. **Steps to reproduce** 1. Provision a worktree with deferred database seeding. 2. Make the seed manifest end in a failed state while the command exits with code 0. 3. Observe that the provision status remains successful on `master`. 4. Start a seed restore against an already-running target embedded PostgreSQL process. 5. Observe that another lifecycle owner can stop the target during restore. **Paperclip version or commit** `51a843e135` **Deployment mode** Local dev with execution workspaces and embedded PostgreSQL. ## What Changed - Add a first-class `workspace_seed` operation for deferred database seeds. - Require terminal, verified seed evidence before the seed operation succeeds. - Surface the seed phase and failure metadata in workspace status and UI state. - Give the seed process exclusive lifecycle ownership of the target embedded PostgreSQL process. - Suppress imported embedded-Postgres exit hooks without removing existing host listeners. - Record a credential-safe shutdown diagnostic in failed seed manifests. ## Verification - The original deferred-seed commit passed 4 server tests, 24 workspace-status UI tests, shared/server/UI typechecks, and the UI token gate. - The original PostgreSQL-lifecycle commit passed 3 lifecycle tests, 3 ownership/diagnostic tests, 1 real embedded-Postgres seed integration, and the affected package typechecks. - No local tests were rerun after the clean cherry-pick because the operator requested the shortest landing path. - Review the automatic PR checks for the clean `origin/master` replay. ## Risks - A live target database now causes an early error instead of being reused. The error includes recovery guidance. - Workspace consumers must handle the new `workspace_seed` operation type. Shared types and UI state handling are updated in this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, high-reasoning mode, with repository, shell, and GitHub tool use. The runtime does not expose a more specific deployment suffix or context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a2bf936f9a |
feat(workspaces): sign the workspace login handoff and gate readiness (#11671)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed worktree services run isolated Paperclip instances with cloned databases. > - A reachable service was reported as ready even when its database, runtime identity, or login path was not usable. > - The first candidate added verified database seeding and managed repair in #11665. > - This pull request consolidates that candidate with signed login handoff and a complete readiness contract. > - Post-QA fixes close five defects in repair identity, repair responses, UI retry, seed journal handling, and seed-source trust. > - The benefit is a workspace that either opens safely or reports one accurate recovery action. ## Linked Issues or Issue Description No public GitHub issue exists for this work, so the problem is described here. **What happened** Managed workspace URLs could return HTTP 200 and report ready while login failed. QA also found cases where repair used the wrong instance identity, returned a generic error, left the UI stuck, rejected a safe journal lag, or trusted a mutable workspace manifest. **Expected behavior** Opening a ready workspace signs the board user in to the correct isolated instance. Provisioning and repair use a registered source and report a structured recovery state. **Actual behavior** Entry depended on a password copied into the clone. Several failure paths could publish stale readiness, hide the repair precondition, or trust state that the workspace could modify. **Additional context** This pull request includes the commits first published in #11665. That pull request keeps the original base head for review history. This consolidated pull request is the merge candidate. Related open readiness work includes #11575 and #11621. ## What Changed - Adds a short-lived, signed, single-use login ticket. It binds the user, workspace, instance, and runtime origin. - Exchanges the ticket through Better Auth. It creates the session and cookie through the supported adapter path. - Adds protected workspace readiness fields for the database, clone data, login handoff, seed phase, and runtime identity. - Fails readiness closed when the guest has no company or execution-workspace binding. - Binds ticket issuance to the exact cloned user and active company membership selected for the handoff. - Verifies every current active board identity through the exact-user handoff before publication or reuse. - Gates managed runtime publication on the readiness contract and the recorded worktree instance identity. - Refreshes runtime work products from the live runtime row after a port change. - Adds one workspace access card with ready, degraded, repairing, and failed states. - Uses the runtime response identity for repair. It returns structured repair precondition errors. - Lets a valid source journal lag converge during provisioning. - Binds seed and repair manifests to a source registered outside the agent-writable worktree. - Clears recovered UI errors so a successful retry can open the workspace. - Makes runtime tests register canonical sources and avoid ports owned by live host listeners. - Keeps Vitest on source suites when compiled `dist` trees exist. - Isolates CLI and adapter tests from ambient AWS and runtime API environment variables. - Preserves a 404 response for cross-company workspace ID lookups before runtime authorization. - Makes concurrent single-flight coverage independent of path-canonicalization scheduling order. ## Verification The following checks passed on the integrated head: ```sh pnpm -r typecheck pnpm build pnpm check:token-gates pnpm --filter @paperclipai/db check:migrations ``` - The server source lane passed 420 files and 4,953 tests. Five tests were skipped. - The CLI lane passed 57 files and 385 tests. - The database lane passed 26 files and 97 tests. - The shared package passed 58 files and 506 tests. - The adapter utility lane passed 640 tests. Four tests were skipped. - The Claude adapter passed 220 tests. One test was skipped. - The Codex adapter passed 323 tests. - The OpenClaw adapter passed 13 tests. - The OpenCode adapter passed 42 tests. - The plugin SDK passed 45 tests. - The workspace runtime suite passed 124 tests. - The caller-scoped readiness and handoff suite passed 52 tests. - The workspace provisioning shell suite passed 7 tests. - The runtime exposure suite passed 17 tests while live host mappings occupied fixed test ports. - `git diff --check` passed and the worktree is clean. The serialized route lane will run in GitHub CI with its normal shards. No deployment or active-workspace migration was performed. ## Risks - This is a medium-risk authentication and runtime-readiness change. - The login ticket uses exact origin, workspace, instance, and user binding. It has a short expiry and a one-time nonce. - Runtime publication is stricter. A real readiness, identity, per-user handoff, or control-plane database disagreement now blocks publication. - This pull request supersedes #11665 as the merge candidate. Close #11665 after this pull request merges. - No new database migration is included. The lockfile and workflow files are unchanged. - Deployment and active-workspace migration are intentionally outside this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Opus 5 (`claude-opus-5[1m]`), 1M context, extended thinking, tool use, and code execution produced the main candidate. OpenAI GPT-5 (`gpt-5`) through Codex, with agentic reasoning, tool use, and code execution, integrated the post-QA fixes and hardened the test gates. The Codex context-window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
120ae5428f |
feat(server): add a one-click relink action for detached custom-image templates (#11641)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip environments can use captured custom images for agent runs > - A configuration fingerprint change can detach a valid custom-image template > - Operators need a safe way to confirm that the image still matches the boot source > - This pull request adds a guarded relink action with drift classification and audit logging > - The benefit is a deliberate relink without a new sandbox boot or provider snapshot ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting environment, server, and UI behavior. **Problem or motivation** A custom-image template detaches when the environment configuration fingerprint changes. The runtime then uses the base image, even when the boot source did not change. The only prior remedy required a full re-capture. **Proposed solution** Add an operator-triggered relink action. Classify configuration drift from a server-owned boot-relevant snapshot. Relink knob-only drift without confirmation. Require explicit confirmation for boot-source or unclassified drift. Guard the route for instance administrators and record a safe activity event. **Alternatives considered** Keep requiring a full re-capture. This adds a sandbox boot and provider snapshot for cases where the image remains correct. **Roadmap alignment** The roadmap has no matching custom-image relink item. This change addresses an environment operation gap. **Additional context** The relink response exposes raw drift values only in the transient 409 response to the instance administrator. The service never persists or logs fingerprints or configuration values. Reserved identity-path segments fail closed. ## What Changed - Add `relinkActiveTemplate` with drift classification and conditional fingerprint update. - Persist a server-owned boot-relevant configuration snapshot during capture. - Add the guarded relink route with strict request validation and activity logging. - Add the relink action and confirmation flow to the environment page. - Add service, route, UI, and OpenAPI coverage. ## Verification - Run the focused service suite: `pnpm vitest run server/src/services/environment-custom-images-service.test.ts`. - Run the focused route suite: `pnpm vitest run server/src/routes/environment-custom-image-routes.test.ts`. - Run the focused UI suite: `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx`. - Run server and UI TypeScript checks. - Confirm the OpenAPI snapshot matches the new route. - Confirm all required GitHub checks pass on commit `e46fdcfe94a719be854adf8849d30714e5b70b93`. - Confirm Greptile reports 5/5 with no unresolved review threads. ## Risks The relink action can keep an image after configuration drift. The service requires explicit confirmation for boot-source or unclassified drift. Reserved path segments produce a safe unresolved marker and never enter stored values. ## Model Used OpenAI GPT-5 Codex. The model used repository inspection, GitHub operations, and PR preparation with tool use and code execution. The runtime did not expose a context-window value or a separate reasoning-mode value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6b8e42168e |
Add governed secret alias confirmation cards (#11486)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need scoped secret bindings to use external services safely. > - Agents could not request an existing secret under a new config name without an internal secret identifier. > - Existing binding proposals were only visible in Settings and did not create an issue-thread approval path. > - A confirmation card could record acceptance without proving that the binding was created. > - This pull request extends the existing secret proposal system with safe source references and governed issue-thread confirmation cards. > - The benefit is a one-click flow that creates the binding or shows a clear failure without exposing secret material. ## Linked Issues or Issue Description Related prerequisite: #11482. **Subsystem affected** Cross-cutting: server REST APIs, shared interaction contracts, database proposal schema, and issue-thread UI. **Problem or motivation** An agent can need an existing bound secret under a second config name. The agent cannot safely discover the internal secret identifier. The existing proposal is also easy for the operator to miss because it only appears in Settings. A generic confirmation can record acceptance without executing the binding. **Proposed solution** Let an agent create a binding proposal from one of its existing config paths. Mint a server-owned, human-only confirmation card on the checked-out issue. Recheck the operator's target-agent permission under the proposal row lock. Execute the existing proposal transaction after card acceptance. Store an `executed` or `failed` result on the card. Render the complete lifecycle in the issue thread and attention resolver. **Alternatives considered** A new alias subsystem would duplicate proposal quotas, expiry, authorization, and binding synchronization. A text-only issue comment would not provide a governed action or an execution result. An agent-supplied card payload would permit metadata smuggling. This change uses the existing proposal transaction and a server-owned payload instead. **Roadmap alignment** This change extends the completed "Secrets Manager with per-agent access" roadmap item. It preserves scoped bindings and audited resolution. The required GitHub search found no other open duplicate issue or pull request. ## What Changed - Added safe source-config-path binding proposals and preserved user-secret ownership checks. - Added a proposal-to-interaction link and an idempotent database migration. - Minted human-only `request_confirmation` cards with server-owned `secretProposal` metadata. - Rejected agent-supplied governed metadata and agent addressees. - Rechecked `agent_config:update` authority under the proposal lock before execution. - Recorded `executed` or `failed` results and posted a failure comment when no binding was created. - Settled failed accepted proposals atomically and mirrored rejection, withdrawal, and expiry in both directions. - Emitted `secret.binding.created` for new agent binding writes. - Added a dedicated issue-thread card for pending, executed, failed, rejected, withdrawn, and expired states. - Showed only the source label, target agent, config path, skeptical justification, expiry, and safe failure code. - Replaced resolved attention-query entries immediately with the stitched server result. - Added focused server, database, UI, and state-transition tests. - Added Storybook fixtures for every review state and documented the API and agent behavior. ## Verification - `pnpm exec vitest run ui/src/components/IssueThreadInteractionCard.test.tsx ui/src/components/AttentionInteractionResolver.test.ts` — 58 passed. - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm build-storybook` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/db check:migrations` - `NODE_ENV=test pnpm exec vitest run server/src/__tests__/issue-thread-interaction-routes.test.ts server/src/__tests__/secret-proposals-routes.test.ts server/src/__tests__/secrets-routes.test.ts server/src/__tests__/agents-service-secret-bindings.test.ts` — 142 passed. - `NODE_ENV=test pnpm --filter @paperclipai/db exec vitest run src/company-secret-proposals-migration.test.ts --silent` — 1 passed. - `pnpm -r typecheck` - `pnpm test:run` — server 4,175 passed, UI 4,109 passed; the CLI AWS-doctor case passes 8/8 with runtime-injected static AWS credential variables unset. - `pnpm build` - `git diff --check origin/master...HEAD` ## Risks - Migration `0221` adds one nullable foreign key and one index. It uses idempotent guards. - The accept route performs a governed write after it records card acceptance. A failed write is visible and settles the proposal as rejected. - Concurrent proposal and card resolution must use proposal-before-interaction lock order. A race test covers direct approval against card rejection. - The new audit event increases activity rows for newly added agent bindings. It does not include secret values or fingerprints. - The card includes only safe proposal metadata. It does not include secret value, fingerprint, version, or internal secret identifiers. - The UI uses the stitched resolution result. Focused tests cover immediate cache replacement and every terminal state. > This work extends an existing completed roadmap capability. The GitHub duplicate search returned no other open related work. ## Model Used - OpenAI Codex with model ID `gpt-5`. The runtime did not expose its context-window size. Reasoning, repository tools, code execution, database integration tests, UI rendering, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b446ff59bf |
refactor(acpx-engine): coordinator-owned ACP run lifecycle with a typed resource ledger (#11576)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters run agent sessions through the ACPX engine > - The ACPX engine handled one run attempt as a long implicit procedure > - That shape made resource ownership, cleanup order, and failure behavior hard to verify > - This pull request gives the attempt a coordinator, a typed resource ledger, separate run sites, and explicit turn and settlement sequences > - The benefit is clear ownership, one cleanup path, safer session reuse, and testable failure behavior ## Linked Issues or Issue Description **What existing behavior does this improve?** The ACPX engine manages startup, turn execution, session reuse, and cleanup inside one large run procedure. **Current behavior** The run procedure owns several resources through implicit control flow. Cleanup and session reuse behavior depend on lane-specific branches and error paths. **Proposed behavior** The coordinator owns the run attempt. A typed ledger records six resources and their states. Host and sandbox run sites own lane-specific acquisition. Turn and settlement sequences expose typed outcomes. The engine emits allowlisted phase telemetry. **Reason and benefit** Explicit ownership makes cleanup and failure behavior easier to inspect. The fault matrix and characterization tests protect the external result while the refactor reduces hidden control flow. **Breaking changes** None to the public adapter contract. The host warm-save path now closes and relaunches the runtime because a transferred runtime could retain a run-scoped credential. A cold session-handshake failure now closes the created runtime. **Additional context** This pull request contains the ACPX engine lifecycle refactor, its tests, and the lifecycle document. ## What Changed - Add a run coordinator for startup, turn execution, settlement, and result reproduction. - Add a typed resource ledger with open, sealed, and consumed states. - Add host and sandbox run sites for lane-specific resource acquisition. - Replace separate runtime maps with a generic session reuse store. - Split session fingerprint identity from the outer session key. - Add typed turn and settlement sequences with one cleanup owner. - Add a closed allowlist for phase telemetry. - Add characterization tests and a 17-case fault matrix. - Add `doc/acp-run-lifecycle.md`. ## Verification - `npx vitest run packages/adapter-utils/src/acpx-engine/` passes 18 files and 286 tests at the submitted commit. - `pnpm --filter @paperclipai/adapter-utils typecheck` reports 0 errors at the submitted commit. - Run the full pull request checks after GitHub starts CI. - Run Greptile review after the pull request opens. ## Risks - The refactor changes internal control flow across the ACPX engine. - Host warm-save behavior now closes and relaunches the runtime. - Settlement changes the handling of a cold session-handshake failure from a leak to a close. - The characterization baselines and fault matrix reduce the risk of an external behavior change. > Paperclip is the open source app people use to manage AI agents for work > The adapter layer runs agent sessions through the ACPX engine > The engine needs explicit lifecycle ownership for reliable cleanup > This pull request adds coordinator-owned phases and a typed resource ledger > The result makes lifecycle behavior easier to test and review ## Model Used OpenAI GPT-5 Codex. Exact model ID: GPT-5. The model used tool execution, repository inspection, and code review support. The implementation author supplied the submitted code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
14fd8aee36 |
fix(server): flag truncated issue descriptions (#4771)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The issue list API is one of the surfaces API consumers use to synchronize issue metadata. > - The list endpoint intentionally returns a bounded `description` preview so large descriptions do not bloat list responses. > - Before this change, that preview looked like a complete field value because the response did not say whether it had been shortened. > - That made round-trip clients vulnerable to accidentally PATCHing a preview back over the full description. > - This pull request keeps the existing preview behavior but adds an explicit `descriptionTruncated` flag. > - The benefit is backwards-compatible visibility into truncated issue descriptions, so clients can avoid data-loss workflows. ## Linked Issues or Issue Description Fixes #4758. Related PR: #4792 also targets #4758, but it includes unrelated logger changes and currently has separate review/security concerns. This PR keeps the fix scoped to the issue-list description truncation API behavior. ## What Changed - Added `descriptionTruncated` to the issue list projection when `description` exceeds the existing 1200-character preview limit. - Exposed `descriptionTruncated?: boolean` on the shared `Issue` type. - Added service tests for truncated descriptions, exact-limit descriptions, null descriptions, and multibyte-safe preview truncation. ## Verification June 18, 2026 refresh after rebasing onto current `origin/master`: - `pnpm install --frozen-lockfile` - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `pnpm typecheck` - `git diff --check origin/master...HEAD` - GitHub PR checks are green on head `12e828e6`. Earlier pre-review verification also included `pnpm test`. ## Risks - Low risk. This is an additive API response field; existing clients can ignore it. - The list endpoint still returns the same bounded `description` preview. Clients that need full text should continue fetching the issue detail, but can now detect when that is necessary. - No database migration or UI behavior change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5, via Codex desktop on April 29, June 15, and June 18, 2026. Used tool-assisted repository inspection, code editing, local test execution, GitHub CLI workflows, and PR review follow-up. Exact context window size is not surfaced by the tool. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots (N/A: no UI change) - [x] I have updated relevant documentation to reflect my changes (N/A: additive API field covered by tests) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Sami Rusani <sr@samirusani> |
||
|
|
c1c46f1e4e |
feat: Claude login on the new-agent page before agent creation (#11347)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Claude local adapter supports subscription login through a sandbox > - The new-agent page must show login before the user creates an agent > - Test results must not expose raw sandbox diagnostics or secret values > - This pull request adds the login UI to both Test lanes and closes the diagnostic boundary > - The branch also adds durable cleanup recovery for failed sandbox teardown > - Reusable sandboxes must retain both their recorded teardown configuration and a valid lifecycle path until destruction succeeds > - The benefit is a usable login flow with fixed public checks, redacted server logs, and recoverable sandbox cleanup ## Linked Issues or Issue Description Related public work: [#9488](https://github.com/paperclipai/paperclip/pull/9488) adds first-class recognition for `CLAUDE_CODE_OAUTH_TOKEN` in headless and remote runs. Related public issue: [#2681](https://github.com/paperclipai/paperclip/issues/2681) requests Claude Code subscription support. This pull request adds the login transport and new-agent UI flow that those changes do not provide. **Subsystem affected:** Claude local adapter, server login probes, sandbox provider setup, cleanup recovery, and the new-agent UI. **Problem or motivation:** The Test lanes did not show the sandbox login panel in all supported cases. Test results also exposed raw probe diagnostics, and JSON escapes could end secret redaction early. **Proposed solution:** Surface the login capability through the bundled provider manifest. Prepare the same probe runtime in the ACP lane. Send diagnostics only to redacted server logs. Keep Test checks on fixed public messages. Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. Consume JSON escapes during redaction. Preserve failed sandbox cleanup state across retries and restarts, and prevent deletion from severing the lifecycle context of a live reusable sandbox. **Alternatives considered:** Keep raw diagnostics in Test checks or trust login URL text from the sandbox. Both choices increase information exposure. Keep separate probe behavior in the ACP lane. That choice would leave the two Test lanes inconsistent. ## What Changed - Surface the sandbox login panel on both Test lanes. - Reconcile the bundled Daytona plugin manifest so `supportsSetupTokenLogin` reaches the UI capability gate. - Prepare the ACP Test lane with the same probe runtime as the CLI Test lane. - Add the `claude_acp_login_probe_unavailable` warning when the ACP probe cannot run. - Send raw sandbox diagnostics only to redacted server logs. - Keep Test checks on fixed public messages in the ACP, managed-config, and CLI paths. - Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. - Redact JSON and escaped-JSON secret values, including escaped quotes and backslashes. - Preserve orphan cleanup records across provider failures, restarts, and unavailable plugins. - Atomically block environment deletion while a live reusable sandbox lease still depends on it. - Verify pending cleanup destroys plugin sandboxes with the provider configuration recorded on the lease, even after the current environment configuration changes. ## Verification - Head under review: `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241`. - Focused environment route/service/runtime coverage passes: 196 tests across 3 files. - `pnpm -r typecheck` passes. - `pnpm build` passes. - The full Vitest run completed with 4,754 passing and 28 failing tests. All 23 source-test failures reproduce unchanged on parent head `58cfe61a33191ce03d965d65085d26064b4888ba`; the other 5 are duplicate executions from stale `server/dist` output. The failures are unrelated macOS path/listener and scheduler-fixture failures, so there is no new bad commit for bisect to localize. - All required CI checks pass for the current head, including build, typecheck/release registry, all server and workspace shards, serialized server suites, canary, and e2e. - A fresh Greptile review for `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241` reports 5/5, “safe to merge,” with no blocking failure remaining. ## Risks - A probe or redaction change could hide useful server diagnostics. - An allowlist change could reject a valid Claude login URL. - Cleanup recovery changes could affect provider teardown ordering. - An environment with a live reusable sandbox can no longer be deleted until the owning issue or execution workspace completes teardown. - The implementation keeps public Test messages fixed and sends detail to redacted server logs. ## Model Used OpenAI GPT-5 via Codex — exact model ID: GPT-5; tool use and code execution enabled; extended reasoning enabled. The implementation author used AI-assisted development. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and documented the result - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation or confirmed no separate documentation change is needed - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3061ce6901 |
feat(sandbox): stream session output by capability, drop three operator flags (#11557)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandboxed agents use provider capabilities to select safe execution paths > - Session output still depends on three operator flags that duplicate capability data > - Duplicate flags can drift from the verified sandbox capability snapshot > - This pull request makes the capability snapshot the only streaming decision and removes the obsolete flags > - The benefit is default streaming with a poll fallback when a capability or stream fails ## Linked Issues or Issue Description **What existing behavior does this improve?** ACP sandbox session-output streaming and sandbox execution configuration. **Subsystem affected** Cross-cutting (multiple of the above): server/, packages/shared/, packages/adapter-utils/, and packages/plugins/. **Current behavior** Session-output streaming requires operator flags in the server and Daytona plugin configuration. Saved configurations can retain a removed key. **Proposed behavior** The verified capability snapshot selects streaming. The Daytona plugin uses persistent sessions by default, keeps bypass commands one-shot, and falls back from the log stream to polling. Removed configuration keys become inert. **Reason and benefit** One capability source prevents configuration drift. The fallback keeps output available when capability resolution or log streaming fails. **Breaking changes** The three operator flags no longer control session-output streaming. Existing saved keys load but have no effect. ## What Changed - Remove `useSessions` and `useLogStream` from the Daytona plugin configuration and manifest. - Remove `streamAgentSessionOutput` from server configuration, shared types, and execution-target plumbing. - Select streaming from `persistentProcessSessions` and `independentControlCommands`. - Keep poll fallback on capability resolution failure and stream failure. - Strip removed keys from strict fake-sandbox and catchall plugin configuration. - Update the sandbox capability documentation and focused tests. ## Verification - `tsc --noEmit` passed in `packages/shared`, `packages/adapter-utils`, `server`, and the Daytona plugin. - Daytona `plugin.test.ts` passed 139 tests. - Server capability, configuration, route, and runtime suites passed 160 tests. - `packages/adapter-utils` `execution-target-sandbox.test.ts` passed 44 tests. - The capability matrix covers stream, poll, and resolution-failure paths. - Removed-key tests cover strict fake-sandbox and catchall plugin schemas. ## Risks - A capability snapshot that lacks either required session capability uses polling. - A log stream failure uses polling and can increase request count. - Existing removed configuration keys no longer change behavior. - The isolated-worktree Daytona Vitest run has a pre-existing missing `packages/adapters/droid-local` reference. CI and standard checkouts use the committed configuration. ## Model Used OpenAI Codex, GPT-5, tool use and code review assistance. The exact runtime context window is managed by the Codex platform. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e71ce9a9d3 |
feat: sandbox provider capability contract with fail-closed effective resolution (#11463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs work through adapters and sandbox providers > - Providers need a clear contract so the server can use only verified capabilities > - A declared capability must not grant a method that the live worker did not verify > - This pull request adds manifest declarations and fail-closed effective capability resolution > - The benefit is safe provider reuse across execution targets and run lifecycles ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Sandbox providers expose different runtime methods. The server needs one safe capability contract that accounts for provider declarations, worker verification, and narrowing configuration. **Proposed solution** Add strict manifest validation for five sandbox capabilities. Resolve effective capabilities as the subset of verified, declared, and narrowed values. Store the result as a frozen execution-target snapshot. **Alternatives considered** Trusting the manifest alone could grant methods that the worker does not support. Trusting only a fixed built-in list would reject valid third-party providers. The intersection rule keeps the verified runtime ceiling and supports both provider types. **Roadmap alignment** This change supports the ACP run lifecycle track and the sandbox provider contract work in the current roadmap. **Additional context** The legacy `supportsReusableLeases` field remains supported. The nested capability validator rejects unknown keys. Missing or unavailable verification resolves all capabilities to `false`. ## What Changed - Add strict `sandboxCapabilities` manifest validation with legacy reusable-lease compatibility. - Carry declarations through the ready-driver projection. - Add fail-closed effective resolution from verified, declared, and narrowed capabilities. - Add narrowing for provider configuration, Kubernetes Job leases, and Daytona sessions. - Add a frozen read-only capability snapshot to execution targets. - Add focused tests and keep existing characterization baselines covered. - Add and update sandbox provider capability documentation. ## Verification - `npx vitest run packages/shared/src/validators/plugin.test.ts` - `npx vitest run server/src/__tests__/plugin-environment-driver-sandbox-capabilities.test.ts` - `npx vitest run server/src/__tests__/sandbox-capability-contract.test.ts` - `npx vitest run server/src/__tests__/environment-execution-target-capabilities.test.ts` - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-characterization.test.ts packages/adapter-utils/src/acpx-engine/turn-characterization.test.ts packages/adapter-utils/src/acpx-engine/settlement-characterization.test.ts packages/adapter-utils/src/acpx-engine/composed-run-characterization.test.ts` - Package typechecks for shared, server, and adapter-utils pass. - Stage-2 security review suites pass with 28 tests. ## Risks The resolver fails closed when verification is absent or unavailable. Providers that rely on undeclared capabilities may see narrower behavior until they expose verified worker methods. The change does not alter the existing native-sync guard. ## Model Used OpenAI Codex, GPT-5, exact runtime model ID `gpt-5`, tool use and code execution. The implementation author used this model to assist with the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4802b1bbc |
feat(runtime-exposure): least-privilege Tailscale HTTPS broker, shared contract, and persisted exposure state (#11524)
<!-- Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts and supervises managed runtime services for a project's execution workspaces, so an agent's branch can be previewed while it works > - Those services only listen on plain loopback HTTP. A person on another device, or on a phone, cannot open the preview > - A Tailscale HTTPS mapping solves this, but `tailscale serve` needs host privileges that the Paperclip server process must not hold > - This pull request adds the foundation only: a separate least-privilege host broker, the shared exposure contract, and the database columns that hold exposure state > - Nothing calls the broker yet, so there is no behavior change. The benefit is that the privileged surface is small, reviewable, and isolated before any lifecycle code depends on it ## Linked Issues or Issue Description No public GitHub issue exists. The change follows the feature request template. **Subsystem affected** Managed workspace runtime services, the shared type and validator package, and the database schema. **Problem or motivation** A managed runtime service binds to loopback only. There is no supported way to reach that preview from another device. Adding HTTPS directly to the server would mean the server process runs `tailscale serve`, which needs privileges far wider than the task requires. A compromised or buggy server could then map any port to the tailnet. **Proposed solution** Split the privileged work into a separate broker process with a narrow protocol, and define one shared contract that the server, the UI, the runtime, and the broker all read. Land this foundation first, with no caller, so the privileged code can be reviewed on its own. **Alternatives considered** - Call `tailscale serve` from the server process. This was rejected because it gives the server unrestricted mapping authority. - Use `sudo` for single `tailscale` commands. This was rejected because the argument list is the only guard, and it is easy to widen by accident. - Use a generic reverse proxy. This was rejected because it does not remove the need for a privileged Tailscale mapping step. **Roadmap alignment** This supports the existing managed workspace runtime capability. It adds no new product surface on its own. **Additional context** The broker is the security boundary of the feature, so it is deliberately the first slice. Three later pull requests build on it: the server exposure lifecycle, the runtime lease and recovery integration, and the leased-port mediator. ## What Changed - Add the `@paperclipai/tailscale-https-broker` workspace package. The broker listens on a unix socket, authorizes each peer with `SO_PEERCRED`, and answers a small request protocol. - Restrict what the broker will map. It accepts only same-number HTTPS-to-loopback pairs inside the Paperclip port range, refuses protected ports, and confirms that the loopback port belongs to a Paperclip-owned listener. - Parse every request with a strict JSON reader that rejects duplicate keys, prototype keys, and unknown fields. - Write an append-only audit record for each broker decision. - Add the shared exposure contract in `@paperclipai/shared`: the `RuntimeExposureConfig`, `RuntimeExposureState`, and `RuntimeExposureStatus` types, their zod validators, the app and HMR port rules, and the loopback-bind helpers. - Persist exposure state on `workspace_runtime_services` with the new `exposure` column, plus the server-private `exposure_handle` and `backend_url` columns that are never serialized to API clients. - Add the `execution_workspace_runtime_leases` table that the later lease slice uses. - Extend the runtime read-model test fixture for the three new columns. ## Verification Focused checks, all run on this branch: - `pnpm --filter @paperclipai/tailscale-https-broker test` — 12 files, 82 tests pass. This covers peer credentials, port policy, protected ports, the serve config writer, the strict JSON reader, argv parsing, and the socket server. - `pnpm --filter @paperclipai/tailscale-https-broker typecheck` — clean. - `npx vitest run --root packages/shared src/runtime-exposure src/validators/runtime-exposure.test.ts` — 3 files, 40 tests pass. - `pnpm --filter @paperclipai/db typecheck` — runs `check:migrations` first. Migration numbering and migration safety both pass. - `pnpm --filter @paperclipai/shared typecheck` — clean. - `pnpm --filter @paperclipai/ui typecheck` — clean. - `npx vitest run --root server src/services/workspace-runtime-read-model.test.ts` — 3 tests pass. - `npx tsc --noEmit -p server/tsconfig.json` — 139 errors, which is exactly the count on `master` before this branch. All 139 come from the unbuilt `@paperclipai/plugin-sdk` package. To confirm the exposure state is inert, start a managed runtime service as usual. The new columns stay null and the service behaves as it does today. ## Risks - Migration risk is low. Both migrations only add a table and three nullable columns. No column is backfilled and no existing column changes. The migration safety check passes. - Behavior risk is low. No code path calls the broker in this pull request, and the shared exposure fields are optional. - The broker is privileged, so it is the real risk surface. It is mitigated by peer-credential authorization, a fixed port range, a protected-port deny list, same-number pair enforcement, listener-ownership checks, strict JSON parsing, and an audit trail. Reviewers should read `packages/tailscale-https-broker/src/authorization.ts` and `src/port-policy.ts` closely. - The broker requires a `tailscale` version floor, which its README records. An older host CLI makes the broker refuse to start rather than map incorrectly. - `pnpm-lock.yaml` changes because a new workspace package is added. The diff is the new importer block, plus one duplicate `tinyexec` entry that pnpm removed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. ## Model Used Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
10d0555189 |
fix(interactions): authorize resolvers consistently (#11376)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue interactions give agents and people a structured decision record. > - Resolver routes used different authorization rules. > - Some routes blocked valid agents, including task watchdogs with normal issue access. > - The API did not show who could resolve a pending interaction. > - This pull request gives every interaction kind one resolver policy evaluator. > - The benefit is a clear decision path with consistent governance and company isolation. ## Linked Issues or Issue Description Fixes: #8087 Refs: #7403 Related PR: #11082 proposes board-only confirmation rules. This change keeps human-only review as an explicit policy. **What happened?** Agents could create issue interactions. Some resolver routes still required board access. This left valid agent confirmations pending. Task watchdogs could see the same problem without board identity. **Expected behavior** Every interaction kind must use one resolver policy contract. The contract must support `anyone`, `not_creator`, and `human_only`. It must also apply all normal governance controls. **Steps to reproduce** 1. Create a `request_confirmation` interaction as an agent. 2. Resolve it with another authorized agent. 3. Observe the board-only denial. **Paperclip version or commit** The problem exists on `master` before this change. **Deployment mode** Local development with `pnpm dev`. ## What Changed - Add canonical policies for `anyone`, `not_creator`, and `human_only`. - Use one server evaluator for every interaction kind. - Apply named addressees, company limits, review rules, and task watchdog scope. - Charge cross-issue resolutions to the existing per-run action limit. - Return the effective resolver audience in attention and interaction data. - Show the audience, governance choices, and denial reasons in the board UI. - Add telemetry, API documents, product documents, and regression fixtures. - Add migration provenance for safe legacy behavior. - Make migration `0218` safe for complete replays and partial prior runs. ## Product Rules - An interaction records a response. It does not grant authority for the next action. - `anyone` lets any authorized issue participant respond. - `not_creator` requires a responder other than the interaction creator. - `human_only` requires an authorized person. - A named addressee, company policy, or governed action can narrow the audience. - These controls cannot widen the audience. - A task watchdog uses the same rules as an ordinary agent. - A task watchdog does not receive board authority. - An agent resolution on another issue uses the shared cross-issue action limit. - Legacy pending interactions keep their earlier restrictions. - The UI shows the effective audience and a permanent denial reason. ## Verification - `pnpm --filter @paperclipai/db check:migrations` - `pnpm --filter @paperclipai/db typecheck` - `pnpm exec vitest run packages/db/src/issue-thread-interaction-resolver-policy-migration.test.ts` - The focused PostgreSQL test applies migration `0218` twice. - The test also completes a partial prior run and preserves existing provenance. - The latest GitHub head has 29 successful checks. - The opt-in Storybook visual check skipped as expected. - Greptile reports 5/5 with no open comments. ## Risks - New interaction writes use `anyone` by default. - Callers must select `not_creator` or `human_only` when they need stricter review. - Legacy pending interactions keep the old creator and human restrictions. - Migration `0218` fills only missing provenance fields during recovery. - Cross-issue resolutions can reach the existing action limit. - The shared evaluator affects every interaction kind. - Route, service, database, shared contract, and UI tests cover these rules. > This work matches the Agent Reviews and Approvals direction in `ROADMAP.md`. It does not duplicate a planned item. ## Model Used OpenAI Codex, GPT-5. The runtime does not expose the exact deployment ID or context window. The agent used reasoning, repository tools, shell commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked public issues or described the issue with the required labels - [x] I have not referenced internal Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented the risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open comments - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9e9f744f58 |
Show blocker links in the task chat (#11456)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip helps operators supervise agent work through tasks and task threads. > - The redesigned task thread shows the current work and its state. > - A blocked task did not show the dependency that prevented progress. > - Operators had to leave the thread to find the direct and final blockers. > - This pull request adds compact blocker links at the top and bottom of the task thread. > - The benefit is that operators can identify and open the relevant tasks without adding a large notice to the thread. ## Linked Issues or Issue Description **What existing behavior does this improve?** The redesigned task thread did not show which task directly blocked the current task or which task ultimately blocked its dependency chain. **Subsystem affected** `server/`, `packages/shared/`, and `ui/` task-blocker presentation. **Current behavior** A blocked task can open in the redesigned thread without a visible dependency link at the top or bottom of the conversation. **Proposed behavior** Show one compact amber row for the direct blocker. Show a second row for the selected final blocker when one exists. Render the rows at both ends of the thread. **Reason and benefit** Operators can see the reason for the blocked state and open the relevant task from the conversation. The compact rows preserve thread density. **Breaking changes** None. The new blocker-attention fields are optional. Existing clients remain compatible. ## What Changed - Added a compact task-chat component for direct and selected final blocker links. - Added the blocker rows to the top and bottom of populated and empty task threads. - Added link-ready blocker-attention details so an intermediate selected task stays on its correct direct chain. - Included blocker-link changes in the thread content key so pinned threads follow a newly added bottom row. - Added component, scrolling, server contract, and Storybook coverage for the new states. ## Verification - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx server/src/__tests__/issue-blocker-attention.test.ts` (38 tests passed) - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` ## Risks - Low risk. The rows only render while the task status is `blocked` and an unresolved blocker is available. - Long titles are truncated to keep each blocker on one line. The full task label remains available in the link title. - Older server payloads keep the original leaf-selection behavior because the new sampled details are optional. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The run used tool-enabled reasoning and code execution. The context-window size was not exposed to the run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e52b8a343f |
fix: ACP run lifecycle corrections — failure settlement, workspace sync-back, lease cleanup (#11454)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters run ACP sessions and manage runtime, workspace, and lease resources. > - Several failure paths left runtime bridges, staged workspaces, or environment leases active after an error. > - These leaks reduce run reliability and can leave later runs without clean resources. > - This pull request closes the failure paths, applies one teardown policy, and adds regression tests. > - The benefit is consistent failure settlement and safer reuse of agent workspaces and leases. ## Linked Issues or Issue Description **What happened?** ACP runs could leave runtime bridges, staged workspaces, or environment leases active after failures. Claude and Gemini ACP runs did not restore the sandbox workspace on teardown. Lease release stopped when one lease returned an error. **Expected behavior** Each ACP failure must return an error result and settle its resources. Teardown must run each step, release leases independently, and restore the host workspace when the sandbox ends. Pending cleanup leases must receive bounded retry attempts. **Steps to reproduce** 1. Run an ACP session that fails after runtime creation or during turn preparation. 2. Run an ACP session that fails during a warm hit or staged runtime handoff. 3. Run lease cleanup with more than one lease when the first release returns an error. 4. Inspect the result phase, teardown calls, workspace state, and lease metadata. 5. Run the regression suites listed in the Verification section. ## What Changed - Settle every ACP failure after runtime creation with an error result and one sandbox.startup span closure. - Close the ACP runtime and remove warm entries after every pre-turn failure. - Run all teardown steps, record teardown errors, release staging leases in finally, and prevent duplicate teardown. - Dispose staged runtimes after seam failures and remove borrowed staged entries with identity guards. - Add fail-open workspace sync-back teardown for Claude and Gemini ACP adapters. - Isolate lease release errors and add bounded retry sweeps for stranded pending_cleanup leases. - Atomically claim pending_cleanup retries and clamp attempt readers to keep the five-attempt bound. - Default absent provider reusableLeases values to false and align the fake provider with its runtime declaration. - Add regression tests for engine, adapter, server, and shared environment behavior. ## Verification - [x] `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts` — 124 tests passed. - [x] `npx vitest run packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/claude-local/src/server/acp.test.ts packages/adapters/gemini-local/src/server/acp.test.ts` — 61 tests passed. - [x] `npx vitest run server/src/__tests__/environment-runtime.test.ts server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts server/src/__tests__/reusable-leases-default.test.ts server/src/__tests__/environment-routes.test.ts packages/shared/src/environment-support.test.ts` — passed. - [x] All listed suites ran from the repository root. - [x] GitHub CI completed successfully for `cfc349c9f232711433897915112a1c52c0e462ca`. - [x] Greptile completed with a 5/5 confidence score and no blocking finding. ## Risks The engine changes affect failure settlement and teardown order across ACP runs. The server changes add retry state to existing lease metadata without a schema migration. The adapter changes restore workspaces after sandbox execution. Regression tests cover the changed paths. GitHub CI and Greptile passed for the current head. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This change fixes runtime reliability and does not duplicate a roadmap feature. ## Model Used OpenAI GPT-5 Codex. The model used tool-based repository inspection, GitHub operations, and code review support. The runtime does not expose a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (for example, `docs/...` or `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eabecc6f77 |
feat(annotations): include issue document annotations in agent review context (#11332)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Reviewers annotate plans and issue documents with inline comments, and assigned agents act on that feedback > - The server already builds a bounded review context from open plan annotations and includes it in agent wake payloads > - Non-plan issue documents did not get the same treatment: their open annotation threads never reached the agent, and the properties pane did not surface their annotations > - This pull request extends the review-context path and the properties-pane UI to issue documents, at parity with plans > - The benefit is that agent feedback on any issue document reaches the assigned agent, not only feedback on the plan ## Linked Issues or Issue Description **What existing behavior does this improve?** The review-context pipeline that delivers inline annotation feedback to assigned agents, and the properties pane that surfaces those annotations to reviewers. **Subsystem affected** The server review-context path (`server/src/services/plan-review-context.ts`, wake payload assembly in `server/src/services/heartbeat.ts`, `server/src/routes/issues.ts`), shared wake-payload types (`packages/shared`, `packages/adapter-utils`), and the issue properties pane (`ui/src/components/issue-properties/`). **Current behavior** A reviewer can annotate any issue document, not only the plan. The agent wake payload includes open annotation threads for the plan document only. Feedback left on other issue documents is invisible to the assigned agent. In the properties pane, the Artifacts tab also gives no way to see or open a document's annotations. **Proposed behavior** Add `buildDocumentReviewContext` beside the existing plan builder. It collects open annotation threads for all non-plan issue documents, applies the same thread, comment, and character budgets across documents, and reports truncation. Include the result as a new `documentReviewContext` field in agent wake payloads and in the issue wake-context route. Keep the plan context on its legacy builder and field so plan-only wakes stay byte-for-byte compatible. Render the new context in the adapter wake-payload text, and surface annotation counts and the annotation panel for documents in the properties pane's Plans and Artifacts tabs. **Reason and benefit** The floating annotation popover and persistent highlight UI landed earlier; this change completes the loop so agent feedback on any issue document reaches the assigned agent, not only feedback on the plan. **Breaking changes** None. The wake payload gains a new optional `documentReviewContext` field; the existing plan context field and its legacy builder are unchanged, so plan-only wakes stay byte-for-byte compatible. ## What Changed - Add `buildDocumentReviewContext` in `server/src/services/plan-review-context.ts`: bounded review context (shared thread/comment/character budgets, per-document legacy limits) over all non-plan issue documents - Include `documentReviewContext` in agent wake payloads (`server/src/services/heartbeat.ts`) and in the issue wake-context response (`server/src/routes/issues.ts`) - Add shared `DocumentReviewContext` / `DocumentReviewContextDocument` types in `packages/shared` - Normalize and render the new context in adapter wake-payload text (`packages/adapter-utils/src/server-utils.ts`), with tests - Show a `DocumentAnnotationsCountChip` and the annotation panel for documents in the properties pane Plans and Artifacts tabs, with tests - Extend server document-annotations service tests to cover the new context builder ## Verification - Run `npx vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/document-annotations-service.test.ts` from the repo root — 104 tests pass - Run `TZ=UTC npx vitest run ui/src/components/issue-properties/IssuePropertiesDocumentAnnotations.test.tsx ui/src/components/IssueProperties.test.tsx ui/src/components/IssueDocumentAnnotations.test.tsx ui/src/components/DocumentAnnotationPopover.test.tsx` from the repo root — 75 tests pass (one pre-existing monitor-row case asserts UTC timestamps, so use `TZ=UTC` locally; CI runs in UTC) - `pnpm run typecheck` in `server/` passes - Manual: annotate a non-plan issue document, then wake the assigned agent with a comment — the wake payload lists the open document annotation threads; the Artifacts tab shows the annotation count chip and opens the panel ## Risks - The wake payload gains a new optional `documentReviewContext` field; consumers that ignore unknown fields are unaffected, and the plan context field is unchanged - The context is new input to agent wakes; shared budgets (same limits as the plan context) bound token cost across all documents - Low UI risk: the properties-pane changes reuse the existing annotation components > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Claude Fable 5), with extended thinking and agentic tool use (Claude Code harness) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f0e6c0f549 |
feat(server): receive and apply the Paperclip Cloud onboarding seed (#11098)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud provisions a dedicated tenant stack for each
customer. During signup it asks for a mission, a name and role for the
first agent, and a first task.
> - Cloud pushes those answers into the new stack at activation, as
`POST /api/companies/:companyId/onboarding-seed`.
> - No route served that path. The tenant answered 404, so Cloud
recorded the push as unacknowledged and retried on every portfolio
fetch.
> - The failure was soft. The answers stayed durable in Cloud and the
stack still activated. But the stack opened on the empty first-run
wizard, and it asked the customer again for what they had already given.
> - This pull request adds the receiving endpoint. It validates the
seed, applies it, and acknowledges it.
> - The benefit is that a seeded stack opens with the mission, the agent
and the first task already in place.
## Linked Issues or Issue Description
No public GitHub issue covers this. The problem is described in-PR,
following the feature template.
**Subsystem affected**
server/ — Express REST API and orchestration services. Also
`packages/db` (one new table) and `packages/shared` (one new validator).
**Problem or motivation**
Paperclip Cloud collects onboarding answers at signup and pushes them to
the tenant stack at activation. The tenant had no route for that
request. It answered 404. Cloud treats a non-2xx as "not yet applied",
so it kept the answers and retried, but the stack itself stayed
unseeded. A customer who had already named their mission, their first
agent and their first task arrived at an empty first-run wizard that
asked for all three again.
**Proposed solution**
Serve `POST /api/companies/:companyId/onboarding-seed`. Validate the
body, apply it to the company, then acknowledge it.
The seed is customer free text, so it is bounded and validated in
`packages/shared` and read from the JSON body only. It is never read
from an `x-paperclip-cloud-*` header. That header set is the trusted
identity envelope: every member is derived server-side from the host
plus verified domain records, and that is exactly what makes it
trustworthy. Mixing user content into it would remove the property. A
test plants a mission on a cloud header and asserts that the body value
wins.
Application reuses the shapes the first-run wizard already produces, so
a seeded stack and a manually onboarded one look the same afterwards:
- The mission becomes the company-level goal. A multi-line mission
splits into a title and a description, as the wizard does.
- The agent becomes the company's first hire. Its free-text role ("Chief
of Staff") lands on `title`. The structural `role` stays `ceo`, which is
what the org chart and the default-instructions lookup read.
- The first task becomes an issue in the Onboarding project, assigned to
that agent.
Cloud retries until it gets a 2xx, and it reads any 2xx as "the tenant
holds this content". So the endpoint is idempotent per `revision`. A new
`company_onboarding_seeds` table records the applied revision together
with the goal, the agent and the issue it produced. A replay of a
revision that already matches is a successful no-op. A later revision —
the customer edited their answers — updates those three rows in place
instead of creating a second agent and a second task. The record is
written last, after every other write has landed, so a partial
application cannot present itself as acknowledged.
Everything is applied before the 200 is sent. This is an ordering
guarantee, not eventual consistency. The tests read the database
immediately after the response, with no waiting and no polling, so a
lazy receiver fails them on a fast machine as well as a slow one. That
matters because the redirect into the tenant dashboard is gated on this
acknowledgement.
**Alternatives considered**
Store the seed and let the tenant UI apply it on first load. Rejected:
the dashboard redirect is gated on the acknowledgement, so a background
apply would let the dashboard open before the agent and the task exist.
The whole point is that it must not.
Reuse `POST /companies/:companyId/agents` and `POST
/companies/:companyId/issues` over HTTP from Cloud. Rejected: it needs
three round trips with no shared idempotency key, and it moves the "did
all of it land?" decision to the caller.
**Roadmap alignment**
This completes an existing Cloud-to-tenant contract. It does not add a
new user-facing surface.
## What Changed
- Add `POST /api/companies/:companyId/onboarding-seed` in
`server/src/routes/onboarding-seed.ts`. It authenticates exactly as
`POST /api/companies/:companyId/logo` does, through
`assertCompanyAccess`.
- Add `server/src/services/onboarding-seed.ts`. It applies the mission,
the agent and the first task, and records the applied revision last.
- Add the `company_onboarding_seeds` table: schema, migration `0216`,
and journal entry. It holds the applied revision and the ids of the
goal, agent and issue the seed produced.
- Add `applyOnboardingSeedSchema` in `packages/shared`. It bounds
mission to 2000, agent name to 80, agent role to 120, task title to 200,
and task details to 2000 — the same limits Cloud enforces before it
sends.
- Mount the router in `server/src/app.ts` and register the path in the
OpenAPI document.
- Add `server/src/__tests__/onboarding-seed-route.test.ts` with 13
tests.
- The seeded agent is created on `claude_local`. This mirrors the
teams-catalog default for agents created server-side, where no human
runs an environment test first. `PAPERCLIP_ONBOARDING_SEED_ADAPTER_TYPE`
overrides it.
## Verification
```sh
pnpm typecheck # whole workspace, passes
npx vitest run \
server/src/__tests__/onboarding-seed-route.test.ts \
server/src/__tests__/openapi-routes.test.ts # 15 passed
```
The suite runs against embedded Postgres with migrations applied, so
migration `0216` is exercised by every test.
The route tests cover:
- the happy path — mission, agent and task all applied, read immediately
after the 200
- replay of the same revision — no second agent, no second task, no
second goal, no second project
- a later revision — the goal, agent and task are updated in place
- a multi-line mission splitting into a goal title and description
- a revision-only seed
- the activity log entry written once, and not again on a replay
- a caller without access to the company — 403, and nothing written
- a body with no revision — 400
- each field bound past its limit — 400
- a mission planted on an `x-paperclip-cloud-*` header — ignored, body
wins
- an existing Onboarding project — reused, not duplicated
Not verified here: the full Cloud-to-tenant walk against a live stack.
That needs a deployed Cloud and a provisioned tenant together, which is
separate staging work.
## Risks
Migration `0216` creates one new table. It adds no column to an existing
table, rewrites nothing, and backfills nothing, so it is safe to apply
online. The migration safety check passes.
The endpoint writes to a company. Access is enforced by
`assertCompanyAccess`, the same gate the company logo write uses, and a
test covers the denial.
Behavioral note for stacks that already hold data. If a company already
has a non-built-in `ceo` agent, a first seed updates that agent's name
and title rather than creating a second lead. Likewise a seed adopts an
existing company-level goal rather than adding a parallel one. This is
deliberate: the seed is the customer's own stated answer from signup,
and two competing missions or two leads would be worse than one updated
in place. In the intended case — a stack that Cloud has just activated —
none of these exist yet.
The seeded agent is created on `claude_local` with an empty adapter
config. It is idle and needs the usual credential setup before it runs.
Seeding it does not start it.
## Update — rebased onto master + review hardening
Master moved on after this PR was cut, so it was **rebased onto
`master`** and
the seed migration was **renumbered from `0212` to `0216`** (the merged
#11101
took `0212_onboarding_first_task_unique`); the drizzle journal was
re-stitched
and `check:migrations` passes.
Two things landed on top of the original receiver:
- **Mission-only walk contract (PAP-67 r17.4).** The tenant now owns the
first
agent and the first task via #11101's server-owned onboarding path,
which
stamps `ONBOARDING_FIRST_TASK_ORIGIN_KIND` and races safely on the
partial
unique index `issues_onboarding_first_task_uq`. A comment in the apply
path
documents why this receiver leaves the first task to that path on the
cloud
walk, and a paperclip-cloud `node:test`
(`src/onboarding/walk-seed.test.ts`)
asserts the walk's seed carries no `agent`/`firstTask`. The receiver
retains
the agent/first-task code for its documented body contract, kept inert
on the
cloud path by the mission-only seed.
- **Three Greptile P1 fixes** (`95622fa37`): concurrent application is
now
serialized under a per-company `pg_advisory_xact_lock` (no duplicate
goal/agent/project/task on overlapping pushes); a revised first task
carries
its resolved `assigneeAgentId`/`goalId`; and the
`company.onboarding_seed_applied`
audit write is best-effort so a logging failure can't leave the entry
permanently absent. Two new regression tests cover the first two.
## Model Used
Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking,
with tool use and code execution. Used for the original codebase
investigation, the implementation, and the tests. The rebase, migration
renumber, mission-only contract, and the three P1 fixes were done with
Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, with tool use
and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
04bf7a6ab5 |
feat(observability): instrument stage.sync host steps and home the agent process span (#11301)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - It runs each agent in a remote sandbox and emits OpenTelemetry spans for the sandbox bring-up and the run. > - A real trace showed two gaps. `stage.sync` had about 3 seconds of unattributed host work before its `pack` span. The persistent agent process showed a `sandbox.exec` span that outlived its parent by about 50 seconds. > - The gaps hide real cost and make the trace read as a sequencing bug, so an operator cannot see where startup time goes. > - This pull request wraps the two pre-`pack` host steps in their own spans. It also homes the long-lived process in a run-scoped `sandbox.agentProcess` span. > - The benefit is that startup time is fully attributed and the process reads as a resource that overlaps the turn, not a child that outlives its parent. ## Linked Issues or Issue Description No public issue exists. This is an enhancement to existing telemetry. It is described inline below, following `.github/ISSUE_TEMPLATE/enhancement.yml`. Prior related work: the merged PR #10999 added the run-time wrapper spans and the telemetry data-contract section this PR extends. **What existing behavior does this improve?** The sandbox bring-up and run OpenTelemetry trace. It closes two attribution gaps in that trace. **Subsystem affected** Observability for sandbox execution. The code lives in `packages/adapter-utils`. The span contract lives in `packages/shared/src/telemetry`. **Current behavior** `stage.sync` opens a `pack` span, but the git enumeration and the baseline content-hash walk that run before `pack` have no span, so about 3 seconds read as a gap. On the streamed process-session path the agent process launches fire-and-forget inside the ~2.3 second `bridge.process-session` bring-up step, so its `sandbox.exec` span parents to that step and then runs about 50 seconds. The child dangles past its parent and overlaps `agent.turn`. **Proposed behavior** Wrap the two pre-`pack` host operations in `snapshot.git` and `snapshot.baseline` spans under `stage.sync`. Wrap the streamed launch in a run-scoped `sandbox.agentProcess` span that parents to the live run root (`task.run` at launch). **Reason and benefit** Startup time is fully attributed. The long-lived process reads as a resource that overlaps the sibling `agent.turn`, not a mis-parented child. **Breaking changes** None. The spans are opt-in and export only when an OTLP endpoint is configured. The span seam is a no-op when no runner is injected. No first-party telemetry event changes. ## What Changed - `sandbox-managed-runtime.ts`: add `snapshot.git` and `snapshot.baseline` spans around the git enumeration and the baseline content-hash walk, nested under `stage.sync`, through a shared `runStepSpan` helper that `pack` now also uses. - `execution-target.ts`: wrap the fire-and-forget streamed launch in a run-rooted `sandbox.agentProcess` span, so it parents to the live run root and holds the inner `sandbox.exec`. The `.then`/`.catch` chain became try/catch inside the span callback, with identical frame-ingestion behavior. - `packages/shared/src/telemetry/README.md`: update the span table and the parenting prose. Add `snapshot.git`, `snapshot.baseline`, `pack`, and `sandbox.agentProcess`, and document the intended `sandbox.agentProcess` / `agent.turn` overlap. - Tests: update the executor span-tree test (`childNames` and parent assertions), update the `sandbox-managed-runtime` span-set and nesting tests, and add two `execution-target-sandbox` tests (the launch opens `sandbox.agentProcess`; it parents to the run root, not the bring-up step). ## Verification - Run `npx vitest run` on the three affected test files. Result: 174 tests pass. This includes the updated executor span-tree test and the new `sandbox.agentProcess` open and parenting tests. - Run `tsc --noEmit` in `packages/adapter-utils`. Result: no errors in the changed source or test files. - The full 37-test streamed process-session suite passes unchanged. This confirms the try/catch restructure preserves frame delivery and exit/error behavior. - Pre-existing and unrelated to this PR (present on `master`): `tsc` errors in `execute.ts` / `execute.test.ts` / `remote-spawn-smoke.test.ts` (`onAgentStderr` / `spawnCwd`), and a `check:forbidden-tokens` failure from internal `PAP-###` ids in `ui/src/components/IssueRecoveryActionCard.test.tsx`. This PR does not touch those files, and its own diff is token-clean. ## Risks Low. The change adds instrumentation on the opt-in span path and does not change control flow on the default path. The one production restructure is the streamed launch, which stays fire-and-forget, so bring-up does not block on it. Only the streamed path gains `sandbox.agentProcess`; the legacy poll path launches the process detached and has no host-side long-lived span to home. ## Model Used Anthropic Claude Opus 4.8 (`claude-opus-4-8`), about 200K-token context, agentic tool use through Claude Code. The trace was reviewed through the Honeycomb MCP. The code was written and tested with the model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e5a7fd7038 |
Add sandbox device-login for the Codex adapter (#11237)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Agent adapters connect Paperclip to tools such as the Codex command line tool. > - A sandboxed Codex agent may start without a credential. > - The operator needs a safe sign-in flow that does not expose credentials to the shared package or the sandbox. > - This pull request adds a company-scoped device-login flow with a temporary Daytona sandbox. > - The flow promotes the credential only after readiness checks pass and removes the temporary sandbox after use. > - The result lets an operator sign in to a sandboxed Codex agent from the agent form. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** A Codex adapter that runs in a sandbox cannot authenticate when the company has no pre-provisioned Codex credential. **Proposed solution** Add a company-scoped device-login session. Start a temporary sandbox, run `codex login --device-auth`, stream the code and URL, verify readiness, promote the credential, and delete the sandbox. **Alternatives considered** Pre-provisioning a credential does not support first-time sandbox login. Keeping the credential in the login sandbox does not provide a durable company credential. **Roadmap alignment** This supports the roadmap item for cloud and sandbox agents. **Additional context** The flow uses a five-minute cleanup reaper, compare-and-set status changes, and a PostgreSQL advisory lock to protect promotion and cleanup. ## What Changed - Add the adapter login-session contract, database table, and migration. - Add company-scoped server routes and a service for sandbox device login. - Add credential promotion, readiness checks, and cleanup after login. - Add restart-safe cleanup for abandoned login sandboxes. - Add sandbox login controls to the agent creation and edit forms. - Keep device-login and vendor identifiers out of public shared and adapter UI symbols. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local exec vitest run` passed with 310 tests at the submitted commit. - The server login route, service, and reaper tests passed with 45 tests at the submitted commit. - The agent form render tests passed with 26 tests at the submitted commit. - The public-symbol leak check passed at the submitted commit. - A live Daytona sign-in flow still requires confirmation by a user with a live sandbox. ## Risks The migration adds a new company-scoped table. A promotion or cleanup race could remove a credential or leave a sandbox active, so the service uses claims, compare-and-set transitions, and an advisory lock. The live Daytona flow needs operator confirmation because local tests do not provide a real browser sign-in. ## Model Used OpenAI Codex, GPT-5, tool use and code execution, extended reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
23a1b025c2 |
feat(server): chunked resumable company import transfers (#11223)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import moves large packages into an instance, and since the upload cap rose to 1 GB, the transport is the weak point: one HTTP request, buffered fully in memory, with no resume > - A dropped connection at 90% of an 800 MB upload starts the whole transfer over, and a server restart loses all progress > - This pull request adds the server side of chunked resumable import transfers: a durable run ledger and routes that accept the same import zip as verified ~32 MB parts spooled to disk > - An interrupted transfer resumes from the parts already uploaded — across dropped connections, page refreshes, and server restarts — and peak upload memory drops from the whole package to one part > - The benefit is that large imports become reliable on real-world connections instead of all-or-nothing ## Linked Issues or Issue Description **What happened?** Large company imports travel as a single HTTP upload. On a slow or flaky connection, any interruption discards all progress and the upload restarts from zero. The server buffers the entire compressed package in memory during upload. A server restart mid-upload loses the transfer entirely. With the upload cap now at 1 GB, these failure modes govern exactly the imports the cap was raised for. **Expected behavior** A large import upload survives interruptions: already-transferred data is kept and verified, only the missing remainder is re-sent, and the server's memory use during upload is bounded by a part, not the package. **Steps to reproduce** 1. Import a multi-hundred-MB company package over a connection that drops mid-upload. 2. The upload fails; retrying starts from byte zero. 3. Repeat on an unstable connection and the import may never complete. ## What Changed - New `company_transfer_runs` table (drizzle schema + migration) and `companyTransferRunService`: one row per transfer with a content-derived idempotency key, per-part completion recorded atomically and idempotently, resume scoped to actor and direction, completed runs short-circuiting retries of identical content. - New transfer routes beside the existing import routes, same authorization: declare a sliced zip (`POST /import/transfers` — validates cap, 64 MB part ceiling, contiguity, size sums, sha256 format), upload parts (`PUT .../parts/:n` — raw body, hash-and-size verified before an atomic write to a disk spool under the instance root; re-uploads are no-op successes), poll resume state (`GET .../:id` — missing parts recomputed from disk), and apply (`POST .../:id/apply` — requires all parts, re-verifies the assembled zip against the whole-file hash fail-closed, then feeds the existing import pipeline through factored helpers rather than duplicated logic). - Hourly sweep fails and cleans spools idle for 24 h; a swept transfer honestly reports all parts missing on resume. - Strict UUID gating on run ids before any filesystem path construction. - The existing single-shot upload path is untouched; clients arrive in the follow-up PR. ## Verification - Transfer route suite (embedded Postgres): create/upload/status/apply round-trip with a real imported company, out-of-order parts, wrong-hash part rejected and unrecorded, re-upload no-op, apply-with-missing-parts rejection, resume after failure with prior progress intact, assembled-hash mismatch failing closed with spool deletion, actor scoping 404s, async-job apply, sweep followed by honest resume. - Ledger suite (embedded Postgres): part idempotency, actor/direction scoping, completed-run short-circuit, cancelled runs staying cancelled. - Existing portability route suite unchanged and green; server + db typechecks clean. Exact counts in the PR checks. ## Risks - New routes are additive; the existing import path is untouched. The transfer routes carry the same board authorization as the import routes they sit beside. - Disk spool: bounded by the existing upload cap per transfer, cleaned on success, failure, hash mismatch, and by the 24 h sweep. Spool paths are strict-UUID-gated. - The apply step still materializes the assembled zip in memory once (same profile as today's single-shot import at apply time); upload-time memory drops to one part. - Known limitation, deliberate: transfers are keyed on content alone, so identical package content cannot be imported twice without re-exporting (surfaced explicitly to the caller). Acceptable for v1; noted for review. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0044fa8904 |
Let tenants edit env vars on managed sandbox environments; add managed-sandbox-only mode (#11200)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments give each agent run an execution target: the local host, SSH, or a sandbox provider > - A managed deployment can provision one platform-managed sandbox environment through the `PAPERCLIP_MANAGED_CONFIG` `environments` section > - That row is fully locked today. A tenant cannot add environment variables for their agents. There is also no way to hide local execution — run selection falls back to the local row > - A platform that manages the sandbox for its tenants needs both: the tenant adds env vars (and nothing else), and local execution is neither visible nor reachable > - This pull request opens exactly one tenant edit (env vars) on the managed sandbox row, and adds an `enableManagedSandboxOnly` mode that hides local and makes run selection fail closed > - The benefit is a complete managed-sandbox experience with no change for self-hosted instances ## Linked Issues or Issue Description **Subsystem affected** Environments (managed sandbox provisioning, environment routes, run environment selection) and the environments UI. **Problem or motivation** Platform-provisioned sandbox environments (`metadata.managedByPaperclip`) reject every write on cloud-managed instances. Agents often need environment variables inside their sandbox. The tenant has no way to set them on the managed row. Separately, an operator cannot remove local execution: the environment list always shows the local row, and run selection falls back to it when no default is set. **Proposed solution** Allow an envVars-only PATCH on the managed sandbox row, and echo those env vars back for editing. Add a managed-tier feature (`enableManagedSandboxOnly`) that hides the local environment from all read surfaces and redirects local-landing run selection to the managed sandbox environment, failing closed when it is unavailable. **Alternatives considered** UI-only hiding of the local row. This was rejected: it does not stop a run from resolving to local, so it is presentation without enforcement. Full unlock of the managed row was also rejected: name, driver, and config stay platform-owned so boot reconciliation cannot fight tenant edits. ## What Changed - `server/src/routes/environments.ts`: the platform-provisioned write floor admits an envVars-only PATCH on the generalized managed sandbox row (sandbox driver, `managedByPaperclip`, not legacy kubernetes-marker rows). Name, driver, config, status, metadata, and DELETE stay rejected. The read floor stops blanking env vars on that row; credential-shaped config keys stay redacted for every actor. Legacy kubernetes-marker rows keep the full floor. - Same file: under `enableManagedSandboxOnly`, the environments list and the by-id read omit the local row for every actor, including instance admins. - `server/src/services/execution-workspace-policy.ts`: `resolveExecutionWorkspaceEnvironmentId` gains the managed-sandbox-only inputs. A selection that lands on the local environment is redirected to the managed sandbox environment. With no active managed row it throws `ManagedSandboxUnavailableError` — never local. Non-local selections (ssh, user-created sandboxes) are untouched. - `server/src/services/heartbeat.ts`: the run path reads the flag, looks up the managed row (`findManagedSandboxEnvironment`, new read-only finder in `environments.ts`), and passes both to the resolver. Mirrors the forced-kubernetes precedent, which keeps precedence when both regimes are on. - `server/src/services/managed-environments.ts`: after a successful reconcile, the instance default environment moves to the managed sandbox row when the current default is unset, local, or dangling. A tenant-chosen custom environment is never overridden. - `packages/shared`: new `enableManagedSandboxOnly` key (schema default false, catalog tier `managed`, cloudDefault false, selfHostedDefault false) and the matching interface field. - UI: managed rows show a "Managed by Paperclip" lock badge; editing one opens a dedicated env-vars-only editor that sends the one PATCH shape the server admits (the old full form failed with a 403 on save). New `ui/src/lib/managed-sandbox-environment.ts` mirrors the local filter for cached lists (applied in the project picker; the agent picker already excluded local). The experimental settings page gains the toggle at its alphabetical card position. `environmentsApi.update` now declares the `envVars` field it already sent. - Tests: environment route floor coverage (envVars-only accepted, mixed bodies rejected, legacy rows still blanked and locked, local hidden and 404 under the flag, self-hosted unchanged), an embedded-postgres service test pinning that boot reconciliation never touches tenant env vars, resolver redirect/fail-closed cases, managed-environments default-stamping cases, and UI lib/settings tests. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/environment-routes.test.ts src/__tests__/environment-service.test.ts src/__tests__/execution-workspace-policy.test.ts src/services/managed-environments.test.ts` — all pass. - `pnpm --filter @paperclipai/shared exec vitest run` — 425 pass (catalog/schema default parity is pinned by an existing test). - `pnpm --filter @paperclipai/ui exec tsc --noEmit` and the affected UI suites (CompanyEnvironments, InstanceExperimentalSettings incl. card-order test, new lib test) — all pass. - Full workspace `pnpm test`: 3,414 passed. 17 files report failures on this machine; the identical 17 fail on a clean `origin/master` worktree in the same environment (git-worktree/skills/embedded-postgres environment dependencies and plugin-SDK zero-test collections). One additional file (`issue-monitor-scheduler.test.ts`) failed one timing-sensitive test in one of two full-suite runs and passes 7/7 in isolation on this branch — a flake in a domain this diff does not touch. The branch introduces no new failures. - Self-hosted zero-delta: every new behavior is gated on the cloud-managed instance check or the new flag, which defaults to false in schema and catalog; pinned by the "does not floor platform-marked rows on self-hosted instances" and flag-off tests. ## Risks - Behavior is opt-in twice over: the write-floor exception applies only to rows the managed-config provisioner stamps, and the hiding/forcing applies only when `enableManagedSandboxOnly` is on (default false everywhere). Self-hosted instances see no change. - The env-vars echo is scoped to the generalized managed sandbox row; legacy kubernetes-marker rows keep the blanket floor because pre-generalization builds may have written platform values there. - Fail-closed run selection means a managed instance with the flag on and an archived managed row (provider plugin down) refuses runs with a precise error instead of running locally. That is the intended posture. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
7ea2068ef8 |
fix(files): only highlight accessible workspace file links (#11090)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task comments can contain references to files in project and execution workspaces. > - Paperclip detected path-shaped inline code and showed it as an actionable file chip. > - The UI did not first confirm that the current board session could open the file. > - Missing, denied, ambiguous, remote, and unsupported files therefore looked actionable and failed after a click. > - This pull request adds an issue-scoped availability check and promotes only confirmed files to chips. > - The benefit is that the task thread shows a file action only when that action can succeed. ## Linked Issues or Issue Description **What happened?** Task comments promoted path-shaped inline code to file chips before Paperclip checked the file. A chip could point to a missing, denied, ambiguous, remote, or non-previewable file. The action then failed after the user selected it. **Expected behavior** Paperclip must show a file chip only after the server confirms that the current board session can open the exact file reference. All other path-shaped text must stay ordinary inline code. **Steps to reproduce** 1. Add a task comment that contains inline code with a missing or inaccessible workspace path. 2. Open the task thread as a board user. 3. Observe that the path looks like an actionable file chip. 4. Select the chip and observe that the file cannot open. **Paperclip version or commit** `19be4cf927` and earlier. **Deployment mode** Local dev and self-hosted server. **Access context** Board user. ## What Changed - Added shared request, response, and validation contracts for batched workspace-file availability checks. - Added an issue-scoped server endpoint that resolves file references with company, issue, workspace, and preview-access checks. - Added bounded batch concurrency and tests for missing, denied, ambiguous, remote, unsupported, and available files. - Added an issue-scoped UI availability registry that deduplicates, batches, caches, and invalidates file checks. - Changed task-comment markdown rendering so only confirmed files get chip styling and file-viewer behavior. - Bound each chip to the exact workspace target that passed the availability check. ## Verification - `pnpm exec vitest run packages/shared/src/workspace-file-resource.test.ts server/src/__tests__/file-resources.test.ts ui/src/components/MarkdownBody.test.tsx ui/src/components/WorkspaceFileMarkdownBody.availability.test.tsx ui/src/lib/remark-workspace-file-refs.test.ts ui/src/lib/workspace-file-availability.test.ts` — 93 passed, 35 skipped. - `pnpm check:token-gates` — clean. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — all server and UI groups passed. One unchanged CLI test saw the run-injected static AWS credentials and expected only its local `AWS_PROFILE`. The same test passed, 8 of 8, after removing only `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` from its process environment. ## Risks - File chips now appear after an asynchronous availability check, so path-shaped text can briefly render as inline code. - Availability results use the existing 30-second file-resource cache window. File-resource invalidation forces a new check. - The endpoint limits each request to 100 references and the client chunks larger sets. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The service did not expose a more specific model ID or context-window size. The agent used high-reasoning mode, repository tools, command execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
815e49bb7c | feat: make chat-style tasks the default experience (#11101) | ||
|
|
b58ce27a02 |
fix: isolate execution workspace summaries (#10790)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip gives operators a summary for each workspace. > - An execution workspace detail page used the parent project-workspace summary slot. > - Two execution workspaces under one project workspace could therefore show the same summary. > - This pull request gives each execution workspace its own summary scope. > - It also limits the summary snapshot and generated issue to that execution workspace. > - The benefit is that a new or parallel execution workspace cannot inherit unrelated status. ## Linked Issues or Issue Description **What happened?** An execution workspace detail page read and refreshed the summary slot for its parent project workspace. Parallel execution workspaces could show the same status and include issues from each other. **Expected behavior** Each execution workspace must have one isolated summary slot. Its generated snapshot must include only issues assigned to that execution workspace. **Steps to reproduce** 1. Create two execution workspaces under one project workspace. 2. Add different issues to each execution workspace. 3. Generate the summary in the first execution workspace. 4. Open the second execution workspace. 5. Observe that the old implementation could reuse the first summary. **Paperclip version or commit** The problem exists on `master` before this pull request. **Deployment mode** The issue affects both local trusted and authenticated deployments. ## What Changed - Added `execution_workspace` to the shared summary-slot scope contract. - Validated execution-workspace ownership and stored generated summary issues on the correct execution workspace. - Limited execution-workspace snapshots to issues with the matching execution workspace ID. - Updated the execution workspace page to use its own summary slot. - Updated Summarizer instructions, routine options, catalog metadata, documentation, and regression tests. ## Verification - `NODE_ENV=test pnpm exec vitest run packages/shared/src/summary-slot.test.ts server/src/__tests__/summary-slots.test.ts ui/src/pages/ExecutionWorkspaceDetail.test.tsx` — 30 focused tests passed; the embedded-Postgres server tests were run outside the process-restricted sandbox. - `pnpm check:token-gates` — passed. - `pnpm --filter @paperclipai/skills-catalog validate` — passed with 17 catalog skills. - [Latest-head GitHub Actions](https://github.com/paperclipai/paperclip/actions/runs/31491475405) — all 22 jobs passed on `beea14cbaf`, including typecheck, build, server/workspace tests, serialized suites, e2e, canary, and aggregate verification. One unrelated adapter cleanup test initially hit an `ENOTEMPTY` temp-directory race; its single permitted rerun passed. - Greptile — 5/5 confidence on `beea14cbaf`, 12 files reviewed, zero comments added, and zero unresolved threads. ## Risks - Low risk. The new scope is additive. - Existing project and project-workspace summary slots keep their current keys and behavior. - A summary generated for an execution workspace now excludes sibling workspace issues by design. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The deployment does not expose a more specific model ID or context-window value. It used agentic reasoning, repository tools, code execution, and GitHub tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
35aaaa0bd0 |
feat(server): preserve task timestamps and hierarchy through company import/export (#11193)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company export/import moves a whole company — agents, tasks, comments — between instances as a portable bundle > - The bundle never carried task timestamps or parent links: the export writes neither, the importer lets database defaults stamp "now", and sub-tasks arrive flattened > - Boards sort by recency, so every imported task showing "created just now" collapses the task list into import order, and the task hierarchy the user built is gone > - This pull request adds created/updated/started/completed/cancelled timestamps and a parent link to the bundle (schema v7), preserves them end to end on import, and keeps comment imports from clobbering a preserved updated time > - The benefit is that an imported company reads like the company the user left: same recency order, same task tree ## Linked Issues or Issue Description **What happened?** After a company import, every task showed as created at import time. Recency sorting collapsed to import order, and parent/child task nesting disappeared. The user called out losing "the meaningful task hierarchy and recency sorting". Cause: the export bundle has no fields for task timestamps or parent links, the importer lets `defaultNow()` win on insert, and the comment importer bumps every touched task's `updatedAt` to now. **Expected behavior** An imported company preserves each task's creation/update/start/completion times and its position in the task tree, so sorting and nesting on the destination match the source. **Steps to reproduce** 1. On a source instance, create tasks over several days, including sub-tasks nested under parents. 2. Export the company and import it into another instance. 3. Every task shows the import moment as its creation/update time and all tasks are top-level. ## What Changed - Export writes `createdAt`/`updatedAt`/`startedAt`/`completedAt`/`cancelledAt` (ISO, only when set) and `parent: <taskSlug>` into each task's bundle extension; a parent outside the export selection drops the edge with an aggregate warning, mirroring the existing blocker-edge warning (`server/src/services/company-portability.ts`). - Bundle schema version 6 → 7. All new fields are optional: v5/v6 bundles import unchanged with a version-aware downlevel warning; bundles newer than the board still fail closed. - Manifest parsing validates the new timestamps like comment timestamps (invalid → warn and ignore, never a hard failure); shared types and the zod validator carry the new optional fields. - Import resolves parent slugs to pre-generated destination ids, drops self-references and cycles from tampered bundles with warnings, and orders rows parents-first because the self-referencing FK is checked per insert chunk. - `importIssues` writes the preserved timestamps (falling back to insert time when absent; `startedAt` stays null unless bundle-carried, per #11191's semantics) and `parentId`. - `addImportedComments` no longer blanket-bumps `updatedAt = now()`; it takes `GREATEST(updated_at, newest imported comment createdAt)`, so a preserved update time never regresses while unpreserved rows keep the old behavior. ## Verification - `pnpm vitest run server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts server/src/__tests__/productivity-review-service.test.ts` — 102 passed, 1 pre-existing opt-in benchmark skip. Includes: full round-trip with exact timestamp equality and a 3-deep parent chain against embedded Postgres; v6 back-compat (defaults + warning); forward-compat rejection (v8); cycle/self-reference/invalid-timestamp tampered-bundle handling; comment-bump preserve-awareness in both directions. - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/shared typecheck` — clean. ## Risks - **Rollout ordering**: a board on the previous build (max schema v6) refuses bundles exported by this build (stamped v7) — the existing newer-than-supported rejection, working as designed. Cross-instance moves need the importing board upgraded first. Called out here so operators aren't surprised during the transition window. - Parent edges from tampered bundles are dropped with warnings rather than failing the import; blocker relations already behave this way. - Timestamps are data-only; no destination schema migration. Stacked on #11191 (its commit is included here) — merge #11191 first; this PR then shows only the v7 changes. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |