mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
303340f1900723dec9bdb8cbdee7586068bd9f20
548
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
303340f190 |
fix(apps): document Railway SSH setup and sync with master
Register the board-only SSH setup operation with its shared request schema and error responses. Resolve current catalog and branding metadata conflicts. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a8d32e5e61 |
feat(sandbox-providers): add CreateOS sandbox provider (#13434)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent work runs in sandboxes that provider plugins supply > - Operators can choose a provider to run agent work > - CreateOS adds another provider with workspace-preserving pause and resume > - This pull request adds a CreateOS provider plugin > - The benefit is that operators can preserve a workspace between runs without keeping its compute active ## Linked Issues or Issue Description Refs #13203 and the earlier closed #13096. This continues the CreateOS contribution from @bhautikchudasama and @ashwaq06. The branch preserves the original implementation commit. Thank you to both contributors. When squash-merging, preserve the original author's credit in the squash commit body: ```text Co-Authored-By: bhautikchudasama <BhautikChudasama@users.noreply.github.com> ``` The original fork rejects maintainer pushes. This branch includes the merge-conflict resolution and review fixes. The request is described below using `adapter_request.yml`. **Agent or provider** CreateOS sandbox API (https://api.sb.createos.sh). **Why this adapter is useful** CreateOS can pause a sandbox and resume it by ID. The workspace survives the pause. This adds a reusable-lease option to the existing sandbox provider system. **How the agent is invoked** Build and install the local plugin as described in its README. Open Instance Settings, then Environments. Select the `createos` driver. Supply an API key and shape. The driver then supplies sandbox leases for agent runs. **Are you willing to implement it?** Yes. This pull request is the implementation. ## What Changed - Adds the `createos` sandbox provider under `packages/plugins/sandbox-providers/createos`. - Calls the CreateOS HTTP API directly. The package adds no vendor SDK. - Implements the environment lifecycle hooks, incremental process output, and binary workspace sync. - Registers the optional bundled provider and its trusted host credential fallback. The fallback is limited to the official API origin; custom endpoints require an explicit key. - Lists the package in the release manifest with `publishFromCi: false` until its first npm publish is bootstrapped. - Waits through delayed pause/resume state updates without duplicate action requests. - Cancels queued API requests promptly while preserving request spacing. - Uses direct CLI invocation in the setup guide so paths and IDs are passed without an extra shell expansion. - Includes current master and retains its existing Git-subfolder containment fix. ## Demo Fresh setup and a run against a CreateOS sandbox. https://github.com/user-attachments/assets/e71b9e06-c006-4fb9-b847-52dfd68f6110 https://github.com/user-attachments/assets/43b5ac75-66bd-4f76-8563-67e4c7759084 ## Verification All 25 jobs in [CI run 34884260542](https://github.com/paperclipai/paperclip/actions/runs/34884260542) passed at commit `f8d0997677024b784fdadf9d44a84c01cb4e813c`, including typecheck, build, native runner verification, server and workspace tests, browser tests, and the canary release dry run. Greptile reviewed the same commit at 5/5 with no unresolved review threads. GitHub reports no merge conflicts. The remaining merge gate is code-owner approval for the new `package.json`, as required by `.github/CODEOWNERS` and the `master` ruleset. Reviewers have been requested automatically. Local checks passed: - Provider: `pnpm typecheck`, `pnpm test` (52 passed, one live smoke skipped), and `pnpm build`. - Host: focused credential and bundled-plugin tests (17 passed), plus CLI invocation safety (39 passed). - Release: package manifest check and release policy tests (18 passed). The full local `pnpm test:run` attempt caught the README command issue; its focused rerun now passes. The full local run stopped after its general-server group: 7,804 tests passed, with unrelated embedded PostgreSQL startup failures and 10 failures in unchanged runtime-skill-cache tests (`EACCES` on directory rename on macOS). It did not reach the later test groups. Local `pnpm -r typecheck` and `pnpm build` reach the runner package and stop because this machine has no Rust/Cargo installation. The corresponding CI checks passed on provisioned runners, as linked above. The live CreateOS smoke requires explicit provider credentials and was not run during this review. It is available with `CREATEOS_LIVE_TEST=1 pnpm test` in the provider directory. The author supplied the demo links above. ## Risks The provider is opt-in and is not installed by default. It is available through a local-path install or explicit image inclusion. npm publication remains disabled until a maintainer bootstraps the package and enables publishing. Sandbox creation has no idempotency key. An ambiguous create response can leave a resource that requires provider-account inspection. Process tracking is in memory; durable lease recovery belongs to the host. The provider does not advertise guaranteed expiry, interactive login, snapshots, duplex channels, or ingress. Live native-runner qualification remains outside this PR's tested claims. ## Model Used Original provider implementation: human-authored by @bhautikchudasama, as reported in #13203. The original description reports Claude Opus 5 assistance. Review and follow-up fixes: OpenAI GPT-6 via Codex, with code review, editing, and tool execution. The precise runtime model variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks; full-suite environment limits documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: bhautikchudasama <bhautikrchudasama@gmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9adeabb590 |
fix(connections): unblock personal MCP auth discovery (#13497)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections let people give agents access to external tools. > - A personal connection needs the current user's authorization. > - A new MCP URL must be probed before Paperclip can discover its sign-in method. > - Requiring a personal grant before that probe prevents sign-in from starting. > - This pull request permits the initial probe for a creator-owned draft with no credentials. > - People can complete personal setup while later requests retain authorization checks. ## Linked Issues or Issue Description **What happened?** Connecting an unknown MCP URL with "Just me" failed with HTTP 502 and "This connection needs the current user's authorization". Paperclip checked for a personal grant before contacting the provider. The health wrapper also changed the expected authorization error into a server error. **Expected behavior** Discover OAuth and start browser sign-in. Create a personal grant after consent. For a public endpoint, discover its tools and create the empty personal grant after a successful probe. Keep missing authorization on later health checks as HTTP 422. **Steps to reproduce** 1. Add an unknown remote MCP URL with no saved credentials. 2. Select "Just me". 3. Check the link. Before this fix, the request fails before sign-in or tool discovery. **Paperclip version or commit** The three original regressions fail against `6cfe4acff` with the service fix removed and pass with it restored. **Deployment mode** The defect was reported in production and reproduced in local server tests with isolated PostgreSQL. Related work: Refs #11831. Refs #11144. Searches found no duplicate fix. ## What Changed - Allow an initial credential-free probe only for the creating user's personal draft with unknown authentication and no supplied credentials. - Leave OAuth grant creation to the callback. Create an empty personal grant only after a public probe succeeds. - Preserve the personal grant for URLs that already contain a credential. - Make empty personal grant creation conflict-safe without overwriting a concurrent grant or duplicating its creation audit. - Commit the empty grant and audit atomically. Retain the established public draft identity after a later catalog failure, so a failed retry cannot remove a successful retry's grant. - Run the following catalog/default-profile step in its own transaction, so failures discard partial catalog, profile, binding, and audit changes without deleting the established identity. - Verify archived personal connections retain their owner: another user is rejected before probing, while the original owner can resume setup. - Preserve `user_authorization_required` and HTTP 422 in health failures. - Cover the connect and OAuth callback routes, real loopback HTTP, credential-bearing URLs, and later health checks. - Document personal setup and the test fixtures. ## Verification - Red/green: the original three tests failed with the exact reported error before the fix and passed after it. - The credential-bearing personal URL regression also failed before its guard was added. - Focused suite: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/generic-mcp-connection.test.ts src/__tests__/tool-access-service.test.ts` passed all 385 tests across the two suites. An earlier run had a socket hang-up in an existing agent-permissions test; the unchanged suite passed on rerun. - The concurrent rollback regression failed before its fix because the successful retry's grant was deleted. It now verifies the grant and draft survive and a later normal health check succeeds. - Database fault injection during profile-entry insertion reproduced partial catalog writes before the transaction fix. The regression now verifies unchanged catalog rows, no partial profile/bindings, a retained grant, and successful retry. - CI's first serialized-server shard 3 attempt failed an existing peer-agent mutation test (the real run-context guard ran despite the test's mock). The test passed in isolation and all 108 tests in that suite passed unchanged locally. The single failed-shard rerun passed without code changes. - Final-commit CI: all 32 applicable checks passed on `639f037987352cab6084c4ebfa5dbf7b0aed6046`, including all 385 affected tests, the full test matrix, browser suite, build, typecheck, release checks, and security checks. The two Storybook-only checks were not applicable and skipped. Greptile is 5/5 with all review threads resolved. [Successful CI run](https://github.com/paperclipai/paperclip/actions/runs/35027478353). - `pnpm -r typecheck` passed. - `pnpm smoke:mcp-fixtures -- --require-paperclip` passed. - `pnpm build` passed. - Full local `pnpm test:run` was attempted: its general-server group finished with 12,372 passed, 4 failed, and 70 skipped tests. The run started before review edits; its two MCP failures used the old cached service (including an insert without the new conflict clause). All 385 focused tests pass on the final code. The other failures were existing workspace-cleanup and runtime-port tests; their unchanged suites passed on rerun (66 passed, and 25 passed/3 skipped). The local command stopped before later groups. The final-commit CI matrix is the full-suite merge gate; this local run is not claimed as green. ## Risks - The initial probe must not become a general authorization bypass. It is restricted to the creating user's draft. Normal health checks retain authorization enforcement. - Public endpoints get a personal grant with no secrets only after they answer successfully. Credential-bearing URLs keep their existing grant. - The concurrent-probe regression seeds the catalog and default profile to isolate grant creation. Existing first-time catalog/profile creation races are outside this change; this does not claim to make the entire setup flow concurrency-safe. - No database migration or UI change is required. OAuth tests use a simulated provider; the public endpoint test uses real loopback HTTP. ## Model Used OpenAI Codex, GPT-6-based assistant for regression tests and PR preparation; a GPT-5-based Codex assistant assisted with the initial implementation. Exact runtime model IDs and context-window sizes are not exposed in this session. Both used reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
544c3476a8 |
feat(server): wrap a bare Cloud UI snippet body in a <script> element (#13496)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A Cloud-managed instance injects an operator-owned HTML snippet before `</body>` through `injectCloudUiSnippet`. > - The Cloud control plane delivers that snippet to each instance as an environment variable through a provider API. > - The provider edge firewall now base64-decodes the request payload and blocks any value whose decoded form contains a `<script` marker. > - A working snippet needs a script tag, so every delivery is now blocked and the operator cannot ship the snippet at all. > - This pull request treats a resolved value that does not start with `<` as a bare script body and wraps it in a `<script>` element at injection time. > - The benefit is that the operator can deliver a tag-free body that the firewall passes, and the instance restores the script element on the page. ## Linked Issues or Issue Description No public issue exists. The problem is described below. Related PRs (searched the PR list; none duplicate this change): - Refs #13168 — added `injectCloudUiSnippet`, the mechanism this extends. - Refs #13245 — added the base64 `_B64` path on the assumption that base64 clears provider WAFs. That assumption no longer holds; this PR is the successor. - Refs #13441 — the in-product feedback approach that the Cloud-owned snippet replaced (closed). **What happened?** `injectCloudUiSnippet` injects `PAPERCLIP_CLOUD_UI_SNIPPET` (or the base64 `_B64` form) verbatim before `</body>`. A working value must therefore contain a `<script>` tag. The Cloud control plane delivers this value as an environment variable through a provider API that sits behind an edge firewall. The firewall now base64-decodes the payload and rejects any value whose decoded form contains `<script`. The delivery request fails, so the snippet cannot reach the instance. **Expected behavior** The operator can deliver the snippet through the provider API, and the instance runs it. **Steps to reproduce** 1. Build a snippet that contains a `<script>` element. 2. Deliver it to a Cloud instance through the provider variable API, in plain or base64 form. 3. The firewall rejects the request. The instance never receives the snippet. ## What Changed - `injectCloudUiSnippet` now wraps a resolved value that does not start with `<` in a `<script>...</script>` element. A value that already looks like markup is injected byte-for-byte, so existing full-`<script>` snippets are unchanged. - Added two unit tests: one for the wrap path (plain and base64), one that confirms `$&`-style replacement tokens in the body survive the wrap. ## Verification - `pnpm exec vitest run src/__tests__/cloud-ui-snippet.test.ts` — 17 pass. - `pnpm exec vitest run src/__tests__/static-index-html.test.ts` — 2 pass. - Confirmed the changed file has no type errors. The full `pnpm run typecheck` needs the Rust runner toolchain (`cargo`), which is absent on this machine; CI runs it in full. - Verified out of band that a tag-free body clears the provider firewall and wraps into valid, executable standalone JavaScript with no premature `</script>` close. ## Risks Low risk. The change adds a branch that only affects values that do not start with `<` — previously injected as inert text, never as a running script. Values that start with `<` keep their exact bytes. The content is trusted operator HTML, consistent with the existing contract. Roll back by reverting this commit. ## Model Used Claude — `claude-fable-5` (Fable 5), extended thinking, with tool use and code execution in Claude Code. A human author reviewed and verified the change before submission. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d49f168381 |
fix: publish sandbox files on legacy and native runners (#13493)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents must publish generated files so users can inspect their results after a sandbox stops. > - Legacy sandbox bridges blocked attachment listing and could not carry multipart binary uploads through the queue transport. > - The native runner has a separate verified file registration path that needs the same durable result. > - This pull request repairs legacy binary transport, makes native download receipts explicit, and reveals new outputs in the task Artifacts tab. > - Users can open generated files from either runner without a transport flag change. ## Linked Issues or Issue Description Refs #13355 for the existing native file publication path. Related filename fixes: #2615 and #4788. Related sandbox persistence work: #13376. This change repairs attachment delivery through the existing API; it does not add workspace persistence. **What happened?** The upload helper first lists task attachments to avoid duplicates. Both legacy bridge allowlists rejected that GET request with 403. A direct multipart upload also failed: the queue bridge accepted only JSON, excluded attachment uploads, and converted bytes to UTF-8 text. Enabling HTTP/2 alone did not fix the missing listing route. These failures occurred before attachment storage. **Expected behavior** Both runners can publish a workspace file, register its work product, bind it to a response, and return a working download. The file stays accessible after sandbox deletion. A new output opens the task Artifacts tab. The agent receives accurate errors and decides how to retry or report a failure. **Steps to reproduce** 1. Run a legacy agent in Daytona with the duplex bridge disabled. 2. Invoke the bundled upload helper with Bash on a PNG or PDF. 3. Repeat with the duplex bridge enabled. 4. Register the same file through the native runner with generic API tools disabled. 5. Retry registration, delete the sandbox, and compare the downloaded bytes with the original file. **Paperclip version or commit** The failing baseline was `f2c5e54dc`. This branch is rebased onto `6cfe4acff`. **Deployment mode** Source checkout with a local API and real isolated Daytona sandboxes. ## What Changed - Allow authenticated attachment listing, upload, and content download through both legacy bridge transports. - Add optional base64 body encoding to queue envelopes. Preserve the existing UTF-8 contract when the encoding field is absent. Decode binary bodies before forwarding them. - Preserve multipart headers. Bound raw bytes, encoded envelopes, and in-flight reservations. Retain timeout and uncertain-write behavior. - Preserve helper deduplication and return structured uncertain-write failures. Document explicit Bash invocation in live skills. - Add attachment IDs and content/download paths to native registration receipts. Reuse verified local and remote file reads, attachment storage, work-product registration, and response binding. - Preserve Unicode upload filenames and provide a valid Content-Disposition header. - Open the task Artifacts tab when new stored outputs arrive, including a closed desktop panel or mobile drawer. Deduplicate upload and registration events by object ID. Preserve manual selection on refetches, edits, and panel remounts. - Remove task artifact filters, the company Artifacts footer link, and the unassigned group heading and timestamp. ### Screenshot  This is the local display fixture. The image was generated separately and published through the attachment and work-product APIs. ## Verification - Post-rebase `pnpm -r typecheck` and `pnpm build` pass. - The post-rebase local `pnpm test:run` passed 12,369 tests before one existing conversation reset test timed out; all 33 tests in that suite pass when rerun with isolated test configuration. The aggregate command stopped before its remaining groups. GitHub runs the complete suite in separate shards. - All [GitHub verification checks](https://github.com/paperclipai/paperclip/actions/runs/35017893350) pass on `b66ac276dd3d5fc738a22ecea783400106a494d4`: 32 successful checks and two configured skips. The native-session recovery assertion initially raced its fire-and-forget Sentry report; all 13 tests pass locally, and the same-commit CI rerun passes all 170 suites (3,079 tests). - Live post-rebase Daytona: all three file-delivery tests pass. They cover the real Bash helper with the queue bridge, the helper with HTTP/2, and native `register_deliverable` with generic API tools disabled. - Daytona cases cover PNG/PDF bytes, spaced and Unicode names, duplicate registration, response binding, authorization controls, and byte-for-byte downloads after sandbox deletion. - Local focused coverage includes transfer bounds, malformed encoding, interrupted transfers, remote path containment, and native file verification. The attachment route suite passes all 32 tests, including an eight-case filename-header matrix for Unicode and special characters, inline and forced downloads, and full and partial responses. - Browser verification confirms image previews, persisted downloads, automatic Artifacts selection, and preserved manual selection after edits and reloads. Desktop/mobile component coverage passes. The latest UI cleanup passes its 10 affected tests and token gates. - Coverage limit: the Daytona tests call the real helper and native registration path directly. They do not replay a complete model-led image-generation task through the browser. Live command (requires a configured Daytona credential): ```sh PAPERCLIP_FILE_DELIVERY_DAYTONA=1 pnpm exec vitest run server/src/__tests__/file-delivery-bridges.test.ts ``` ## Risks - Binary queue bodies use more memory because base64 adds encoding overhead. Transfer and process limits must remain aligned. - An interrupted write can have an unknown result. The bridge reports this state and preserves stable retry identities. - New artifacts intentionally change the active task tab. Existing history and repeated updates must not take focus again. - Transport flag defaults, server authorization, frozen skill snapshots, and completion policies remain unchanged. No schema migration is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, code execution, and browser testing. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6cfe4acff7 |
fix: preserve README images in npm package (#13488)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its CLI is published as the paperclipai npm package, with the root README shown on the package page > - The root README uses repository-relative image paths so images render correctly on GitHub > - npm resolves those paths under the package repository directory, which is cli, so the image requests point to missing cli/doc/assets files > - This pull request prepares the generated npm README by converting only image src and srcset asset paths to stable raw GitHub URLs > - The benefit is that the same source README remains correct on GitHub and the published npm README displays its images ## Linked Issues or Issue Description **Where is the issue?** The issue is in the root README image assets and the npm packaging step in scripts/build-npm.sh. The affected public page is https://www.npmjs.com/package/paperclipai. **What's wrong?** The npm build copies README.md into cli/ before publishing. npm resolves relative image paths beneath the package repository directory, so doc/assets/banner.jpg becomes cli/doc/assets/banner.jpg. Those files do not exist, and the images render as broken on npm. **Suggested fix** Keep the root README paths relative for GitHub. Rewrite repository-relative image paths only in the generated npm README copy to absolute raw.githubusercontent.com URLs. ## What Changed - Added a small npm README preparation script that rewrites relative image src and srcset asset paths. - Updated scripts/build-npm.sh to use the preparation step when generating the npm package README. - Added a regression test for src, srcset, immutable refs, absolute URLs, and non-image Markdown links. - Pinned release-build image URLs to the source commit, while preserving tarball builds by passing their known source refs. ## Verification - Passed: node --test scripts/prepare-npm-readme.test.mjs - Passed: bash -n scripts/build-npm.sh scripts/e2e-install-lifecycle.sh scripts/e2e-update-migrations.sh - Passed: focused README and E2E migration harness tests - Passed: git diff origin/master...HEAD --check - Generated README asset URLs were checked against raw GitHub and all seven returned HTTP 200. ## Risks Low risk. The change affects only the temporary README generated for npm packaging. It does not change the GitHub README or runtime code. The generated npm README depends on the public raw GitHub asset URLs remaining available. ## Model Used OpenAI Codex, GPT-5. Tool-enabled repository inspection, code execution, browser verification, and git/GitHub operations were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes: # / Refs: # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (e.g. docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4cdc7b231 |
fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane prepares task workspaces before it starts an agent. > - Workspace preparation reads Git state so it can preserve edits and exclude private files. > - A failed scan was treated as a non-Git folder and lost its actual failure code. > - The resulting generic setup failure could not recover, even when the cause was temporary. > - This pull request keeps the cause and uses the existing bounded retry schedule before provider startup. > - Tasks can recover without human intervention, while permanent failures and exhausted retries stop with useful guidance. ## Linked Issues or Issue Description **What happened?** A Git scan error during managed repository preparation became `Configured repository folder is not a Git checkout`, followed by generic `setup_failed`. The agent never started. Generic recovery could not distinguish a temporary timeout from a bad workspace configuration. **Expected behavior** Keep the closed scan error code. Retry temporary timeouts and queue saturation under the existing shared budget. Preserve edits, exclusions, ownership, and pause gates. Stop permanent failures and exhausted retries with a specific explanation. Do not replay historical generic setup failures. **Steps to reproduce** 1. Configure a task project with a local Git source that must be copied into its managed repositories. 2. Make the ignored-file scan exceed its timeout before the agent starts. 3. Before this fix, the snapshot returns null and the run ends as non-retryable `setup_failed`. 4. Use the disposable browser fixture in `tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and test the full recovery path. **Paperclip version or commit** Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`. **Deployment mode** Built from source. The defect is in core workspace setup, not a specific model provider. Related work: Refs #13442 (managed repository preparation), Refs #11572 (bounded Git scheduler), Refs #12997 (separate adapter startup retry work), Refs #13469 (separate terminal-workspace scan performance work). ## What Changed - Return the non-Git fallback only for repository discovery. Propagate failed scans of a confirmed repository. - Replace full ignored status output with an ignored-only directory listing. Preserve NUL-delimited paths and exclusions. - Preserve typed, sanitized scan errors through workspace preparation and persist pre-provider failure details. - Retry only timeouts and queue saturation, using the existing durable two-retry budget and issue gates. Prevent generic recovery from adding another budget. - Show workspace-specific failure copy and actionable exhausted-recovery notices. - Add red-green unit tests, real-database restart and retry-boundary tests, and opt-in browser acceptance fixtures with real Git subprocess timeouts. - Document the recovery contract and browser verification procedure. ## Verification - Red: injected scan failures returned null instead of rejecting; setup lost the timeout code; task-thread and recovery notices had generic copy. - Green: 119 focused adapter/backend tests, 20 recovery-boundary tests, and 136 task-thread tests. - `pnpm -r typecheck` — passed. - `pnpm build` — passed on the final production code. - `pnpm check:token-gates` — passed. - The initial local `pnpm test:run` overlapped source edits and was interrupted after two late-added assertions saw pre-fix behavior; it is not counted as a green full run. A fresh final-head run passed all 229 tests across the six affected adapter/backend/UI suites. The clean latest-head CI full test matrix passed: all five general-server shards, all five serialized-server shards, and all three general-workspace shards. - Latest-head CI also passed all three browser shards and their aggregate gate, typecheck and release registry, build, runner verification, canary dry run, policy, Docker context integrity, and security gates. Greptile: 5/5, with the review thread resolved. - Browser: created a task in a disposable instance. A real Git timeout scheduled recovery, the next run completed through the run-scoped API without manual Retry, and Done survived reload. The deterministic process worker checked preserved source edits and excluded private files; no model calls were made. - `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec playwright test --config tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6 minutes). The persistent case made exactly three failed attempts, never started the worker, showed the cause-specific notice, stayed stopped for another scheduler tick, and retained Blocked after reload. - Extra red-green coverage: 50 recovery tests passed after fixing an exhausted-bootstrap classification that incorrectly implied unknown provider actions. Missing or uncertain evidence still retains the safety hold. - Verified the documented Git executable override during repository seeding. ## Risks - A confirmed repository scan failure now fails closed instead of falling back to directory sync. This prevents unfiltered copying but makes previously hidden errors visible. - Temporary host problems can create up to two additional setup attempts, 30 seconds apart. Permanent scan errors do not auto-retry. Generic recovery cannot reset this budget. - The durable retry path still enforces ownership, pause, and work eligibility. Integration tests cover restart, duplicate promotion, pause, exhaustion, and non-retryable categories. - No schema migration, new runtime setting, new retry budget, production deployment, or historical task replay. ## Model Used OpenAI Codex, GPT-5-based coding agent, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cceeb0aa66 |
test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4510bf7c9e |
ci: use code-owner-reviewed master for trusted PR workflow (#13470)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request CI uses a trusted workflow on the AWS runner fleet. > - The caller used a fixed SHA that also needed runner-group admission. > - A mainline pin update left CI queued because the group still allowed older SHAs. > - This pull request calls the trusted workflow on master, which requires code-owner review. > - New merged workflow versions can use the existing master runner-group entry. ## Linked Issues or Issue Description Related: #12968. That Dependabot PR proposes another SHA rotation. This change keeps this first-party workflow on master instead. **What happened?** CI run 34975562974 stayed queued because its trusted workflow SHA was absent from the runner-group allowlist. The fleet itself was healthy. **Expected behavior** New reviewed versions of the trusted workflow on master should receive runner access without a separate SHA allowlist update. **Steps to reproduce** Change the caller to a new trusted workflow SHA without adding that SHA to the restricted runner group. Its jobs remain queued. The master reference removes that recurring synchronization step. ## What Changed - Call `paperclipai/paperclip/.github/workflows/pr-trusted.yml@master`. - Exclude this exact first-party workflow from Dependabot updates. - Update the existing E2E shard workflow tests for the master caller contract. - Document the runner-group entry, required code-owner review, and old-reference retention. ## Verification - `actionlint .github/workflows/pr.yml` passed. - `node --test scripts/__tests__/e2e-shard.test.mjs .github/scripts/tests/cloud-runner-routing.test.mjs .github/scripts/tests/pr-runner-rust-cache.test.mjs .github/scripts/tests/pr-dependency-cache.test.mjs` passed: 35 tests. - Parsed Dependabot YAML and checked the exact workflow exclusion. - `git diff --check` passed. - Live GitHub checks confirmed `.github/**` has code owners, CODEOWNERS has no errors, and the active master ruleset requires code-owner review. This covers the trusted workflow, caller, and CODEOWNERS itself. - The approved organization setting now allows `paperclipai/paperclip/.github/workflows/pr-trusted.yml@refs/heads/master`. All previous references and other runner-group settings remain intact. - Full application typecheck, tests, and build were not repeated locally for this workflow-only change. This PR's CI and review are pending. ## Risks New versions of the trusted workflow take effect for new callers after merge to master. Keep code-owner review and master protection enabled. Existing administrator pull-request bypasses remain unchanged. Third-party action pins, runner routing, and infrastructure are unchanged. Older callers still use their SHA pins; their allowed refs remain in place. ## Model Used OpenAI GPT-6 through Codex performed implementation and orchestration; its exact runtime variant and context-window size were not exposed. OpenAI `gpt-5.6-luna` with high reasoning inspected workflow assumptions and applied the approved runner-group setting. Both used code and tool access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4cc387f907 |
fix(ci): remove npm propagation from cloud readiness (#13456)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud needs a verified image and matching database
migrator before it can deploy a merge.
> - New npm package versions can take minutes to become downloadable
after the package build finishes.
> - The direct producer now publishes signed archives and a complete
dependency lockfile for each master commit.
> - This pull request makes readiness verify those artifacts and removes
the duplicate automatic npm migrator run.
> - Deployment still requires all source checks, exact image identity,
migration compatibility, and pinned dependencies.
## Linked Issues or Issue Description
Refs: #13455, #13454, #13192
**What existing behavior does this improve?**
The time from a master merge to the `Cloud deployable v1` signal.
**Current behavior**
Readiness polls npm metadata for the new DB and shared versions. An
automatic dispatcher also starts a separate npm-only migrator workflow.
A measured source built its packages at 06:22:41 UTC on 2026-09-15, but
both npm archives were not downloadable until 06:31:56 UTC.
**Proposed behavior**
Wait for the successful exact-source direct producer, verify its signed
manifest and all pinned downloads, and publish readiness only after the
existing source and image jobs pass. Keep manual npm migrators and
branch previews available.
**Reason and benefit**
Remove new-version npm propagation from merge-to-deployable time. The
gain depends on whether image building or source verification finishes
later; it is not a fixed subtraction from every run.
## What Changed
- Require a successful producer from the canonical repository, exact
commit, master ref, expected workflow, and approved event.
- Verify the manifest's GitHub attestation with the hosted GitHub CLI.
Enforce the exact source SHA, master workflow identity, and hosted
runner.
- Download and validate both archives and the complete dependency
lockfile after publication succeeds. Reject invalid signatures,
inaccessible objects, corrupt bytes, and source mismatches.
- Remove automatic npm-only migrator dispatch. Retain manual release and
branch-preview publication.
- Document the cloud feature-switch prerequisite and coordinated
rollback.
## Verification
- `node --test .github/scripts/tests/*.test.mjs`: 405 pass.
- Focused readiness, routing, preview, and artifact tests: 249 pass.
- Workflow lint and `git diff --check`: pass.
- `pnpm test:release-registry`: 139 pass after installing this
worktree's dependencies.
- All latest-head GitHub CI checks passed. Greptile is 5/5 with no
unresolved comments.
- Application source is unchanged. Common-source local typecheck and
build passed. The full local application suite has the documented macOS
read-only-directory rename limitation from #13454 (13 failures in two
unchanged suites); Linux CI is the final application gate.
- Live readiness verification of master
|
||
|
|
9ed55f6931 |
fix: allow concurrent agent runs on one OpenAI or xAI subscription connection (#13452)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts local command-line sessions and stores provider credentials through managed connections > - One OpenAI Codex or xAI Grok subscription connection held a credential lease for the full agent run > - A second run then waited for the first run, and concurrent runs could overwrite a newer credential > - The write-back must compare fresh credentials while it holds the row lock > - This pull request removes the run lease, keeps the revocation guard, and bounds Codex timestamps against the host clock > - The benefit is safe concurrent use of one subscription connection with newest-credential selection ## Linked Issues or Issue Description **What happened?** A managed OpenAI Codex or xAI Grok subscription connection held a credential lease for the full agent run. A second run waited for the first run to finish. The write-back gate also rejected any row change before it compared credential freshness. **Expected behavior** Concurrent runs should start on one subscription connection. The server should keep the newest valid credential and reject a credential write after a person revokes the connection. **Steps to reproduce** 1. Start two runs that use one OpenAI or xAI subscription connection. 2. Let both provider tools refresh the credential. 3. Finish the runs in either order. 4. Confirm that the newest valid credential remains in the connection. **Paperclip version or commit** `ec25bf1e4a81d1729a6d7276e6a486587cff4a0b` **Deployment mode** Local dev (`pnpm dev`), built from source. **Agent adapter(s) involved** Codex. The server path also covers xAI Grok subscription connections. **Database mode** Embedded PGlite for local development, and external Postgres for deployments. **Access context** Both board and agent runs can use managed connections. **Additional context** The provider command-line tool refreshes credentials inside the sandbox. The server copies the result back after the run. Two long runs can still refresh one token hours apart, so the provider can reject the second refresh. The server cannot observe that provider call. ## What Changed - Remove the full-run credential lease for OpenAI Codex and xAI Grok subscription connections. - Lock and re-read the connection row before credential write-back. - Accept only a strictly newer credential, while keeping the connection revocation guard. - Reject Codex freshness timestamps more than five minutes ahead of the host clock. - Add tests for both completion orders, xAI cleanup, revocation, the Codex time bound, and agent hiring. - Update the connection and run-log documentation. ## Verification - `server/src/__tests__/ai-connections.test.ts` passes with 42 tests. - `packages/adapters/codex-local/src/server/codex-auth-merge-decision.test.ts` covers the five-minute boundary and the one-millisecond overflow. - `packages/adapters/codex-local/src/server/codex-auth-merge.test.ts` passes. - `server/src/__tests__/agent-hire-ai-connections.test.ts` covers OpenAI and Anthropic. - The project type check reports no new error in changed files. - GitHub Actions must pass on this pull request. ## Risks The write-back now permits concurrent runs, so the provider may reject a later refresh when both long runs use one token. The server keeps the revocation guard and rejects future-dated Codex timestamps. No schema change occurs. ## Model Used OpenAI Codex, GPT-5. Context window and exact deployment build are not exposed in this run. The model used tool calls, code inspection, and test verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
da77a0c28c |
feat(ci): publish immutable cloud migrator artifacts (#13455)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys images and a matching database migrator.
> - New migrator versions must currently become available on npm before
cloud can use them.
> - npm can serve package metadata while the named archive still returns
404.
> - This pull request publishes immutable migrator archives and a
complete dependency lockfile through the existing artifact store.
> - Cloud can install these exact packages without waiting for their new
npm versions.
> - This producer change prepares a separate cloud consumer and
readiness cutover.
## Linked Issues or Issue Description
Refs: #13454
**What happened?**
A recent master run built both packages by 06:22:41 UTC on 2026-09-15.
Both archives became downloadable from npm at 06:31:56 UTC. Fresh
metadata requests did not remove the delay.
**What did you expect to happen?**
Cloud should be able to install the verified migrator as soon as its
package build and artifact upload finish.
**Steps to reproduce**
Compare package build completion, npm publication, version metadata
availability, and tarball download availability for a fresh full commit
SHA.
**Version**
Master commit `08adcc70d5ec45b7ced9619a3dc10c1d1bec397d`.
## What Changed
- Add a master-only workflow that builds the DB and shared archives
without publication credentials.
- Resolve the dependency lockfile from local archives, then pin those
archives to content-addressed URLs.
- Publish the complete bundle to a separate prefix in the existing
S3/CloudFront artifact store. Write the commit manifest last and verify
public downloads.
- Add a dedicated OIDC role policy. Only canonical master can assume it.
Writes require `If-None-Match: *`; the role cannot overwrite or delete
objects.
- Add source, integrity, lockfile, publication, and real npm install
tests. Document the format and staged rollout.
- Attest the validated manifest with GitHub/Sigstore before S3
publication. The signature binds every package and lockfile hash to the
exact master workflow and source commit.
## Verification
- `node --test scripts/cloud-migrator-artifacts.test.mjs`: 7 tests pass,
including real `npm ci` with an empty cache and no new-version metadata
lookup.
- `pnpm test:release-registry`: 136 tests pass.
- `actionlint .github/workflows/cloud-migrator-artifacts.yml` and `git
diff --check`: pass.
- Ran the workflow's filtered install and package build against the
exact master source. Built and validated the dependency lockfile from
those real archives.
- Latest-head application tests passed, including reruns of two failures
in unchanged application tests. The final CI aggregate passed. The
application source is unchanged. Common-source local typecheck and build
passed; the full local suite has the same documented macOS
read-only-directory rename limitation as #13454 (13 failures in two
unchanged suites).
- The dedicated role and additive bucket read permission are configured.
IAM simulation allows only conditional writes in the intended prefix;
overwrite without the condition, other prefixes, and deletion are
denied.
- Published the verified master
|
||
|
|
058c55bf8f |
fix(ci): overlap preview migrator publication waits (#13454)
## Thinking Path
> - Paperclip manages AI agents and their work.
> - Cloud deploys an application image with a migrator from the same
source commit.
> - The migrator consists of the DB package and its matching shared
package.
> - npm can accept a package several minutes before readers can see it.
> - The publisher waits for shared visibility before it submits the DB
package.
> - This change submits both validated packages before polling their
visibility.
> - Their availability delays can then overlap while the final gate
still requires both packages.
## Linked Issues or Issue Description
Related: #13334. No duplicate open migrator-publication fix was found.
**What happened?**
The cloud migrator publisher waits up to ten minutes for the shared
package before submitting the DB package. The two registry delays
therefore accumulate. A shared visibility timeout can prevent the DB
package from being submitted at all.
**Expected behavior**
Submit both valid packages, then wait for both to pass the existing
exact-source metadata checks. Report which package remains unavailable.
**Steps to reproduce**
Return 404 for shared metadata after npm accepts the shared archive.
Observe that the old publisher never submits the DB archive until shared
becomes visible. The new regression test requires both submissions
before the first visibility sleep.
**Paperclip version**
Base commit
|
||
|
|
c276d3fdc3 |
feat(observability): report terminal run failures to Sentry (#13446)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server records the state of each agent run. > - Terminal run failures need clear error tracking for operators. > - The server did not report terminal `failed` or `timed_out` transitions to Sentry. > - This pull request reports each genuine terminal failure transition with safe diagnostic data. > - The benefit is faster diagnosis without changing run control flow or exposing credentials. ## Linked Issues or Issue Description **What happened?** The server wrote terminal run failures but did not report them to Sentry. Operators could not see these failures in error tracking. **Expected behavior** The server should report each genuine transition to `failed` or `timed_out` to Sentry. **Steps to reproduce** 1. Run an agent task that reaches a terminal failure state. 2. Inspect the Sentry events for the server. 3. Observe that the terminal run failure has no matching Sentry event. **Paperclip version or commit** The change targets the current `master` branch. No public GitHub issue or pull request covers this change. ## What Changed - Add `captureRunFailure()` as a fail-open Sentry entry point. - Add `reportRunFailure()` to filter status, resolve the adapter, redact text, and report the failure. - Call `reportRunFailure()` beside each of the eight terminal status writers. - Report six diagnostic values: the instance host, task identifier, run identifier, error message, error code, and agent adapter. - Group events by error code and agent adapter while keeping the redacted message in the event. - Report only genuine transitions and avoid duplicate finalization events. - Keep Sentry failures outside run control flow. ## Verification - `pnpm vitest run server/src/services/__tests__/run-failure-report.test.ts server/src/__tests__/run-failure-sentry.test.ts server/src/__tests__/native-session-resumption.test.ts` - `pnpm vitest run server/src/services/execution-control-reconciliation.test.ts` - `pnpm --filter @paperclipai/server typecheck` - The full continuous-integration suite must run on this pull request. ## Risks - The report path can add diagnostic events when Sentry is configured. - The report path returns without action when Sentry is not configured. - Redaction runs before length limits and before the event leaves the process. - The change has no migration and no schema change. ## Model Used OpenAI Codex, GPT-5, tool use and code review support. The exact context window and reasoning configuration are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
667c79ded2 |
fix: prevent retry-exhaustion events from exhausting attention-feed memory (#13451)
## Thinking Path > - Paperclip manages AI agents and their work. > - Its attention feed shows failed runs whose retry budget is exhausted. > - Startup retention builds this feed before startup completes. > - The feed joins every exhaustion event to the full run context, then removes duplicate runs in JavaScript. > - Recovery can revisit an exhausted run and append the same event again. This multiplies the data loaded into memory. > - This pull request selects one small row per run in PostgreSQL and makes exhaustion writes idempotent. > - Existing duplicate events can stay in the database without multiplying run contexts in server memory. ## Linked Issues or Issue Description Refs #13367. This fixes the repeated-event allocation path in the retention feed. Other full-feed sources and sweep cadence remain separate concerns. **What happened?** The attention query loaded one full run context for every matching exhaustion event. Deduplication ran only after the driver had loaded those rows. A run with thousands of exhaustion events therefore produced thousands of context copies. The startup retention sweep can exhaust the server heap while reading this result. **Expected behavior** The query should return one row per exhausted run and only the context fields that the feed needs. Repeated checks of the same exhausted retry budget should reuse the original event. **Steps to reproduce** 1. Create a failed run with a 32 KB context and 2,500 matching exhaustion events. 2. Build the attention feed, including dismissed items, as startup retention does. 3. Inspect the database result before JavaScript feed processing. The old join returns 2,500 copies of the run context. 4. Call bounded retry scheduling repeatedly for a run at its retry limit. The old writer appends another exhaustion event on every call. **Paperclip version or commit** The attention query was introduced by #9380 and is present in stable `v2026.831.1`. The retention caller is also present in that stable release. The later recovery callback added by #13075 provides a repeated path into the exhausted-budget writer. This change is based on `c0fda8fac` after rebasing onto current master. Related PR: [#12162](https://github.com/paperclipai/paperclip/pull/12162) changes retention cadence and newer-run suppression queries. This change addresses the exhaustion-event join and duplicate event writes. ## What Changed - Select the newest company-scoped exhaustion event per run with a PostgreSQL `DISTINCT ON` subquery before joining run data. - Project only `issueId` and `taskId` from run context. Preserve JSON types, fallback behavior, run ordering, company filters, and run/agent status filters. - Reuse an exhaustion event for the same run, reason, attempt, and retry limit under the existing run-row lock. Recognize historical events without a migration. - Skip sequence allocation and live publication when a receipt already exists. - Document the run-log behavior and add PostgreSQL regression tests. ## Verification - Focused attention, retry-scheduling, and event-sequencing suites: 72 tests pass on the rebased head `39ce77158`. - The large-history fixture verifies two database result rows under 4 KB, newest-message selection, task-ID fallback, and company/status filtering. This assertion runs before feed deduplication. - Concurrent event tests verify one receipt across two database clients, reuse of historical receipts, distinct reason/attempt/budget keys, and stable event sequences. - Repeated scheduling through a new service instance produces no extra run-log or live events. - `pnpm --filter @paperclipai/server exec tsc --noEmit`: passes. - `pnpm -r typecheck` and `pnpm build`: pass on the rebased head, including native runner checks. - Greptile: 5/5 on `39ce77158`, with no review threads. - [GitHub CI](https://github.com/paperclipai/paperclip/actions/runs/34929002304): all 25 workflow jobs pass, including all server/workspace test shards, serialized server suites, browser tests, typecheck, build, native-runner verification, and the release dry run. Security checks also pass. The branch has no merge conflicts. - `pnpm test:run`: stopped with unrelated failures. Two chat integration cases passed when rerun separately (2 passed, 993 unselected). Three skill-cache cases failed on both this branch and the unmodified parent commit, with `EACCES` during directory rename. The full suite is not reported as green. ## Risks - No schema migration or data cleanup is required. The query still scans matching event history in PostgreSQL; its result size now scales with exhausted runs. - This does not bound every source in the attention feed or change retention scheduling. - An exhaustion receipt is emitted once per retry decision. Consumers that observed repeated copies will now receive one event. - Native source-event replay handling stays on its existing path. - A deployed application boot has not been verified. ## Model Used OpenAI GPT-6, used through Codex with reasoning, repository inspection, code editing, and test execution. The runtime does not expose a more specific model snapshot or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the focused tests locally and they pass (full-suite limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ee81cee76d |
fix: allow concurrent agent runs on one Anthropic subscription connection (#13445)
## Thinking Path > - Paperclip manages AI agents and their work. > - Agent runs use provider connections and subscription credentials. > - File-backed subscriptions need a lease while a run writes new credentials. > - Anthropic subscriptions pass a token and do not write a credential file. > - The current lease blocks concurrent Anthropic runs without protecting shared state. > - This pull request limits the lease to file-backed subscriptions and tests both paths. > - The result lets Anthropic runs share one connection safely while file-backed credentials remain serialized. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change improves concurrent use of one subscription connection. Anthropic runs no longer wait for a lease that protects no shared file. **Subsystem affected** server/ — REST API and orchestration services, with related server tests and connection documentation. **Current behavior** The server takes a connection lease for every subscription invocation. The lease blocks a second Anthropic run while the first runtime remains open. **Proposed behavior** The server takes the lease only when the subscription writes a provider authentication file back to the grant. Anthropic runs can proceed at the same time. **Reason and benefit** Anthropic passes its token through an environment variable and writes no file. Removing this unnecessary wait improves concurrency and preserves credential safety for OpenAI and xAI. **Breaking changes** None. OpenAI and xAI keep the existing lease and retry behavior. ## What Changed - Limit the AI credential lease to subscriptions that write a provider authentication file. - Add coverage for concurrent Anthropic runs and hired-agent runs. - Keep the existing OpenAI contention assertions as a sensitivity control. - Update the AI connection documentation to describe the lease boundary. ## Verification - Run `server/src/__tests__/ai-connections.test.ts`. - Run `server/src/__tests__/agent-hire-ai-connections.test.ts`. - Run `server/src/__tests__/heartbeat-ai-subscription-contention.test.ts`. - Run the server type check. - Review the pull request checks after GitHub starts continuous integration. ## Risks The main risk is an incorrect provider classification. The write-back condition remains unchanged, and OpenAI contention tests keep the file-backed lease behavior under test. No schema or API contract changes occur. ## Model Used OpenAI GPT-5. Exact runtime model ID and context window are not exposed in this execution. The model used tool calls, shell commands, and GitHub operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a2e7ffdc34 |
fix(runtime): validate sandbox paths and preserve live controller leases (#13432)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can run in remote sandboxes. > - Connection checks must use the selected execution target. > - Recovery must respect the controller that owns an active run. > - A host path or PID does not describe a remote sandbox. > - This pull request checks sandbox paths on the target and preserves live controller leases. ## Linked Issues or Issue Description **What happened?** Selecting an AI account for a sandbox agent could fail because the Claude ACP environment check tried to create the sandbox directory on the Paperclip host. The recovery sweep could also interrupt a sandbox run while its controller lease was still valid. It treated a PID absent from the local host as proof that the run had stopped. **Expected behavior** ACP checks directories on the selected execution target. Recovery leaves a run with a live controller lease alone. Its final database write rejects a stale snapshot after renewal, a claim, a controller change, or a runtime change. **Steps to reproduce** 1. Test a Claude ACP sandbox agent with a directory that cannot be created on the host. The check fails before this fix. 2. Give a running sandbox task a valid controller lease and a PID absent from the host. Run the stale-lock sweep without an in-memory handle. The sweep interrupts the run before this fix. 3. Renew or replace the controller between the sweep's read and write. The old snapshot must not end that controller's run. **Paperclip version or commit** Rebased onto `origin/master` at `0e9b24c8216171c26c8358ba387d77858e02c7a9`. All seven regression cases still fail against this base. Refs #13438, which supplies the managed hiring and task-connection behavior, and #13433, which preserves non-assignee subscription comment wakes. This PR preserves both upstream changes and addresses the two remaining sandbox failures. ## What Changed - Resolve and create Claude ACP test directories through the execution-target helpers. - Preserve active legacy controller leases during stale-lock recovery, including finalization after a task becomes terminal. - Recheck the controller, lease, runtime mode, and native ownership in the terminal database write. - Add two sandbox-directory cases and five database-backed controller-lease cases. - Document the target used for ACP directory checks. ## Verification - Red: all seven new cases fail against `0e9b24c82` without these two implementation changes. - Before the final upstream sync, 192 focused tests passed. Full `pnpm test:run` coverage completed using the repository's group/shard runner: all general server and workspace groups passed, and all 147 serialized server suites passed across the initial run and isolated continuations. Five cold-import timeout suites passed with `--experimental.fsModuleCache`; their assertions and deadlines were unchanged. - The hiring routes, default-selection service, and upstream hiring tests match `origin/master` exactly. The two remaining fixes are unchanged by the final rebase. - On final head `bd28d5cefbdf7084acc3759199cd9661907e7a26`, all 372 focused tests pass across 16 suites covering both upstream changes and these fixes. Two timeouts in the combined run (database setup and an existing ACP case) pass in isolated reruns with fresh test homes and temporary directories. `pnpm -r typecheck` and `pnpm build` also pass on this head. - All 32 active checks pass on final head `bd28d5cef`, including the full test matrix and browser shards, in [CI run 34904204849](https://github.com/paperclipai/paperclip/actions/runs/34904204849). Two Storybook checks are skipped by path filters. - Greptile reviewed final head `bd28d5cef` at 5/5 with no findings or unresolved review threads. ## Risks - A failed remote directory check still blocks connection adoption. - A live controller retains finalization authority after its task becomes terminal. Cleanup waits for ownership to expire and must pass the final ownership check. - No schema or credential-storage changes. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository inspection, code execution, and API tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f1905d34d |
fix: provision all project repositories for local and sandbox tasks (#13442)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Projects can now attach several source repositories. > - Task preparation still treated these sources as alternative workspaces. > - Sandbox sync preserved Git history only for the selected repository. > - A task needs every attached repository to complete work across the project. > - This pull request prepares all distinct project repositories and preserves their separate Git histories through sandbox restore. ## Linked Issues or Issue Description **What happened?** A user reported that a project with two repositories received only the first repository in Daytona. Repository-only project rows also reached the agent with null local paths. Managed checkouts with matching repository names could resolve to the same directory. **Expected behavior** Local and sandbox tasks receive every distinct repository attached to their project. Repository-only sources work without preconfigured local folders. Each repository keeps its own Git history and working files. **Steps to reproduce** 1. Create a project with two repository sources and no local folder paths. 2. Assign a task to the project and run it in Daytona. 3. Inspect the task workspace and the repository paths exposed to the agent. 4. Observe that the original implementation supplies only the selected checkout. Related change: #13010 added multiple repository selection. The open repository-catalog proposals #11234 and #11228 cover a different data model. This fix uses the existing project workspaces. ## What Changed - Materialize each additional distinct repository as an editable checkout inside the task root. Seed configured local sources with their current working files and retain task edits across runs. - Pass materialized repository paths to local agents and native sandbox task prompts. Apply existing run-scoped Git credentials to each remote clone. - Preserve each repository's Git history, dirty files, and restore baseline during sandbox staging and durable recovery. Apply each repository's ignore rules and the operator's workspace exclusions. - Keep same-name managed repositories in separate directories. Report additional clone failures before the task starts. - Add task-level, checkout, sandbox round-trip, environment-hint, and recovery-descriptor regression coverage. Document checkout and restore behavior. ## Verification - Red: the original implementation fails the sandbox test because the second repository has no Git directory. It also fails the same-name checkout test and both real-database task tests because repository hints have no local path. - Green: focused tests pass for one and two repository-only sources, local source edits, clone failures, per-repository credentials, separate Git histories, ignored files, and recovery from remote or durable seed state. - Live Daytona smoke passed with two disposable repositories through the production provider sync functions. Both repositories arrived with Git history. Commits from both restored locally. Ignored files stayed excluded. The disposable sandbox was deleted. - Passed on final commit `93ab76763`: `pnpm -r typecheck` and `pnpm build`. - Final focused coverage: 254 assertions across the six changed test areas passed across the serial run and an isolated rerun of the existing process-kill timing test. The live Daytona smoke also passed. - The local `pnpm test:run` overlapped source edits and retained stale transformed code. Its first phase reported 12,240 passed assertions, nine failed assertions, three hook failures, and one worker error; later phases did not run locally. This run is not claimed as green. Fresh focused tests verify the changes, and every general/workspace and serialized-server CI shard passes on the final commit. - Final CI is green on `93ab76763`: all test shards, all three browser shards, typecheck, build, runner verification, canary dry run, and security checks. The initial unrelated chat-delivery browser timing failure passed in the final CI run. Optional Storybook visual checks were skipped. - Greptile is 5/5 on the final commit with no unresolved review threads. Its checkout-race finding was reproduced with a failing test, fixed, and rechecked. ## Risks - Additional repositories need disk space and clone time. Access failure for an attached repository stops preparation. - Additional checkouts live under `.paperclip-repositories/` and keep independent histories. Changes stay in those task copies; they do not overwrite configured source folders. - Detached or reconfigured repository copies are retained under `.paperclip-runtime/detached-repositories/`. Sandbox recovery retains per-repository merge baselines. - No database migration, UI contract change, or new credential delegation is required. Referenced projects retain their separate read-only behavior. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and tool use. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e9b24c821 |
fix: resume subscription comment wakes and preserve retry status (#13433)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Concurrent task runs can share a managed AI subscription with one credential lease. > - #13438 added durable retries for busy subscriptions and task-lock checks. > - A comment wake can run for an agent who is not the task assignee. Such a run never owns the task lock, so the new checks suppress its retry. > - This pull request preserves those comment wakes while keeping the lock checks for assignee runs. > - It also shows the subscription wait in task status and preserves the failure count through the database projection. ## Linked Issues or Issue Description Refs #13438. **What happened?** When a subscription is busy, a non-assignee comment wake is cancelled without a successor. Repeated subscription waits also display an inflated attempt count because the execution query omits the preserved failure count. **Expected behavior** An eligible comment wake waits and resumes without claiming the assignee's task lock. The task shows “Waiting for AI subscription”. Waiting does not consume provider-failure retries. Assignee retries still stop when the task lock is cleared or transferred. **Steps to reproduce** 1. Configure an agent with a managed subscription and concurrent runs. 2. Hold its credential lease in one run. 3. Mention the agent on a task assigned to another actor. 4. Release the lease and inspect whether the comment wake has a scheduled retry. The test uses a real embedded Postgres database and a real credential lease. Provider execution uses a fixture. The reassignment test applies a database mutation after the real checkout. No live provider account is needed. **Paperclip version or commit** Based on `f912ecaac` from #13438. Before the reconciliation fix, the rebased regression at `2063cfd13` failed because the comment wake had no scheduled retry. The original configuration failure was reproduced on `5282cabde` before #13438 merged. **Deployment mode** Built from source with embedded Postgres for local verification. ## What Changed - Record non-assignee comment-wake authority at admission while holding the task and run locks. A later reassignment cannot grant this exception. Preserve it through repeated subscription waits. - Keep master’s lock checks for assignee retries, including the check inside the scheduling transaction. - Show the subscription wait and preserve its failure count in the execution projection query. - Limit pre-provider wait receipts to fresh executions. A persisted native execution input must retain its recovery path. - Add lease, retry, projection, cancellation, reassignment, pause, revocation, service-recreation, and ownership regression coverage. - Use master’s 60–120 second retry interval and document the resulting behavior. ## Verification - Red: after rebase, the comment-wake test failed with no scheduled retry; the other 50 contention and retry-scheduling tests passed. - Red/green: reassigning the task immediately after real checkout reproduced an unwanted successor for both assignment and comment wakes. Both tests pass after recording authority at admission. - Green: all 285 targeted tests pass across 12 suites, including #13438’s four cancellation-race phases, its hiring/connection suite, run-dispatch integration tests, and both reassignment regressions. - `pnpm -r typecheck` passes after rebase. - `pnpm build` passes on the final revision. - The full general and serialized test suites pass in [CI run 34899514681](https://github.com/paperclipai/paperclip/actions/runs/34899514681) on `99f1e4c1d`. Full-suite verification ran in CI; local verification used the 285 targeted tests. - All 32 active PR checks pass, including build, typecheck, native runner verification, canary dry run, and all browser shards. Two Storybook checks are skipped by path filters. - Greptile reviewed `99f1e4c1d` at 5/5 with no outstanding findings. All review threads are resolved. - `git diff --check` and `node scripts/check-module-boundaries.mjs` pass. ## Risks - The non-assignee exception requires authority recorded under admission locks or its server-created subscription retry. Tests cover caller-supplied flags, assignment and comment reassignment races, repeated comment waits, and assignee lock protections. - A wait uses master’s 60–120 second interval. A long-held lease can produce multiple wait records. - Existing native sessions can have provider effects. They must not receive a fresh-execution receipt that permits replay. - No schema migration or credential permission change is required. ## Model Used OpenAI Codex, model `gpt-6-astra`, with repository search, code editing, shell execution, and test tools. The session does not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f912ecaacf |
fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents hire other agents and assign tasks to them. > - Managed AI connections must follow those hires across legacy and native runners. > - Missing accounts should pause task execution and let the user connect from the task. > - Subscription contention must wait without asking for new credentials. > - This pull request fixes these paths and the native tool and Daytona staging failures found during live tests. > - The result is a working hire, subtask, and connection setup flow on local and remote runners. ## Linked Issues or Issue Description **What happened?** A managed Claude or Codex agent could hire a teammate without a usable AI binding. Cross-provider hiring could fail before the user had a chance to connect the new provider. First-time task setup did not show the existing AI credential form inline. A busy subscription could request a new connection. Native API replies could stop the parent after a hire had already committed. Fresh Daytona sandboxes could fail to extract read-only skill directories created on macOS. **Expected behavior** Compatible hires inherit the managed connection choice. A hire for another provider uses the responsible user's default. If that account is missing, the hire succeeds and the task asks for a connection. Completing setup in the task resumes work automatically. Explicit child auth settings and existing unmanaged login paths keep precedence. Shared-account access checks remain in force. **Steps to reproduce** 1. Connect a Claude or Codex parent with a managed AI account. 2. Ask it to hire one agent of each provider and create a self-assigned subtask. 3. Assign work to both hires without connecting the second provider first. 4. Connect the missing provider from its task card. 5. Check that all tasks finish and same-provider work uses the original account. 6. Repeat with native runners and fresh Daytona sandboxes. The opt-in browser suite in `tests/hiring-ai-connections/README.md` performs these steps. **Paperclip version or commit** The live failures were reproduced from `f2c5e54dc`. The branch is rebased onto `5282cabde`. **Deployment mode** Isolated local development instance. Legacy CLI and native runners. Local execution and ephemeral Daytona sandboxes. Related work: Refs #13247 for managed AI connections. Refs #13268 for legacy credential-reference inheritance, which this branch preserves. Refs #13432 for a concurrent managed-inheritance fix. This PR also covers cross-provider task setup, subscription waits, native API replies, and Daytona extraction. It permits missing responsible-user defaults at hire time; restricted shared selections still fail. ## What Changed - Apply managed connection defaults to both agent creation routes. Preserve explicit auth choices and legacy credential-reference inheritance. - Allow hires before their responsible user connects the provider. Keep approval gates, company boundaries, and shared-account access checks. - Reuse the production AI credential form inside the pending task card. Resume the task after setup. - Retry subscription lease contention without consuming the provider-failure allowance or creating a connection request. - Require task execution-lock ownership when scheduling, promoting, and dispatching subscription retries. Recheck ownership under the issue row lock. - Rename the HTTP operation identity at the native tool boundary so it cannot override the runner's operation identity. - Delay directory permission restoration during Daytona extraction. Preserve the final read-only modes. - Add database-backed regressions, real browser acceptance tests, and Storybook states. Document setup and run-log behavior. ## Verification - Six real browser scenarios passed: both parent providers on legacy local and legacy Daytona; native Codex locally; native Claude on Daytona. Each scenario hires both providers, completes a self-subtask and assigned work, and connects the missing provider inline with automatic continuation. - Successful runs verify the account, responsible user, runner mode, and Daytona lease. All 18 test sandboxes were deleted. - Live authentication used API keys. Subscription inheritance, lease contention, and retry have integration coverage. Fresh subscription OAuth sign-in was not automated. - Red/green tests reproduced missing bindings, missing inline forms, subscription contention, native API reply failure, and GNU tar permission failure. - Seven Storybook browser checks passed. They cover both providers, method selection, narrow layout, completion, cancellation, and invalid credentials. - Full local suite coverage completed before rebase. Initial timing and fixture startup failures passed unchanged on isolated reruns. The first full command did not exit cleanly; the remaining workspace and serialized groups were completed separately. - After rebase, 107 hiring/auth/retry tests and 59 native API, task-card, and Daytona tests passed. The full workspace typecheck, production build, and token gates passed again. Storybook build passed before rebase. - Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and four explicit cancellation-race cases passed. Eight cross-provider cases cover stale auth keys on both creation routes and both runner types. Server typecheck and build passed. - Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks passed. Two optional Storybook jobs were skipped. The full server, workspace, browser, native runner, build, typecheck, and release checks passed. - Three unchanged tests initially failed on a busy port, a chat row-lock race, and preview-server readiness. Each affected job passed after one CI rerun. Isolated local checks also passed: 41 credential tests, the chat-concurrency case, and 25 preview-runtime tests. - Greptile reviewed the final commit at 5/5. Both review threads are resolved. GitHub reports no merge conflicts. ## Risks - A missing personal account now defers authentication to the first task. Explicit incompatible bindings and restricted shared accounts still fail at hire time. - An inherited personal default uses the responsible user's existing authorization to install access for the new agent. It never copies credentials or another user's identity. - Subscription contention retries after a delay and rechecks task eligibility. It does not consume the provider-failure budget. - Native hiring uses the existing managed API-tools opt-in. Remote native runners require a matching Linux binary and provider pack, as documented in the acceptance README. - No schema changes. Live tests make paid provider calls and remain opt-in. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository inspection, code execution, browser automation, and API tools. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5282cabde8 |
fix(ui): make every agent reachable from Chats (#13420)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Experimental Agent Chat gives each person a persistent conversation
with each agent.
> - The sidebar only showed starred or previously visited agents.
> - A person could not start a conversation with an agent absent from
that list.
> - This pull request adds a searchable agent picker and a default
sidebar entry.
> - People can now reach every company agent while keeping their
frequent conversations close.
## Linked Issues or Issue Description
Refs #13283, which introduced the existing task-backed agent chat
surface. This change adds discovery to that surface after maintainer
review of the Storybook design.
**What happened?**
With Agent Chat enabled, a person with no starred or recent agents had
no agent shortcut. The sidebar also lacked a direct way to start a chat
with another agent.
**Expected behavior**
Show the first-created company agent by default. Let the person search
all company agents and open a persistent conversation.
**Steps to reproduce**
1. Enable Agent Chat in a company with agents.
2. Use an account with no starred agents or recent chats.
3. Look for a chat entry in the sidebar, or try to chat with an agent
absent from its shortcuts.
**Paperclip version or commit**
The existing behavior is present on master at
|
||
|
|
728f7185f6 |
feat: add native in-app announcements with persistent dismissal (#13403)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Self-hosted boards need a way to show occasional product announcements. > - An app release should not be required to publish or withdraw a card. > - Native card controls keep publishing consistent; the hero can use a static image or isolated HTML/CSS animation. > - This pull request renders a validated JSON feed with native components. > - It stores dismissals per account on each instance, so a closed card stays closed across companies and browsers. > - Named staging feeds let authors test content before production publication. ## Linked Issues or Issue Description **Subsystem affected** Board application shell, announcement delivery, and user preferences. **Problem or motivation** Operators need a small, optional announcement card. Users need reliable dismissal state. Authors need to test remote content without changing the production feed. **Proposed solution** Add one non-modal AnnouncementWell. Fetch validated JSON and content-addressed media through the instance server. Keep card controls native, with optional sandboxed HTML/CSS animation in the hero. Use stable announcement IDs for dismissal, an explicit empty manifest and quiet 404 handling. Provide a staged publishing helper and isolated test-drive guide. **Alternatives considered** Hosting the entire card as a page would move navigation and dismissal into remote content. This change limits HTML to a scriptless, isolated visual hero and keeps controls native. Browser-only storage would lose dismissals across browsers, so the instance stores account preferences. **Roadmap alignment** ROADMAP.md has no overlapping announcement feature. A GitHub title search found no related announcement pull requests. This work implements a maintainer-requested feature. ## What Changed - Add shared feed types, strict validation of every object, supported routes, expiration and version checks. - Add a board-only current-feed API, constrained media proxy, and idempotent dismissal API. Store the first dismissal and its company audit entry in one transaction. - Cache upstream data for one hour. Use conditional requests, request deduplication, response limits, public destination checks, and a three-second deadline. Treat a remote 404 as an empty feed with a fifteen-minute retry cooldown. - Keep announcement visibility stable when focus moves to browser chrome or another app pane; only tab visibility starts a return check. - Add a responsive native announcement card. Respect onboarding, dialogs and toast placement. Sync pending dismissals across tabs and retry after reconnect or return. - Add idempotent migrations for dismissals and validated publication IDs, design-guide examples, static and animated Storybook examples, and focused tests. The publication registry supports offline retries without accepting caller-invented IDs. - Add HTML/CSS animated heroes with static posters, automatic playback, reduced-motion handling, strict DOMPurify validation, an empty iframe sandbox and CSP that blocks scripts/network resources. - Add validated staging publication, content-addressed assets, an empty production manifest, preview fixtures, and authoring/operator documentation. ## Verification - The preceding implementation passed 98 targeted shared/server/publisher/route/OpenAPI/UI tests and 127 tests including the master rebase. The playback-control removal passes all 21 announcement UI tests, covering the rendered sandbox, fallback, reduced motion, dismissal and slow/stale state lookups. The preceding shared/server tests cover HTML validation and response sandbox headers. - The playback-control removal passes UI typecheck, production UI build, Storybook build and token gates locally. Browser verification confirms the animated card has only its dismiss button and two links, with no page errors. The full canonical CI matrix passed on current head `00e416431edb610861599d50490270bbd0f3c6b6`: 32 successful checks and two optional Storybook deployment checks skipped. This run needed no retries. Greptile reviewed this same head at 5/5 with no outstanding findings. - The local canonical general-server run passed 12,063 tests before reporting embedded-PostgreSQL startup failures in an unrelated fixture. All 31 tests in that fixture passed across isolated retries. The UI group passed 6,219 tests and other workspace groups passed 3,201; two CLI database-startup failures also passed individually. Serialized server suites were verified by the full CI matrix rather than repeating them locally. No source changes were needed for these environment failures. - The real S3/CloudFront staging manifest and both media asset headers were verified. Production remains empty/unpublished. The guide distinguishes the preview host's disabled edge cache from production cache requirements. - In the isolated test-drive, the animation visibly moves without playback controls. A 390×844 browser viewport keeps the card above navigation. Reduced motion makes no animation request. Both themes render correctly and browser page errors are empty. Browser fault injection verified that scripts cannot execute and CSS cannot make network requests; a missing animation leaves its poster and controls. - Refresh leaves the animated card visible. Closing it persists after reload and the API returns null. Earlier live checks verified dismissal across browsers, company-relative CTA navigation, modal deferral/restoration, and new-ID eligibility after restarting the same database. - The deployed empty feed and a real remote 404 return HTTP 200 with null from the board API, with a usable dashboard and no announcement popup or browser warnings. - Authoring documentation covers staging, animated HTML constraints, test-drive, withdrawal, ID reuse and cache-refresh steps. ## Risks - Animation supports self-contained visual HTML/CSS and inline SVG, without JavaScript or external resources. A static image is required. Older builds that do not recognize the optional animation field quietly hide that unsupported feed. - The default feed makes an outbound request from an instance when a board is used. Operators can disable it. Requests contain no account IDs, company data, cookies or interaction events. - Feed publication and withdrawal can take about 65 minutes to reach returning users because of CDN and instance caches. Expiration also removes visible cards locally. - Dismissals follow an account within one instance. No-login instances share the existing local-board identity. Separate installations do not share state. - Both tables are additive. A unique key prevents duplicate dismissals; the transaction prevents duplicate first-dismissal audit entries. The publication registry retains only validated IDs. AGENTS.md and the implementation spec document the required exception to company scope for these instance-level records. - Publication was limited to separate public staging prefixes on the existing preview host. Production remains empty/unpublished. No AWS policies or infrastructure were changed. ## Model Used OpenAI GPT-6 through Codex. The exact runtime model ID and context-window size are not exposed in this session. Capabilities used: reasoning, code editing, shell execution, tests, browser interaction, and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1d57863d2 |
fix: make connection checks and task handoffs reliable (#13404)
Preserve connection-probe outcomes through cleanup, reduce unrelated startup work, and report selected Claude authentication accurately. Make artifact download actions match their labels. Route delegated feedback through its active child, retain accepted messages across completion, and avoid redundant worker runs for proven closing notes. Preserve explicit follow-ups, human input, company boundaries, source provenance, and mixed issue references. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2d47a8058d |
fix(apps): update connector artwork and theme fallback (#13361)
> Awaiting author review. Do not merge until the author explicitly approves. ## Thinking Path > - Paperclip helps people manage AI agents for work. > - Connector screens need recognizable app artwork. > - Several bundled marks have inconsistent artwork or dark-theme behavior. > - The shared logo component should retain the existing frame. > - This change replaces selected artwork and fixes local theme fallback. > - Users see consistent connector icons across shared component callers. ## Linked Issues or Issue Description **Current behavior** Some connector marks use outdated artwork or unsuitable theme variants. A remote dark logo can override a canonical local mark that works in both themes. **Proposed behavior** Use the selected bundled artwork with the existing gray rounded frame. Use local light artwork in both themes unless a distinct local dark variant exists. **Subsystem affected** Connector artwork, the shared AppLogo resolver, and its validation. **Breaking changes** No connector capability, permission, credential, or catalog activation changes. No duplicate PR was found in the earlier search. ## What Changed - Update 46 artwork files and only the app definitions whose logo paths need to change. - Keep a compact public manifest of identities, paths, visibility, and aliases. - Prefer local artwork in both themes and preserve the existing frame and padding. - Add locally runnable artwork safety checks and a light/dark Storybook gallery; leave PR workflows unchanged. - Document artwork conventions. Source research is kept outside the public manifest. ## Verification - Artwork check: 69 identities pass. - Node artwork validation tests: 13 pass (rechecked September 14). - Focused resolver, component, and catalog tests: 40 pass (rechecked September 14). - Maintainer decision: dedicated icon-validation CI is not required; the workflow remains unchanged. Greptile acknowledged 5/5 with no remaining code concerns on September 14. - UI typecheck, token gates, and Storybook build pass. - Earlier full build and repository typecheck passed. The broad local test run was stopped after workspace-runtime dependency fixture failures outside this change. Current GitHub CI remains the full-suite gate. - Visual approval remains outstanding. Review the canonical icon registry in both themes at 24–48px. ## Risks The SVG checker rejects common active features; it is not a general sanitizer for arbitrary uploads. Optical balance still requires human review. Remote fallback for unknown brands retains existing behavior. This change does not add an asset importer or new connector capabilities. ## Model Used OpenAI Codex, assisted by GPT-5 and GPT-6 with code execution and browser tooling. Exact hosted model IDs and context window sizes were not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Artwork comparison Before and after for every affected identity. Images follow your GitHub light/dark theme and are pinned to the base and PR commits. This compares artwork; the existing gray rounded product frame and padding are unchanged. | Connector | Before | After | |---|:---:|:---:| | AgentMail | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail.svg" alt="AgentMail" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg" alt="AgentMail" width="48" height="48"></picture> | | Airtable | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg" alt="Airtable" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg" alt="Airtable" width="48" height="48"></picture> | | Asana | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg" alt="Asana" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg" alt="Asana" width="48" height="48"></picture> | | ClickHouse | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg" alt="ClickHouse" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse.svg" alt="ClickHouse" width="48" height="48"></picture> | | Cloudflare | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg" alt="Cloudflare" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg" alt="Cloudflare" width="48" height="48"></picture> | | Cloudinary | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary.svg" alt="Cloudinary" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary.svg" alt="Cloudinary" width="48" height="48"></picture> | | Discord | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg" alt="Discord" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord.svg" alt="Discord" width="48" height="48"></picture> | | GitHub | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github.svg" alt="GitHub" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github.svg" alt="GitHub" width="48" height="48"></picture> | | Gmail | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg" alt="Gmail" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg" alt="Gmail" width="48" height="48"></picture> | | Google Calendar | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg" alt="Google Calendar" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg" alt="Google Calendar" width="48" height="48"></picture> | | Google Chat | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg" alt="Google Chat" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg" alt="Google Chat" width="48" height="48"></picture> | | Google Docs | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg" alt="Google Docs" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg" alt="Google Docs" width="48" height="48"></picture> | | Google Drive | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg" alt="Google Drive" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg" alt="Google Drive" width="48" height="48"></picture> | | Google People | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg" alt="Google People" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg" alt="Google People" width="48" height="48"></picture> | | Google Sheets | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg" alt="Google Sheets" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg" alt="Google Sheets" width="48" height="48"></picture> | | Google Slides | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg" alt="Google Slides" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg" alt="Google Slides" width="48" height="48"></picture> | | Google Workspace Search | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg" alt="Google Workspace Search" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg" alt="Google Workspace Search" width="48" height="48"></picture> | | Hugging Face | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg" alt="Hugging Face" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg" alt="Hugging Face" width="48" height="48"></picture> | | Jam.dev (library only) | New library entry | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev.svg" alt="Jam.dev" width="48" height="48"></picture> | | Jira | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira.svg" alt="Jira" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg" alt="Jira" width="48" height="48"></picture> | | Linear | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg" alt="Linear" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear.svg" alt="Linear" width="48" height="48"></picture> | | Manufact | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg" alt="Manufact" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact.svg" alt="Manufact" width="48" height="48"></picture> | | Microsoft Teams | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg" alt="Microsoft Teams" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg" alt="Microsoft Teams" width="48" height="48"></picture> | | Miro | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro.svg" alt="Miro" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg" alt="Miro" width="48" height="48"></picture> | | Mixpanel | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg" alt="Mixpanel" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel.svg" alt="Mixpanel" width="48" height="48"></picture> | | Netlify | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg" alt="Netlify" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify.svg" alt="Netlify" width="48" height="48"></picture> | | Notion | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion.svg" alt="Notion" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg" alt="Notion" width="48" height="48"></picture> | | PagerDuty | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg" alt="PagerDuty" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg" alt="PagerDuty" width="48" height="48"></picture> | | PostHog | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog.svg" alt="PostHog" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog.svg" alt="PostHog" width="48" height="48"></picture> | | Postman | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg" alt="Postman" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg" alt="Postman" width="48" height="48"></picture> | | Shopify | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg" alt="Shopify" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg" alt="Shopify" width="48" height="48"></picture> | | Slack | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png" alt="Slack" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg" alt="Slack" width="48" height="48"></picture> | | Stripe | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg" alt="Stripe" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg" alt="Stripe" width="48" height="48"></picture> | | Supabase | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg" alt="Supabase" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg" alt="Supabase" width="48" height="48"></picture> | | Telegram | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg" alt="Telegram" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg" alt="Telegram" width="48" height="48"></picture> | | Todoist | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg" alt="Todoist" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg" alt="Todoist" width="48" height="48"></picture> | | Wix | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg" alt="Wix" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix.svg" alt="Wix" width="48" height="48"></picture> | | Zapier | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg" alt="Zapier" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg" alt="Zapier" width="48" height="48"></picture> | |
||
|
|
4cd7b40255 |
fix: recover new messages after historical native runs stop (#13405)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native recovery must distinguish a fresh user request from replay of
failed work.
> - Older runs can lose their process fields before a local stop receipt
exists.
> - A suspended durable session can still prove that the exact runner
and provider session are idle.
> - The message admission path ignored that evidence and kept new user
messages blocked.
> - This pull request uses the existing exact-state verifier for those
historical runs.
> - The user can start one fresh turn while the old history and unknown
outcomes remain intact.
## Linked Issues or Issue Description
**What happened?**
A user sent a new message after a native run exhausted recovery.
Paperclip saved the message but said the previous run had no verified
stop record. The old runner was suspended, with no active provider turn
or pending output. Its process fields had been cleared before stop
receipts were added.
**Expected behavior**
A new user message starts a fresh turn when the exact retained session
proves it is suspended and the other execution gates pass.
**Steps to reproduce**
1. Retain a failed native run with a terminal controller, cleared
process fields, and no process receipt events.
2. Retain its exact suspended runner state and idle provider state. Keep
its recovery hold.
3. Send a new user comment. Before this fix, admission returns no
successor.
**Paperclip version or commit**
Reproduced on master at
|
||
|
|
5d459199b0 |
feat(apps): add governed Railway operations and container access
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d351e08dee |
fix(ui): stop Recent Tasks storage feedback across tabs (#13402)
Version recent-task snapshots separately from activity, reject stale cache updates, and publish only on query or membership changes. Migrate history into an isolated storage namespace while preserving restart retries. Add regression coverage for conflicting caches, migration, comments, unchanged writes, and storage implementations on Node 24 and Node 26. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f2c5e54dca |
fix(runner): preserve handoff work and publish requested files (#13355)
## Thinking Path > - Paperclip lets people manage AI agents and their tasks. > - A task keeps its instructions, progress, and files when its assigned agent changes. > - The replacement runner lost the interrupted run's context and could overwrite an existing draft. > - A saved message also stayed attached to the former agent and could reopen the task after the replacement finished. > - File tasks could report Done with only a local path that the user could not open. > - This pull request transfers handoff context and saved messages, and makes requested files accessible through the existing attachment contract. > - Users can change agents and collect completed work without repeating instructions or confirming bookkeeping. ## Linked Issues or Issue Description Refs #13338. Builds on merged #13354 for queue admission and #13353 for remote workspace retry. #10123 concerns restricted recovery-model escalation; this change instead covers ordinary native handoff and file completion. **What happened?** Codex wrote a draft before a user assigned the task to Claude. The replacement lacked continuation context and replaced the draft. A queued user message could later restart the former agent and reopen the completed task. Separately, a runner could finish a requested file but return only a machine-local path. Remote native runs had no bound file publication tool. **Expected behavior** The replacement reads and preserves existing work, receives saved messages once, and keeps each message's author. The former agent stays stopped. A requested file has a working attachment or accessible work product before Done. Text-only tasks do not require attachments. **Steps to reproduce** 1. Ask Codex to save three newsletter names and then wait. 2. Queue an instruction to keep those names and expand the draft. 3. Use Interrupt and assign to select Claude. 4. Verify the original names survive, the result has a working download, and only the source and replacement runs exist. 5. Ask either provider for a Markdown checklist and open the file from its completed response. ## What Changed - Carry the exact same-task interrupted run's summary, semantic receipts, and history into handoff context. Tell the replacement to inspect existing files before editing. - Adopt saved ordinary task comments into the successor's receipt under the task lock. Preserve authors and separate mention, chat, and interaction contracts. - Prevent a former-assignee comment wake from reopening a completed task or starting a stale execution. - Reject workspace-only, fabricated, and cross-task file completion references with actionable runner feedback. New file output also needs a matching current-run publication receipt and asset filename/size/hash, or an accessible work product registered by the current run. Prior output can remain context alongside a current file, or be verified and re-registered internally. Authorized chat attachment reuse retains its verified current-run clone receipt; older receipt shapes require an intact matching source. - Bind remote file reads to the active environment runner and reuse the existing attachment and work-product publication path. - Enforce workspace confinement, regular single-link files, stable identity, a 10 MiB limit, and exact size and SHA-256 checks. Rotate the native session fingerprint for the updated tool contract. - Contain rejected remote signals and protocol-failure cleanup, including logging failures. Preserve the original cleanup rejection for its owner; a rejected operation never supplies stop acknowledgement or cleanup proof. - Allow exactly one maximum-size base64 file through the native SSH command adapter, preserving a finite output cap. - Document handoff and accessible file completion rules. ## Verification - Each observed bug has a failing regression before its fix. Final post-rebase integration passed 732 tests across 13 files before the final receipt and signal guards; final affected results are below. - Publication provenance and compatibility: 8 provenance regressions and 2 compatibility regressions failed before their fixes; the final four affected suites pass 53 tests, including mixed old/new references and real authorized chat reuse. Controls cover old attachments and work products, filename/size/hash/origin mismatch, missing/wrong receipts, current-run publication, same-run durable proof, internally re-registering preserved bytes, and no-new-file follow-ups. - Remote signal rejection: the real Node subprocess previously exited 1 when the production launcher signalled a deleted sandbox. It now stays alive for both a failed signal and failed logging; all 349 executor tests pass. The failed signal still provides no termination proof. - Remote file reader and SSH command boundary: 33 tests passed, including real Linux descriptor reads and the actual SSH adapter subprocess output cap (network executable replaced by a deterministic fixture). Exact 10 MiB bytes pass, one byte beyond the encoded cap fails. - Live local Claude and Codex Stop journeys preserve the saved file, deliver queued instructions once, and reach Done with two total runs. The handoff journey preserves the original names and download with exactly two runs. Both local providers deliver exact checklist files without a completion confirmation. - Live combined Daytona verification passed: the original failed task's Retry reused its sandbox; a selected Git subfolder produced an exact downloadable file; warm and deliberately resumed Claude runs took about 33 seconds. Codex produced a 240-byte download in 33.1 seconds after 134 seconds of contention/backoff. Both cloud downloads retained exact bytes after the two owned sandboxes were deleted. - The sandbox-deletion retest identified a separate ignored promise in protocol-failure cleanup. Two real Node subprocess regressions failed under fatal unhandled-rejection policy before the fix; all 36 protocol, lifecycle, and integrity tests now pass. The original close promise still rejects to its owning runtime. The final live retest passed: a normal Claude Daytona task completed in 132.352 seconds, then its sandbox was deleted. Thirteen samples over 361 seconds confirmed the same controller stayed healthy, the task stayed Done with unchanged run IDs, and its attachment retained exact bytes. The post-deletion browser download passed with zero page errors; all five owned sandboxes are confirmed absent. - Full repository typecheck and build passed on final commit `de64d16f1`. Final-head Greptile is 5/5 with no unresolved threads. [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/34736623758) passed: 32 successful checks and two conditional skips. The earlier mixed-source full local test invocation was deliberately stopped before rebase, so no pristine green full local aggregate is claimed. Its known failures passed in later affected suites. ## Risks - Handoff may adopt only ordinary comments from its validated former owner. Other delivery contracts must remain independent. - File verification fails closed if a remote file changes during reading. The runner must retry publication or explain a blocker. - The updated session fingerprint starts a fresh provider process where needed to install the new tool contract. - No schema migration or historical status reconciliation is included. ## Model Used OpenAI `gpt-6-astra` through Codex, with reasoning, code execution, browser testing, and tool use. The context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
827ba8a434 |
fix: cancel stalled sandbox startup without waiting for setup (#13352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox runs prepare credentials and files before an agent starts. > - Stop must work during that preparation. > - ACPX registered cancellation, but did not handle it while a setup command was waiting. > - Daytona cleanup waited for that same command before stopping the sandbox. > - This change stops the run's sandbox first and requires proof before abandoning setup. ## Linked Issues or Issue Description **What happened?** Stop left a sandbox run active when remote credential setup stalled. The run kept its connection lease until the sandbox was stopped separately. **Expected behavior** Stop terminates the selected run's sandbox, prevents later setup from launching the agent, and lets run cleanup finish. It must not report success without proof from the provider. **Steps to reproduce** 1. Start an ACPX agent in Daytona. 2. Hold a command during remote credential or file setup. 3. Select Stop before the agent starts. 4. Before this fix, cleanup waits for the held command and never reaches sandbox stop. Related: #13351 exposed this during connection acceptance testing. #12150 addresses scheduler load and session initialization limits, a separate startup problem. ## What Changed - Handle cancellation during ACPX sandbox preparation with a host-owned stop callback. - Pass an explicit active-work cancellation flag through environment cleanup. - Stop Daytona before draining setup commands. Keep normal graceful cleanup. - Require an exact run and lease termination receipt. Keep ownership of outstanding requests when stop cannot be verified. - Reject late setup work and defer sandbox resume until old requests settle. - Add regression tests and document the cancellation boundary. ## Verification - Red: both the stalled ACPX setup test and the Daytona cancellation test failed before the fix because Stop never reached the provider. - Green: adapter and Daytona suites passed, along with cancellation boundary and database-backed receipt/isolation tests. - Two real Daytona probes ran a five-minute setup command. Cancellation returned matching stopped receipts in 6.01 and 6.533 seconds. No later setup command ran. Both test sandboxes were deleted. - Repository typecheck and build passed locally. The full required CI suite passed, including all server/workspace tests, browser shards, Paperclip Runner verification, and canary packaging. The duplicate local full-suite run was interrupted after CI passed; it is not claimed as a completed local pass. - No UI changes. The live probe uses the actual adapter cancellation boundary and Daytona plugin; it is not a browser acceptance test. ## Risks - Provider stop failures remain unacknowledged. The adapter keeps ownership while its original requests remain active. - A retry can receive a settling-work error until old provider requests finish. - Cancelled sandboxes are stopped and retained under their existing provider expiry policy. - Local execution and cancellation after the agent turn starts keep their current behavior. - Other providers must return an exact termination receipt to permit early setup cancellation. No database migration. ## Model Used OpenAI GPT-6 through Codex, with code execution and tool use. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c9e3bb7ca4 |
fix: preserve queued work after native Stop and honor steering support (#13354)
## Thinking Path > - Paperclip lets people manage AI agents and their tasks. > - The runner owns execution, while the task keeps user instructions and status. > - Stop must stop the current response without losing instructions that the user already sent. > - The queue stored its original reason inside saved context. Recovery checked the outer deferred reason and left the message waiting. > - Claude also exposed Steer through a shared method even though its driver did not support it. A rejected request could remove its own error row. > - This pull request keeps queued work until execution has stopped, uses the driver's real capability, and preserves completion event order. > - Users can continue work without repairing task state or repeating messages. ## Linked Issues or Issue Description Refs #13338. Related recovery work: #13353. **What happened?** A message sent during a native run stayed queued after Stop. Claude exposed an unsupported Steer action. A steering failure could hide the queued row and its error. A terminal event could also precede the final provider result, and subtree Stop omitted the board actor. **Expected behavior** Stop ends the current execution. Once Paperclip proves that execution has stopped, it delivers the saved instruction once through normal admission. Pause and recovery holds still prevent dispatch. Unsupported controls stay disabled, and a rejected action leaves an actionable error visible. Final results precede terminal events. **Steps to reproduce** 1. Start a Claude or Codex task that writes a file and then waits. 2. Send a follow-up instruction while it runs. 3. Press Stop. Check that the queued instruction runs once and preserves the file. 4. Check Claude's Steer control and simulate a server rejection on the only queued message. ## What Changed - Recover saved native comments after acknowledged Stop using their original wake reason. - Require durable remote termination receipts or verified local process termination before dispatch. - Preserve actor identity, queued-message deduplication, Pause, and recovery gates. - Derive steering support from the driver descriptor and reject unsupported calls. - Keep the queue mounted until a steering request succeeds so its error remains visible. - Emit provider results before terminal events and pass the board actor into subtree Stop. - Document Stop and steering behavior. ## Verification - Final-head continuation suite: 104 passed, after failing regressions for saved wake reasons, cleanup proof, and deduplication of every queued message. Steering UI: 121 passed. Driver capability: 26 passed. Runner backend/transport: 205 passed; Rust library: 285 passed. - Real Claude and Codex browser journeys both preserved the saved file, delivered the queued instruction once after Stop, and reached Done with exactly two total runs. The process Stop browser fixture also passed. - Local full repository typecheck and build passed on `afaa35139`; server typecheck and the affected 104-test suite passed after the final queue changes. Token gates passed. Final-head CI verifies the complete integrated source. - Local aggregate evidence has explicit limits: the general-server invocation overlapped the queue fixes and finished with 11,989 passed, 2 failed, and 80 skipped; both failures are covered by the final 104-test pass. The UI and CLI then passed all 6,184 and 485 tests; the complete 145-file serialized rerun passed all 2470 tests. The shared-package lock fixture passed unchanged on rerun, but the package phase subsequently stopped at an embedded-Postgres bootstrap resource failure. No single pristine green local full aggregate is claimed. - Greptile reviewed `fa66e2bd5` at [5/5 with no unresolved findings](https://github.com/paperclipai/paperclip/pull/13354#issuecomment-5650334242). [Final-head CI completed successfully](https://github.com/paperclipai/paperclip/actions/runs/34733781888/attempts/2): 33 successful checks, 2 conditional skips, including all server, package, UI, browser, runner, typecheck, and build gates. The first attempt hit a preview-readiness/port-collision fixture; its unchanged local control passed 25 tests with 3 skips, and one supported unchanged CI retry passed the affected shard and aggregate gates. ## Risks - Queue recovery must never overlap an old execution. Unknown cleanup state remains blocked. - Driver descriptors are now authoritative for steering; a wrong descriptor disables the action instead of attempting it. - No schema migration or historical status reconciliation is included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, browser testing, and tool use. The exact hosted model ID and context-window size are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6809314a3f |
fix: recover sandbox workspace setup and retries (#13353)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - Sandbox tasks need a workspace and a provider session before work can start. > - A selected Git subfolder is a valid workspace, but it is not a repository fetch source. > - A resumed sandbox must keep one resource identity when the provider fills in its region. > - Workspace reuse does not prove that a provider session has started. > - A recorded continuation must not leave an old failure blocking Retry for a newer attempt. > - This pull request fixes those setup and recovery boundaries while preserving ownership checks. ## Linked Issues or Issue Description **What happened?** A Daytona task failed before provider startup when its project folder was inside a parent Git repository. After the folder was repaired, Retry resumed the same sandbox and verified its workspace sentinel, but workspace preparation reported that the lease was no longer active. A later attempt could fail because the resumed workspace had no provider checkpoint. The old run also kept a recovery projection after the server recorded its explicit successor, hiding Retry for the new failure. **Expected behavior** Sync the selected folder without importing parent files or history. Keep the existing sandbox identity stable. Allow a new provider session only with proof that its exact session has never started. Let the current failed attempt retain Retry once the old recovery has a recorded successor. **Steps to reproduce** 1. Select a subfolder of a Git repository as a project workspace and start a Daytona task. 2. Leave the target region unset. Release its reusable sandbox after setup fails, then resume and realize the workspace. 3. Retry a provider setup that failed before any session directory or checkpoint was created. 4. Record an explicit successor for a native failure, fail that successor during setup, and inspect Retry in the task thread. **Paperclip version or commit** Reproduced against `8d1f0c20a` with new failing regressions before each production fix. **Deployment mode** Local controller with a Daytona environment. Automated tests use deterministic provider fixtures and real filesystem operations. Related: #13338, #13163, #13264, #13349. The open work-folder stack in #13264 includes a broader fresh-session authority change. This patch addresses the independently reproduced startup failure with existing durable bootstrap proof and atomic directory creation. It does not include the work-folder migration or credential changes from that stack. ## What Changed - Classify only the selected Git repository root as a fetch source. Sync subfolders as directories and preserve their enclosing Git ignore rules on upload and restore. Recognize the shared scheduler’s completed non-repository result so ordinary folders still sync; timeout, cancellation, and output-limit failures remain closed. - Remove the placement target from Daytona account cache identity. Keep API endpoint, credential digest, company, environment, driver, and sandbox ID boundaries. - Preserve closed-lease admission until a sentinel-verified resume reopens that same resource. - Show the already-recorded explicit successor of a resolved recovery. Keep the old failure evidence and unresolved holds. No historical status writes occur. - Permit a resumed workspace to create a new session directory only with matching durable identity, zero connections and events, untouched bootstrap commands, no backup, and an absent remote session. Claim the directory atomically. Existing, partial, or ambiguous state still blocks startup. ## Verification - Final head `c43c7403f`: [CI completed successfully](https://github.com/paperclipai/paperclip/actions/runs/34733846820/attempts/2), with 32 successful checks and 2 conditional skips. This includes every server, UI, package, serialized-route, browser, native-runner, typecheck, and build gate. Greptile reviewed the same head at [5/5 with no unresolved findings](https://github.com/paperclipai/paperclip/pull/13353#issuecomment-5650270168). - New regressions failed before each of the four production fixes. Git/archive/restore suites: 148 passed. The scheduler-wrapped non-Git regression also failed before its fix; 135 affected Git/sync/Codex tests then passed. - Daytona plugin: 230 passed, 6 opt-in live tests skipped. Native executor, projection, and TaskChatThread: 494 passed, including existing/partial state, wrong identity, prior connections or turns, backups, unavailable proof, and unresolved recovery controls. - Local full-repository typecheck, build, and token gates passed on the final head. Local CLI: 485 passed. Complete single-worker package rerun: 3,224 passed, 19 skipped. - Local verification is an aggregate with recorded retries, not one pristine green invocation: the general-server run began on `e721a920a` and finished with 12,002 passed, 3 failed, 70 skipped. Its real Codex scheduler failure is fixed above; the socket and workspace-runtime timeout failures passed unchanged in focused reruns. Both Inbox failures passed unchanged in the full 27-test Inbox file; database/shared-package failures passed in the single-worker package rerun. The supplemental local serialized-route rerun remains in progress; all five corresponding final-head CI lanes passed. - The first final-head CI attempt hit a Daytona fixture-readiness race and a signoff-browser heartbeat receipt timeout. One supported unchanged failed-job rerun passed both and the aggregate gates. Live combined user-journey verification is tracked in the related follow-up; this PR's provider tests use deterministic fixtures and real filesystem checks. ## Risks - A selected subfolder uses directory sync and does not carry parent Git history. Its ignored files stay local. - The target region remains a creation setting and part of workspace reuse policy; it does not split the account identity of an existing sandbox. - Incomplete or conflicting provider state still fails closed. This change does not erase a session, infer completed work, bypass a user decision, or replay uncertain actions. - A later remote setup failure can leave a claimed partial session directory. It remains blocked rather than being overwritten. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact hosted model ID and context-window size are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d1f0c20af |
fix: let responsible users choose either AI subscription or API key (#13351)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - AI Connections select the account used for each run. > - A responsible-user binding must follow the person whose work the agent performs. > - The saved sign-in method currently blocks users with another method for the same provider. > - This pull request resolves a personal default by company, user, and provider. > - Each user can use a subscription or API key with the same bot and model. ## Linked Issues or Issue Description Refs #13247, #13248, #13346, #13347. **What happened?** A bot configured with a Claude subscription rejects another responsible user’s Claude API key. Inline repair also limits that person to the original sign-in method. **Expected behavior** The same bot uses each responsible user’s default Claude account, whether it is a subscription or API key. Explicit shared account selections remain fixed. **Steps to reproduce** 1. User A connects a Claude subscription and creates a bot using the responsible user’s connection. 2. User C connects a personal Claude API key. 3. User C runs the same bot. Before this fix, credential resolution fails. ## What Changed - Add a personal provider-default table. Preserve legacy per-method preferences and backfill the most recently updated preference, including unavailable defaults. Repeated migration does not replace a selection. A database trigger propagates old-server default updates without treating new accounts as replacement defaults. - Resolve responsible-user bindings by provider. Retain the method as a wire compatibility hint for old servers. Explicit selections still require the exact method and grant. - Use the selected account’s method for credential isolation, refresh locking, and run attribution. - Update onboarding, agent setup, the picker, and inline task repair. Keep existing authentication components and harness/model settings. - Add mixed-method runtime, migration, repair, and Storybook coverage. Include upstream’s duplicate Anthropic option fix through the base branch. ## Verification - Focused resolver, migration, connection-intent, onboarding, agent setup, model, and connector UI suites: 331 tests passed. - Onboarding and new-agent regression suites passed during the initial focused run. - UI typecheck, token gates, and Storybook build passed. - Live browser checks passed for Claude and Codex API-default execution, switching both back to subscriptions, and both existing shared-account bots. Bot configuration remained unchanged. - One Daytona startup command stalled before Claude launched. The test run was cancelled, its sandbox stopped, and the same account/task passed on retry. Startup cancellation remains a separate environment finding; this PR does not change that command transport. - Browser review: all eight assertions passed in the new mixed-method story, including shared selection, return to responsible-user selection, and unchanged harness/model. - Repository build and typecheck passed after refreshing upstream dependencies. Final resolver and historical rollback verification: 38 tests passed. All latest-head CI gates passed, including server/workspace/serialized suites, browser E2E, build, typecheck, and runner verification. The extra serial local full-suite run was stopped after equivalent CI passed; focused local checks completed. ## Risks - Users with both historical method defaults get their most recently updated preference as the initial provider default. They can change it explicitly in Connections. - A revoked or unavailable default blocks. Connecting an additional account does not silently replace it. - Existing legacy authentication is unchanged. Managed responsible-user bindings intentionally stop pinning a method. - Live staging: the same Claude and Codex bots completed real API-key runs after changing only the personal default, then completed subscription runs after restoring the original defaults. Read-only database verification confirms unchanged bot configuration and actual method attribution. Distinct-user concurrency is covered by automated real-database tests with synthetic credentials, not two live human logins. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, code execution, and browser interaction. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
422287eecd |
fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task messages, provider execution, and task outcomes. > - First-time user tests exposed gaps in recovery, completion permissions, message delivery, and Stop behavior. > - These gaps left usable output hidden, completed work waiting for bookkeeping, or safe work unable to continue. > - This pull request fixes the shared lifecycle and receipt paths while preserving process ownership and action checks. > - Users can continue work with accurate task state and durable messages. ## Linked Issues or Issue Description **What happened?** A stopped local Codex execution could remain blocked even after its processes had stopped and its complete transcript proved that no external action needed replay. Claude under Conservative permissions could fail to call task completion tools. Recovery could reuse an assistant item ID and overwrite prior output. A delivered comment could remain marked uncertain after navigation. Stop could look like Pause or a new recovery incident. Workspace contention could look like cancellation. A direct reply reopening Done could enter a clarification loop. **Expected behavior** Recover automatically only with verified termination and complete action receipts. Preserve answers and messages. Keep task completion available under Conservative permissions without broad tool access. Show crashes as Blocked, actual human decisions as In Review, and ordinary workspace contention as waiting. Stop the current response and allow a new direction. **Steps to reproduce** 1. Create ordinary response tasks with local Codex and Claude Code, then send follow-up messages through the task composer. 2. Interrupt a disposable local Codex runner during text-only work. Verify automatic continuation and retained output. 3. Stop a response, send a new request, answer a clarification, and reopen completed work with another message. 4. Navigate or reload while a comment submission is pending. Confirm the exact persisted request receipt settles it without removing newer draft text. 5. Run two tasks in a shared Daytona workspace. Confirm waiting does not appear as failure. **Paperclip version or commit** Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`. Current integration base: `6cef9743c`. Both operator-interruption and workspace-waiting guards are preserved; native restart and legacy permission rules remain documented. **Deployment mode** An isolated source-built test-drive instance, with real local Codex and Claude Code providers and disposable Daytona environments. Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163. This PR addresses additional failures from ordinary task journeys, including controller restart handoff and repeated warm sandbox setup. Historical task status reconciliation is excluded. ## What Changed - Persist runner ownership immediately at spawn and resume an explicitly adopted runner even when the controller crashed before the first driver checkpoint. Detach the controller safely across graceful restarts, including session startup. Prevent an old finalizer from suspending or signaling an adopted runner. Checkpoint idle warm sessions before shutdown. Preserve the same run and queued follow-up messages. - Scope saved legacy queue successor checks to the queue owner while preserving ordinary task locks, operator identity, assignment gates, and exactly-once delivery. - Preserve managed Codex credential files when an old session is detached for restart; normal owned cleanup still copies refreshed auth back and removes the scoped copy. - Reuse the bound warm shared sandbox and fully verify an existing staged provider pack before using it. This avoids repeated uploads when the pack is already valid. - Add a narrow local Codex replacement path with stopped-process proof, a closed transcript inventory, exact completion receipts, and fresh-session lineage. Preserve no-replay holds when evidence is incomplete. Recovery may clear only the same run's recorded Blocked status version; manual re-blocking and dependency changes invalidate that receipt, while queued comments do not. Later blocks stop scheduled, queued, and final dispatch; queued/final checks re-read dependencies even when the task status stays In Progress. - Permit only task delivery and human-input tools through the isolated Claude runner's exact task bridge. - Scope assistant item identity to the provider turn and ignore only authority-free Codex skill-change notifications during startup. - Reconcile composer submissions by client request ID across response loss, navigation, and reload. Retain text typed during delivery. - Keep acknowledged run-only Stop neutral and show workspace contention as waiting. Project exhausted native failures as Blocked. - Restore the guarded task-page retry action for failed legacy runs, including the server-supported explicit new-attempt path for stopped conversation adapters. Preserve native/process recovery holds and avoid promising Retry while a decision or execution gate hides it. - Refresh delivered artifacts and handle direct user replies that reopen completed work without a clarification loop. - Check the embedded PostgreSQL PID, data directory, and actual port before connecting or migrating. - Document accepted behavior and add focused regressions at lifecycle, route, transcript, and UI boundaries. ## Verification - Final head `fece606ac2` passes the complete GitHub CI matrix: **34 green checks, two expected Storybook skips, no failures or pending checks**, including `ci / verify`, `ci / e2e`, full runner verification, typecheck, build, every server/workspace shard, and all browser shards. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34727183287). Greptile is **5/5 with no open findings**. The final two commits only refine test fixtures; both affected suites pass 24/24 locally and in CI, with server typecheck green. - Complete local Vitest coverage uses the canonical groups/shards: all 635 general server suites, all 145 serialized suites, and all workspace packages. The aggregate began on `0a8001c18` while the final queue fix arrived: 23,903 passed, five failed, 87 skipped. The five port/socket/timing failures passed unchanged in follow-ups (60 tests in the exposure/file suites and 412 tests covering the serialized failures and unrun tails). The final queue/operator-identity suites separately passed 52/52. This is aggregate coverage plus explicit reruns, not a pristine single-command final-head run. - After integration with current master, queue/operator-identity/continuation suites passed 162/162 and affected UI suites passed 140/140. ACP Stop/continuation and legacy task/Inbox/message browser suites passed 9/9, including both task recovery Retry and thread Try again, automatic saved-message delivery, exactly one new run, Done, and retained output after reload. The default process Stop/Pause/Resume browser case passed (the native-provider case is opt-in and skipped by default). The complete Board attachment/receipt browser suite passed 11/11 on a disposable instance, covering both composers, exact receipts after lost responses, no replay, bound attachments, and newer drafts after reload. - Blocking-intent regressions cover pre-existing Blocked, a mismatched run/cause, an explicit manual re-block, changed dependencies, a queued comment after failure, and a block arriving between scheduling and provider dispatch. The negative cases reproduced before the fix. All 478 affected executor/recovery/dispatch tests passed; both database suites ran separately after availability-probe skips in the first combined command. The final late-dependency check passed all 143 affected recovery/dispatch tests (zero skips) after two new negative cases reproduced the bug. - Focused runtime regressions cover awaited runner ownership publication, authenticated adoption before the first checkpoint, old-finalizer detachment, idle and busy warm-session shutdown, rejected checkpoint propagation, provider-pack verification, and managed-Codex credential preservation. Four managed credential detachment cases reproduced the bug before the fix; normal owned cleanup still succeeds exactly once. - Live local Claude: SIGKILL 2.6 seconds into startup recovered the same run automatically in 53 seconds, then a normal follow-up completed in 24 seconds. SIGTERM 2.5 seconds into startup preserved the same run (54 seconds) and its queued follow-up (21 seconds). Answers remained visible and the task reached Done. - Live Claude Daytona: a warm follow-up retained its sandbox and fell from 121 seconds to 44 seconds. A separate cold turn took 127 seconds; after controller shutdown and checkpointing, its follow-up completed in 33 seconds with the same sandbox, workspace, native session, and runner. Both answers remained visible and the task was Done. - Other live journeys covered task completion and follow-up with local and Daytona Codex, local Codex crash recovery, Stop then new direction, clarification response, live artifact refresh, and shared-workspace waiting. - Validation limits: the opt-in native composer Stop/Pause→subtree Resume fixture exposes terminal/result ordering and subtree-cancellation attribution bugs that can leave a child task blocked; that new finding is assigned to a separate follow-up and is not claimed fixed here. Default CI skips this optional native-provider fixture. Managed-Codex credential handoff and the queue-agent integration use automated regression evidence. Cold custom provider-pack uploads still add startup latency. ## Risks - Automatic replacement remains deliberately narrow: local Codex, verified stopped identities, unchanged retained state, and a complete text/completion-only turn. Unknown actions, partial history, or changed ownership remain blocked. - Claude completion permission handling changes an upstream package patch. The exact isolated task bridge must remain pinned; unrelated tools keep their existing permissions. - New task failure projection changes user-visible status. No historical status backfill or database migration is included. - This is a broad lifecycle fix across server and UI. Live proof covers graceful local Claude restart during startup and idle Claude Daytona session recovery across controller shutdown. Live abrupt SIGKILL during local Claude startup also recovered the same run. Unknown ownership or missing action evidence still blocks reuse. Cold custom provider-pack uploads still add startup latency; this change avoids unnecessary repeat uploads. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, browser automation, and tool use. The exact hosted model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
13bae6fa21 |
fix: reject unsupported REST tool connections without stdio validation (#13346)
## Thinking Path > - Paperclip manages AI agents and their connections. > - Connection checks must use the configured transport. > - The tool service treated every remaining transport as local stdio. > - Anthropic's old REST method therefore failed with a templateId error. A REST connection with a valid stdio template could incorrectly pass. > - Anthropic now has a supported AI-account flow. This pull request removes its obsolete REST setup option and limits stdio checks to stdio connections. > - Users can connect an AI account, and existing unsupported connections receive an accurate error. ## Linked Issues or Issue Description Related: #13248 added the supported AI-account flow. Searches for related REST health and templateId bugs found no duplicate fix. **What happened?** The Anthropic REST API-key connection showed `Local stdio MCP connections must use an approved templateId`. Health checks and catalog discovery both fell through to the local stdio path. A REST connection with an approved template could report success and expose the template's catalog without a REST integration. **Expected behavior** Only local stdio connections use command templates. Unsupported transports return an accurate HTTP 422 error. New Anthropic accounts use the supported runtime authentication flow. **Steps to reproduce** 1. Check out the test-only commit `924e6e85a` in a separate worktree and install dependencies. 2. Run `pnpm exec vitest run packages/shared/src/app-definitions.test.ts server/src/__tests__/tool-access-service.test.ts -t 'unsupported REST|obsolete Anthropic'`. 3. The tests exercise saved Anthropic REST configuration and an unsupported REST connection containing an approved stdio template. They cover health checks and catalog discovery separately. 4. Run the same tests on the fix commit. They pass. The full affected files also pass. **Paperclip version or commit** Reproduced against master `6cef9743c`. **Deployment mode** Server transport handling. Reproduced with an isolated embedded PostgreSQL test database. No provider account or live credentials are required. ## What Changed - Restrict stdio health checks and tool discovery to `local_stdio`. - Return and audit `tool_connection_transport_unsupported` with HTTP 422 for unsupported tool transports. - Remove Anthropic's obsolete REST method from the generated catalog and its durable ingestion source. Keep its subscription and API-key AI methods. - Cover the reported error, false-success case, rejected obsolete setup, connection removal, and the UI's AI-account submission path. - Replace impossible reconnect forms for removed methods with supported setup, while preserving connection removal. - Preserve AI-versus-tool intent isolation for legacy requests and reject new unsupported Anthropic tool requests. - Document recovery for existing unsupported connections. ## Verification - Clean-worktree red/green: the same command failed all six regression cases at `924e6e85a` and passed all six at `4d3de9de0`. The failing run includes the reported templateId error. - Green: all 555 tests across the six affected test files passed. - Recovery UI red/green: three added cases failed before the recovery fix and passed afterward; all 200 tests across setup, detail, and advanced controls passed. - After the recovery UI update, UI typecheck/build and token gates passed again. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm check:token-gates` — passed. - Catalog regeneration — passed with the documented `PAPERCLIP_CONTENT_TEMPLATES` override for the local capture corpus. - Full CI on `69fb31fd4` — passed all general and serialized test shards, browser shards, typecheck, build, runner verification, and canary dry run: https://github.com/paperclipai/paperclip/actions/runs/34726975425. - The local serial `pnpm test:run` was stopped after the fixture correction superseded that run; full-suite verification above comes from CI. All 555 affected tests passed locally, including all 17 connection-intent tests after the correction. - Greptile — 5/5, successful check on final commit `69fb31fd4`, no unresolved findings. - No live Anthropic validation was performed. The UI regression uses a fake key and a mocked AI-account response. ## Risks Existing obsolete REST connections remain in needs-attention state. Users must add an account through the supported flow and remove the old connection. Credentials and grants are not transferred automatically. Removal remains covered. The specialized AgentMail and Composio paths keep their existing behavior. There are no schema or permission changes. ## Model Used OpenAI GPT-6 through Codex. The exact serving model ID and context-window capacity are not exposed in this session. Used reasoning, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e704e1c9ae |
fix: keep AI subscription locks safe through transaction pooling (#13347)
## Thinking Path > - Paperclip manages AI agents and their credentials. > - AI Connections must preserve account ownership from login through execution. > - A subscription environment check can pass and leave a session advisory lock on a pooled database backend. > - The next execution then reports that the subscription is busy. > - Onboarding can also lose its completion callback when the connection list refreshes before the login poll. > - This change keeps each credential lease in one transaction and keeps the login controller mounted until completion. ## Linked Issues or Issue Description Refs: #13247, #13248, #13344. **What happened?** A real hosted subscription login saved successfully and passed its environment check. The first task then failed with an AI connection busy error. A connection-list update could also leave onboarding on Connecting. **Expected behavior** A completed environment check releases its credential lease. A saved login completes onboarding without another sign-in. **Steps to reproduce** 1. Run Paperclip with a transaction-pooling PostgreSQL proxy. 2. Connect a Claude subscription through onboarding and reuse it for an agent. 3. Start a task after its environment check succeeds. 4. For the UI race, refresh the managed connection list before the login completion poll. ## What Changed - Hold the grant lock inside one transaction. Roll back on cleanup or failed acquisition. - Disable the lease transaction's idle timeout so long provider executions retain the lock. - Keep the onboarding login controller mounted while its authorization URL is active; restore saved-account retry if later agent creation fails. - Add lock-lifetime and Claude/Codex completion-order regressions. - Document the pooling requirement. ## Verification - 104 focused connection and onboarding tests passed. - Workspace typecheck, production build, Storybook build, and token gates passed. - All 34 latest-head checks are terminal: 32 passed and two Storybook workflows skipped by their normal filters. CI covers all server/workspace/serialized test shards, browser tests, typecheck, build, runner verification, and canary packaging. The additional local full-suite run is still in progress. - Real browser sign-ins saved Claude and OpenAI subscriptions on both new and existing staging stacks. On the patched new stack, both providers completed tasks and immediate repeat executions with the same saved accounts. Read-only database inspection confirmed live transaction-scoped locks and no remaining locks after completion. On the final deployed head, both shared accounts on the existing stack also completed actual tasks. Both personal accounts passed a fresh environment check followed immediately by execution after a deployment restart. ## Risks - Each active subscription reserves one database connection and holds an otherwise idle transaction until cleanup. This is required to pin the backend through a transaction pool. - Old session locks from earlier versions may need operator cleanup after active runs drain. This change does not bypass existing locks. - No database migration, credential routing, or legacy authentication change. ## Model Used OpenAI GPT-6 in Codex, with code execution and browser tools. The exact deployment model ID and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6cef9743c0 |
fix: deliver saved user messages after recovery stops (#13327)
Deliver saved user messages after legacy recovery stops. Validate undelivered comments and queue ownership under the task lock, preserve the operator identity checks from #13315, and prevent duplicate successors. Add a recovery notice with Retry and inline errors, plus service, route, component, and browser coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b2acc674be |
fix: enable isolated subscription login on authenticated self-hosted instances (#13344)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - AI Connections store reusable provider credentials with company and owner boundaries. > - Self-hosted instances can require board authentication while running agents on the local server. > - The login UI treated those instances as unsupported and displayed instructions without a command. > - Removing that restriction must not expose the server operator's existing CLI account. > - This pull request enables owner-scoped login attempts and reuses the existing login UI. > - Users can connect Codex and Claude subscriptions on an authenticated self-hosted instance. ## Linked Issues or Issue Description **What happened?** On an authenticated self-hosted instance, OpenAI subscription setup displayed terminal instructions with no command and a disabled Connect button. Claude also could not complete the local connection flow. **Expected behavior** An authorized board user can prepare an isolated sign-in attempt, sign in on the server, and save the verified account as a Connection. This must not import another user's or the operator's ambient credentials. **Steps to reproduce** Run Paperclip in authenticated mode with a local environment. Open Connections, choose OpenAI or Anthropic, and select Subscription. The previous UI never enabled local login preparation. Related: #13247, #13248, and #10751. This fix preserves the restriction on remote access to the operator's ambient Claude login. ## What Changed - Allow company-authorized users to create, check, cancel, and complete their own isolated local login attempts. - Keep ambient Claude credential import restricted to the local operator. - Support isolated Claude credential files without falling back to the host account or mutating process-wide environment variables. - Use Codex device authorization so sign-in does not depend on a browser callback to the remote server's localhost. - Gate server-host login on authenticated public deployments unless a trusted runtime host is configured. Publish the capability through health so setup shows supported alternatives. - Read isolated Claude credential files through bounded, descriptor-bound opens with ownership, permission, and symlink checks. Try the alternate filename after malformed JSON. - Reuse shared login instructions and lifecycle hooks in onboarding, agent setup, and Connections. Show health-query failures explicitly. - Document authenticated self-hosted behavior and add authorization, isolation, lifecycle, and UI regression tests. ## Verification - Passed 71 focused tests across connection routes, credential isolation, legacy compatibility, the shared login hook, and agent setup. - Passed 89 onboarding regression tests. - Passed `pnpm -r typecheck`, `pnpm build`, Storybook build, and `pnpm check:token-gates`. - Completed real Codex device authorization and Claude browser authorization on an authenticated Linux self-hosted instance. Both accounts were detected automatically and saved as Connected. Both completed attempt directories were removed. - These live checks cover login, credential validation, and connection creation. They do not establish a new model execution or long-running refresh result. - Review follow-up: 68 focused checks passed after rerunning one route socket error; the full route/health rerun passed all 49 tests. The 70-test onboarding suite also passed. Final workspace typecheck, production build, and Storybook build passed again. - The broad local run exposed an instance-name assumption in two new assertions. The fixture now uses an explicit non-default instance, and all 32 connection tests passed with a different inherited instance name. The superseded broad run was stopped; this is not a claim that the full local suite completed. Full CI results will be recorded before merge. ## Risks - The server must have the provider CLI installed. Users still run the displayed command on the server that hosts Paperclip. - Authorization checks must keep login attempts scoped to the company, owner, provider, and reconnect target. Regression tests cover cross-user and cross-company access. - Existing local-trusted Claude behavior stays available. Authenticated remote users cannot use its ambient import path. - No database migration, dependency change, agent binding change, or provider routing change is included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and browser tools. The runtime does not expose a more specific model version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
df984cbc2c |
fix: dispatch queued legacy messages with operator identity and task permissions (#13315)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task conversations save messages that arrive during an active turn. > - Legacy adapters deliver these messages in a later turn. > - A run can stop before the saved queue is delivered. > - The Interrupt button previously required an active run, so it could not release this queue. > - This pull request lets a board operator send the saved queue after the run stops and retries queues missed during finalization. > - Manual dispatch must use the clicking operator and must not require permission to create agents. > - The task can continue without a duplicate message or a second execution owner. ## Linked Issues or Issue Description **What happened?** A legacy task retained a queued message after its run stopped. Interrupt was disabled because the queue had no active target. Finalization and deferred message admission can also leave a queue without a successor. **Expected behavior** Interrupt sends the saved messages when no runner is active. Messages that arrive during normal completion are delivered automatically. An uncertain previous execution still requires proof that its process or sandbox stopped. **Steps to reproduce** 1. Queue a user message during a legacy conversation turn. 2. Let the turn stop or simulate a server restart before queue promotion. 3. Open the task with a deferred queue and no active run. 4. Try Interrupt. Before this change, the button is disabled. **Paperclip version or commit** Reproduced against `8f40b4ad4`. **Deployment mode** Legacy conversation adapter. The same persisted queue state is covered with an isolated PostgreSQL fixture. Related public work: #13275 adds active legacy interruption. #13291 addresses automatic sandbox conversation recovery. This change handles explicit saved-queue delivery and late queue promotion. ## What Changed - Accept a null Interrupt target while retaining queue identity, revision, company, and assignee checks. - Save the operator's request on the existing queue. Reuse normal admission after verified stop, including older messages, different authors, and queues whose original wake came from the system. - Strip interruption authority from caller-supplied wake payloads. Only the board queue route can persist that authority. - Retry durable interruption requests after restart and deferred queues after legacy cleanup. - Let an explicit Interrupt retry cleanup for its stopped run, including old ephemeral leases that recorded success without a provider stop receipt. Preserve retained resources, other lease owners, and the automatic retry limit. - Preserve the server's waiting explanation when normalizing and combining queue entries. - Revalidate the consumed board queue receipt at dispatch so a different message author does not cause setup failure. - Use the Interrupt user's execution identity for the new run. Preserve original message authors. Validate the receipt independently at startup and inherit the resulting identity on retry. - Persist authenticated board authority for ordinary manual wakes too. Adopting someone else's queued messages cannot switch a manual run to that author's permissions. Strip caller-supplied authority markers and retain private conversation ownership checks. - Keep the clicking user when a manual wake is merged into an older deferred receipt. Update its requester and payload in the same transaction. - Use the same current-queue/revision API on task details and pipeline conversations; show Interrupt after a legacy target stops. - Keep manual wakes out of active runs, including unscoped agent wakes. They receive their own execution identity; a matching receipt requester is not sufficient because an exact retry can retain a different originating identity. - Authorize both existing-agent wake endpoints with `agent:wake`, available to active non-viewer company members. Keep `agents:create` for hiring. Validate the stored task and current assignee before an exact task retry. - Reject viewer Interrupt requests before saving intent or stopping execution. Keep external chat retry authorization and per-action agent/user permission checks. - Preserve edits and discards until dispatch. Prevent another queue promotion when the same agent already has a successor. Keep independent reviewer recovery available. - Suppress cancelled/failed run toasts for intentional operator interruption. Keep ordinary runtime error notices. - Add UI, route, admission, restart, successor ownership, and toast regression tests. Document the behavior. - Reuse the existing socket reservation helper for both credential-quorum test cases after CI exposed an ambient-port collision. This changes test preparation only; production credential staging is still called exactly once. ## Verification - Failing regression tests reproduced the message-author identity bug and an operator's `agents:create` rejection before the fixes. - All 316 focused tests pass across eight route, queue, identity, authorization, continuation, and responsible-user suites, including the 44-test rerun of queue admission and actual startup after the final manual-wake restriction. Regressions reproduce cross-user merging both with and without a task, and same-requester receipt ambiguity. The cross-company existence guard also passes both tests. - Startup integration tests reach adapter execution under the clicking operator and retain that identity through follow-up. Coverage includes mixed authors, adopted queues, system-origin queues, restarts, forged or stale receipts, viewers, suspended memberships, changed assignees, private conversations, and caller-supplied authority markers. - The earlier queue/cleanup/UI regression suite passed 402 tests. The final review corrections pass another 180 tests across queue admission/persistence, real heartbeat startup, UI API, conversation rendering, and pipeline suites. Regression tests reproduced both review findings before correction. The final head has a 5/5 review with no unresolved threads. Full CI passes on `c2002979c`, including every general and serialized server shard, all browser shards, Paperclip Runner verification, typecheck, build, canary dry run, and the aggregate gates. - Full `pnpm -r typecheck`, `pnpm build`, and UI token gates pass after the final application changes. CI identified an outdated task-page API mock after the shared helper extraction; the fixture now exercises the real helper, and all 131 task-page/API tests pass. The final application build passes with the additional manual-wake restriction. - CI exposed a pre-existing port collision in the Codex credential-quorum fixture. It reproduced locally; both listener cases now use the existing bounded reservation helper. All 41 credential tests pass on rerun. One intervening local run hit a separate ambient bind collision in the two-occupied-port case. - The full local `pnpm test:run` attempt was stopped after host contention caused focused-suite timeouts. The affected focused tests passed on rerun. An expiring trace fixture and a missing private-conversation state were corrected. The successful full CI run is the complete-suite verification. - Hosted Interrupt previously cleared the original queue and produced exactly one successor with neutral interruption feedback. It exposed the dispatch authorization defect. Retry on that earlier build was rejected for missing `agents:create` before creating another run. - Deployed the final application build (`38257f391`) to the scoped hosted instance and verified readiness. The latest PR commit changes only the credential test fixture; application code matches that deployment. A live Retry by the same operator without `agents:create` created one successor attributed to that operator, passing the former dispatch permission gate. Startup then stopped at `configuration_incomplete` because that operator has not configured their required personal Claude Code OAuth secret; the post-deployment run page confirms the operator identity and no provider work started, and the My secrets UI still shows the token as not set. Provider execution remains unverified pending that credential. No permission grants or credentials were changed. ## Risks Queue admission and finalization can race. The task lock, durable queue receipt, current comment IDs, and successor guard prevent duplicate dispatch. Process and lease stop checks, task pauses, approvals, ownership, and budgets remain in force. The API change only allows null on legacy Interrupt; native steering still requires an active run. No schema migration is required. Active non-viewer board members can now invoke existing agents without agent-creation permission. Agent self-invocation rules, raw provider-trace admin access, task retry scope, external chat authorization, and action-specific user/agent permissions remain enforced. ## Model Used OpenAI GPT-6 through Codex. The session does not expose an exact backend model ID or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d6232e7b0 |
feat: reuse provider sign-in across AI connection workflows (#13248)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect provider accounts during onboarding and agent setup. > - They should reuse and manage those accounts through the existing Connectors interface. > - A second login wizard would diverge from the established provider workflows. > - This pull request composes the existing sign-in components into Connections and agent configuration. > - Users can select accounts without changing their agent's harness or model. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add compact AI-account management to the existing Connectors pages. - Reuse AgentProviderConnection, AdapterLoginPanel, AdapterLoginChrome, and authentication controllers. - Add the shared connection picker to agent setup/settings and task requests. - Preserve onboarding's sequence and reuse existing accounts. - Add local-login recovery, retry, cancellation, and React StrictMode handling. - Add interactive Storybook scenarios, design-guide examples, and app acceptance checks. This is part 2 of the AI Connections change. The runtime foundation in #13247 is merged. This PR now targets master. ## Verification - Updated against master `47ded8bf9`, including the landed runtime foundation and upstream task-search changes. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All CI test, browser, build, packaging, and runner jobs passed on final head `dd17d3211931dd70aaa6ea619d83a7f9966dd18e`. The fresh Greptile review is 5/5, the security scan passed, and there are no unresolved review threads. The final CI aggregate gates passed. - Local general-server coverage passed 11,804 tests; three port-collision failures passed in an isolated 25-test rerun. All 6,111 UI tests passed. CLI coverage passed 484 tests; its remaining doctor test requires port 3199, which is occupied by an unrelated report server on this Mac. The complete CLI suite passed in CI. - Live browser testing verified Codex API-key reconnect inside a task card on desktop and phone. Real provider runs resumed and completed with unchanged connection/grant identity and agent routing. - Tested opening, cancelling, reopening, and completing connection creation. A regression confirms Connect another account cannot submit the new-agent form or copy provider keys into agent settings. - Added shared inline repair, automatic local sign-in checks, and responsive connection dialogs. Standalone Daytona installation ignores workspace configuration and suppresses dependency scripts. Its standalone build also passed with CI's exact pnpm 9.15.4. - Destructive live tests are excluded by default. Explicit opt-in, local deployment checks, and matching disposable fixture identities are required before any mutation. ## Risks - Local Codex/Grok creation requires the connection-specific terminal login command. - Browser sign-in uses the existing supported-environment controllers. - This update verifies live local Claude detection and Codex API-key task repair. New subscription authorization/refresh and independent-human/native-runner isolation were not reverified in this update. - No agent automatically adopts managed Connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
47ded8bf97 |
feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need credentials for a specific provider and sign-in method. > - Connections already owns accounts, grants, and access permissions. > - AI authentication should use those same boundaries. > - This pull request adds the storage, API, adoption, and runtime foundation. > - Legacy agents keep their authentication until they explicitly adopt a managed connection. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add AI-purpose/runtime-auth contracts and an additive, idempotent migration. - Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and catalog entries. - Store credentials on grants. Resolve responsible-user defaults or explicit permitted grants. - Isolate managed credentials and provider sessions across accounts. Block missing credentials without ambient fallback. - Keep imported legacy secrets unchanged during reconnect. Use independent local Codex/Grok sign-in attempts for rotating credentials. - Add authorization, migration, concurrent refresh, retry, cancellation, and legacy-compatibility tests. This is part 1 of a two-PR stack. The app UI follows in #13248. Merge the foundation first. ## Verification - Updated against master `04e364236`, preserving upstream provider login and connector workflows. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All current-head CI checks passed on `2a996560a`, including all server/workspace tests, browser shards, runner verification, typecheck, build, and canary dry run. Greptile reviewed that commit at 5/5 with no unresolved threads. Earlier local full-suite attempts hit the Mac PostgreSQL shared-memory limit; the complete suites passed in CI. - Renumbered the additive AI migration to `0276` after upstream migrations and regenerated its snapshot. Existing legacy agents retain their configuration. - Added local login status checks, owner-scoped retry, managed OpenCode remote homes, credential-aware model discovery, and task connection-repair delivery. ## Risks - Managed credential failures intentionally block execution. They do not restore legacy fallback. - Preview-era copied Codex/Grok subscriptions require independent reconnect. - The integrated branch has live provider acceptance coverage. This update verifies local Claude detection and Codex API-key task repair; it does not add a new subscription authorization/refresh or Daytona stress pass. - Runtime-auth connections must stay excluded from tool and channel handling. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1652a5c7f2 |
fix: improve task search relevance with a PostgreSQL rubric (#13335)
Unify full and quick task search around PostgreSQL term coverage, explicit relevance bands, and conservative typo recovery. Preserve matching evidence and navigation, and add a judged corpus, regression tests, and documented performance measurements. Validation: local typecheck, build, focused PostgreSQL tests, and task-list tests pass. All final-head CI gates pass and Greptile is 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
4d317274ce |
feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Channels connect external conversations to company tasks and agent execution. > - Slack, Discord, and AgentMail already provide durable delivery and access controls. > - People also need to reach an agent from Apple Messages and send photos. > - Photon provides shared Pro DMs, dedicated numbers, and authenticated event recovery. > - This pull request connects Photon to the existing channel services. > - People can message an agent while Paperclip retains task ownership and approval authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: channel services, shared contracts, database constraints, Apps, and agent Channels UI. **Problem or motivation** Paperclip has no iMessage channel. A person cannot use Apple Messages to start a task, send a photo, or answer an agent's pending question. **Proposed solution** Add experimental **iMessage Photon** with Pro-compatible shared DMs or a dedicated Photon Cloud number per agent channel. Reuse channel admission, identity links, task generations, publication, and interaction continuation. Keep groups disabled for shared allocation. Dedicated lines support groups that an operator explicitly enables. Require a fresh linked message and a published agent response before setup completes. **Alternatives considered** Shared allocation has no owned phone number, so it reserves one project and allows DMs only. Dedicated allocation reserves one stable number. Local Mac access needs a separate deployment model. The upstream Photon Chat SDK adapter does not persist the poll mappings and send receipts required here. This change uses the lower-level SDK without adding another agent runtime. **Roadmap alignment** This extends Connected Apps and agent communication through the existing channel subsystem. It does not add a parallel tool connection or agent loop. GitHub searches for Photon and iMessage found no matching provider implementation. **Additional context** This ships behind the existing experimental channel gate. Dedicated-line release qualification remains incomplete. Real Photon Pro DMs passed task/reply, native poll, text answers, confirmation rejection, media, restart, pause, reconnect, revocation, and removal tests. An operator-supplied iPhone camera HEIC also passed the full round trip. Dedicated groups remain unqualified. See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the implementation plan](doc/plans/2026-09-11-imessage-photon.md). ## What Changed - Add the provider catalog entry, shared setup contracts, and a forward migration. A global partial index reserves the dedicated number or shared project until its endpoint is archived. - Add Cloud project inspection, vaulted project credentials, selected-line token renewal, and a leased receiver. Persist checkpoint updates under the receiver lease. Shared project replay accepts sparse increasing sequences only after a complete recovery barrier. - Connect DMs and enabled groups to existing task generations, sender authorization, ordered delivery, and publication services. Keep each iMessage conversation on its task after completion; only explicit `/new` or `/close` releases the binding. Publish committed inbound comments live and label their human bubbles “Sent from iMessage” in both task-chat renderers. - Persist immutable text/file send identities, upload receipts, poll IDs, option IDs, per-person drafts, and canonical interaction continuation proofs. - Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG previews, and related Live Photo companion video retention. - Add the three-step setup flow and channel management surfaces with official branding. Preserve the experimental gate and existing pause/disconnect behavior. - Add interactive production-component Storybooks for setup, access, recovery, and ongoing conversations. Add provider, integration, catalog, and browser regression coverage. Document setup, recovery, supported boundaries, and qualification gaps. ## Verification - Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and receive native Codex replies in Apple Messages. Unlinked senders cannot start work. - Three real follow-ups each reopened the same completed task. Incoming bubbles appeared on its open page without reload and showed “Sent from iMessage.” The third follow-up ran after restarting the server on `4d7222110`; the agent correctly repeated its previous reply from before the restart. - Native polls after restart, sequential text drafts, required-field correction, explicit submission, approval rejection with a required reason, and native continuation passed against Photon. - PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC passed in both directions. The camera photo produced a 3024×4032 JPEG preview. The native agent described it and returned the received HEIC byte-for-byte. - Pause/resume, reconnect, identity revocation, removal, `/status`, `/new`, `/close`, and stale answers after close passed live. Messages suppressed by pause did not become work on resume. Removal stopped intake and removed credential bindings. - All 304 focused tests passed on `4d7222110`. These cover Photon unit/integration behavior, both task-chat renderers, live comment hydration, completed-task continuity after restart, enabled groups, duplicate delivery, and explicit reset/close. The selected Teams completion-boundary regression also passed. Full workspace typecheck/build and token gates passed for the conversation fix; the final UI changes passed their affected typecheck/build and tests. - All 26 new Photon Storybook Playwright cases passed in light and dark themes, including the complete shared-DM setup journey and 390px mobile follow-ups. UI typecheck and the Storybook build passed. These stories use simulated Photon responses and do not replace the live evidence above. - The full chat-adapters browser suite previously passed all 39 cases. Migration checks passed, and migration 0275 applied to the isolated live instance with the earlier Photon migration already applied. - The local full Vitest run was previously interrupted by the host's embedded-Postgres shared-memory limit; it is not a full-suite pass. All 30 applicable CI checks passed on preceding head `7a5419cac`, with two skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit required-story discovery guard to the 26 passing Storybook cases. Greptile rates this final head 5/5 with no unresolved review threads. All 30 applicable CI checks passed, with two optional checks skipped. - A repeated live send key suppressed the duplicate but returned gRPC 6 / SDK `internalError` without an original receipt. Paperclip keeps unknown delivery unresolved. This provider behavior is covered by a regression test. - See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package versions, redacted live evidence, deterministic coverage, and remaining qualification gaps. ## Risks - Dedicated group qualification remains unrun; groups are disabled for the approved Pro scope. Real iPhone camera HEIC passed transport, preview generation, agent inspection, and return. Keep the channel experimental; the dedicated-line release matrix remains incomplete. - Shared recovery and attachment aliases were verified against the live gateway. Duplicate writes currently return an error without the original receipt; unresolved sends require operator resolution. The implementation fails visibly on invalid replay ordering, a reset cursor, or changed identity. - The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF binaries have not been executed in this work. Linux musl has no packaged converter. Unsupported conversion retains the original and reports the missing preview. - The migration adds a global reservation across companies for Photon numbers and shared projects. Paused and revoked endpoints keep that reservation until removal. - Integration touches shared channel services. Existing provider browser coverage passes; broad repository verification is recorded above. - `pnpm-lock.yaml` is intentionally excluded under repository policy. The repository bot owns lockfile updates. The additional Superagent supply-chain scan is neutral/inconclusive because these new dependencies are not yet in the committed lockfile. Its security scan passed; all required CI checks pass. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code execution, browser testing, and tool use. The exact served model identifier and context-window size are not exposed in this session. No sub-agents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
44f6312cd8 |
fix(ci): reuse one available Cloud registry cache (#13334)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud needs a verified image for each merged source commit. > - Fresh builders restore compiled native dependencies from registry caches. > - The current workflow imports up to eleven historical cache manifests at once. > - Live builds missed native layers that a fresh builder reused from one manifest. > - This PR selects the nearest available cache and tests reuse across fresh builders. ## Linked Issues or Issue Description Refs #13329 and #13330. A search of open cache PRs found no duplicate of this change. **What existing behavior does this improve?** Remote Docker cache reuse on fresh Cloud image builders. **Current behavior** [Cloud run 34714483272](https://github.com/paperclipai/paperclip/actions/runs/34714483272/job/103609096836) imported the previous cache manifest successfully but rebuilt `cargo-chef` and Rust dependencies. The dependency compile took 3m43s. The preceding image build had already exported those layers. A controlled [fresh-builder diagnostic](https://github.com/paperclipai/paperclip/actions/runs/34715336530) used the same source and registry cache. The single-manifest job reused both layers immediately. The multiple-manifest job rebuilt them and failed the cache assertion. Both jobs used GitHub-hosted runners with read-only access. **Proposed behavior** Inspect cache manifests in first-parent order and import only the nearest available one. Keep full-SHA cache exports, the ten-commit search bound, and the legacy fallback. If caches cannot be read, permit a cold build. **Reason and benefit** Avoid the observed cache misses without changing image contents or builder sizes. Expected savings include about four minutes of native tool/dependency compilation when those inputs are unchanged. The final merge-to-deployable gain still needs a post-merge measurement. **Breaking changes** No image, artifact, deployment, or runner-routing contract changes. ## What Changed - Select one available ancestor cache after Docker login and Buildx setup. - Preserve separate writable cache tags for each full source SHA. - Test cache ordering, missing caches, registry errors, and workflow integration. - Add the selector tests to the existing release-registry suite. - Export a local test cache, remove the first builder, and verify a source rebuild on a fresh builder. - Document cache selection and the stronger Docker check. ## Verification - Passed 456 focused workflow, routing, readiness, preview-artifact, and cache-selector tests. - Passed shell syntax, ShellCheck for the changed probe, actionlint workflow validation, and `git diff --check`. actionlint's shell checks were disabled for the workflow validation because unchanged migration-label commands trigger existing SC2012 notes. - The fresh-builder registry diagnostic proves the single-cache behavior. The [permanent two-builder probe passed](https://github.com/paperclipai/paperclip/actions/runs/34715771048/job/103612624090), including a changed real binary and dependency-declaration invalidation. - Passed all 35 latest-head checks (green or intentionally skipped), including full typecheck, test, build, and browser suites in [PR CI run 34715771217](https://github.com/paperclipai/paperclip/actions/runs/34715771217). - The real selector CLI inspected registry metadata and chose the nearest available ancestor cache. - Fresh Greptile review is 5/5 with no open findings. The PR title was corrected to meet the source-change naming rule; the review check passed after that correction. - Local full-suite runs and Docker builds are unavailable because the local Docker daemon is unresponsive after disk exhaustion. CI provides the Linux verification. ## Risks - Missing or unreadable caches cause a slower cold build. The selector logs that condition and preserves image publication. - Inspecting several missing ancestors adds lookup time. Each lookup has a ten-second timeout and the search is bounded. - The Docker test now exports a local cache. It removes the first builder before starting the second to release disk space, then cleans up its builders and files. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0ce7df2648 |
ci: keep Cloud readiness markers out of the builder queue (#13330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud consumes versioned readiness markers for each merged source commit. > - Those markers can be published only after the required checks and artifacts pass. > - The marker jobs currently wait for the same AWS runner capacity as builds and tests. > - A busy builder pool can delay readiness after all required work has finished. > - This PR moves the small readiness jobs to GitHub-hosted runners while retaining every dependency gate. ## Linked Issues or Issue Description Refs #13326 and #13328. **What existing behavior does this improve?** Time from completed Cloud verification to a deployable marker, and AWS capacity occupied by artifact polling. **Current behavior** In [Cloud readiness run 34711557083](https://github.com/paperclipai/paperclip/actions/runs/34711557083), all builds and tests finished at 18:43:05 UTC. The source marker did not start until 18:44:21, and the deployable marker did not start until 18:44:37. Merge-to-deployable was 11m33s, although the prerequisite work finished in 9m44s. **Proposed behavior** Run the artifact wait and both versioned marker jobs on `ubuntu-latest`. Keep the compute jobs on the approved post-merge AWS fleet. **Reason and benefit** Avoid builder-pool queue delays after verification finishes. This also removes the long artifact-wait job from AWS capacity. Expected savings depend on queue depth: the observed run had over 90 seconds of avoidable marker waiting. The marker commands themselves take only seconds. **Breaking changes** Runner placement changes for three bookkeeping jobs. Marker names, exact-source artifact checks, required verification, and image verification stay the same. **Additional context** Searched the related runner and Cloud readiness work. This addresses queue time observed after the parallel verification change. ## What Changed - Place the artifact wait, source-verification marker, and deployable marker on GitHub-hosted runners. - Keep all existing job dependencies, source guards, permissions, and commands. - Extend routing regressions to enforce this placement and retain fail-closed readiness gates. - Document why readiness bookkeeping uses separate runner capacity. ## Verification - Passed 433 workflow, routing, source-verification, and Cloud readiness tests with `node --test .github/scripts/tests/*.test.mjs scripts/cloud-source-verification.test.mjs scripts/cloud-readiness.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Passed `actionlint` and `git diff --check`. - Passed all latest-head CI gates in [run 34712624340, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/34712624340), including typecheck, build, browser, Runner, and all general/serialized tests. - Attempt 1 had one localhost readiness timeout in an unchanged test. A single targeted retry passed all 166 files (3,225 tests passed, one existing skip), including all seven tests in that file. No timeout, assertion, or application source was changed; the retry is documented in the PR comment. - Local full-suite verification is limited by local disk exhaustion; the focused checks above pass. - Fresh Greptile review is 5/5 with no unresolved findings. ## Risks - GitHub-hosted capacity can also queue, but these jobs no longer compete with AWS build/test demand. The change does not reserve instances or change box sizes. - Readiness must still fail if any prerequisite fails. The existing `needs` relationships and success-only execution are preserved and tested. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7435b2ee9c |
ci: cache compiled Docker Rust dependencies separately from source (#13329)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys images that contain the native Rust Runner. > - The image already builds that Runner before copying ordinary app source. > - A Rust source change still invalidates its entire compiled dependency layer. > - Compiled dependencies can survive source changes when their recipe is unchanged. > - This PR adds a separate locked dependency build before compiling the real workspace. ## Linked Issues or Issue Description Refs #13195. A search of related Docker and Cargo cache PRs found no duplicate dependency-recipe change. **What existing behavior does this improve?** Docker image build time after Rust source or embedded protocol changes. **Current behavior** The `runner-build` stage compiles dependencies and workspace code in one layer. In Cloud readiness run 34698143548, that stage took about 3m48s when its cache was unavailable. **Proposed behavior** Generate a recipe with pinned cargo-chef 0.1.73. Build locked release dependencies in `runner-deps`, then copy and compile real Rust source and embedded protocol inputs in `runner-build`. Source edits can reuse the dependency layer from the existing registry cache. **Reason and benefit** Reduce dependency recompilation during source changes and merge bursts. Expected savings are roughly 2–4 minutes when the old native layer would miss but dependency layers are available. Full cold builds also pay for the recipe tool installation. Ordinary app-only cache hits gain little from this change. **Breaking changes** None to the shipped application or image tags. The recipe tool and compiled dependencies remain in build stages. ## What Changed - Install a pinned recipe generator with its locked dependencies and the existing package-owned compiler. - Add recipe planning and compiled dependency stages. Use the same release profile, package, binary, and lockfile enforcement as the real native build. - Remove generated source stubs before copying actual source. Preserve protocol inputs, timestamp normalization, binary staging, and application checks. - Add Docker cache wiring regressions and update the Docker cache documentation. - Run a two-build probe in Docker Runner check. It requires dependency reuse, changed real binary metadata after a source edit, and a changed recipe after a dependency declaration edit. It uses a disposable tracked-source context and exports only small metadata files. ## Verification - Passed all five Docker build-stamp and dependency-cache tests with `pnpm exec vitest run server/src/__tests__/docker-build-stamp.test.ts`. - Passed the local ARM64 `docker buildx build --target runner-build --progress plain`. Local Docker then hit storage errors during a runtime probe; cache invalidation verification continues on GitHub-hosted Linux. - Passed `bash -n scripts/check-docker-runner-cache.sh`, `actionlint`, and `git diff --check`. - Passed a [Linux AMD64 cache probe](https://github.com/paperclipai/paperclip/actions/runs/34711042199) against the PR source: dependencies compiled in 3m49s for the baseline and were `CACHED` after a source edit; real source compilation took about 37 seconds. Binary metadata changed and dependency declaration changes altered the recipe. The permanent probe is also running in latest-head Docker Runner check. - Passed latest-head [Docker Runner check](https://github.com/paperclipai/paperclip/actions/runs/34711145160), including the permanent source/dependency invalidation probe. - Passed full [PR verification](https://github.com/paperclipai/paperclip/actions/runs/34711145352/attempts/2): typecheck, all grouped tests, native verification, build, release dry run, and browser checks. One unrelated signoff-policy browser test failed waiting for a heartbeat run on attempt 1; only that failed shard and dependent checks were retried, and passed. - Latest-head Greptile is 5/5 with no unresolved findings. Full local tests/build were limited by local disk exhaustion; Linux CI completed those checks. ## Risks - The two-build CI probe has a 20-minute job limit to cover the cold build and source rebuild. It adds no AWS routing. - A fully cold build must install cargo-chef and populate the dependency layer. Both become reusable registry layers; no Actions cache is added. - The recipe and final build must keep the same compiler, build profile, package, binary, and directory layout. A source-change rebuild probe checks real cache reuse and binary invalidation. - Dependency or compiler changes still require rebuilding dependencies. Existing image verification and full-SHA publication gates remain unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed50a39c3f |
fix: preserve NUL characters in run-event payloads (#13325)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d2e940f4c1 |
ci: run release Runner protocol and Rust checks in parallel (#13326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys verified images from merged source commits. > - Cloud readiness waits for every release verification check. > - Runner verification currently runs long TypeScript tests before Rust checks. > - These checks can run on independent runners with their own build directories. > - This PR runs them in parallel while preserving all checks and the shared dependency cache. ## Linked Issues or Issue Description Refs #13194. Related prior work: #13142 and #13259. A search found no duplicate parallel release-check change. **What existing behavior does this improve?** Time from merge to Cloud source verification and deployment readiness. **Current behavior** Recent successful runs take roughly 13 minutes from merge to deployable. In run 34705914878, Runner verification took 11m23s. Protocol tests finished before Rust tests and API authority checks started. **Proposed behavior** Run protocol and Rust verification in two matrix jobs. Cloud readiness still requires both jobs to pass. **Reason and benefit** Remove the serial dependency between independent checks. Expected improvement is about 2–3 minutes on a typical cached run, until the image build or server tests become the longest job. This is an estimate; post-merge timing will confirm it. **Breaking changes** Individual release Runner job names gain a lane suffix. Cloud source and readiness marker names stay the same. PR runner routing is unchanged. ## What Changed - Split release Runner checks into protocol and Rust lanes. Keep every constituent of `check:all` exactly once. - Restore the existing Rust dependency cache in both lanes. Allow only the Rust lane to save it after warming both build profiles. - Add coverage and cache authorization regressions. Document the parallel verification and single cache writer. ## Verification - Passed 477 workflow and source-verification tests with `node --test .github/scripts/tests/*.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Passed `actionlint`, `git diff --check`, and the private AWS routing regression suite. - Passed local `pnpm -r typecheck` and the standalone `check:runner && check:api-authority` lane, including all 1,671 API tests before the protocol lane had built TypeScript output. - The broad local protocol run under Node 25 had four failures. The two affected files passed under CI's Node 24.19.0: 67 passed, 6 platform skips. - Local `pnpm test:run` aborted when disk space ran out; local `pnpm build` could not run afterward. These are local verification limits. [Linux CI run 34710421424](https://github.com/paperclipai/paperclip/actions/runs/34710421424) passed full typecheck, all grouped tests, native verification, build, release dry run, and browser checks. Native protocol CI passed 1,986 tests, plus 1,671 API tests and the Rust suites. - Latest-head Greptile is 5/5 with no open findings. All 33 current-head checks are successful or intentionally skipped. ## Risks - Uses one additional short-lived verification runner per release verification. The existing AWS exact-master restriction remains in place. - The Rust lane warms debug dependencies so its cache save also serves protocol tests. Both lanes always rebuild workspace code. - A workflow regression could omit a check. The new coverage test compares the matrix checks directly with `check:all`; Cloud readiness depends on the complete reusable workflow. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c9021c6721 |
fix: require explicit native completion reviews (#13314)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs report their outcome through paperclip_finish. > - The server previously turned incomplete reports into human approval requests. > - Those requests could block a later successful run, even when no person had requested review. > - This pull request creates review cards only for explicit attention requests and withdraws proven old fallback cards. > - Agents receive useful completion feedback, while explicit approval gates and task state protections remain in force. ## Linked Issues or Issue Description Related: #13266 removed reviews caused by policy upgrades. This change removes the separate completion fallback. **What happened?** An agent reported needs_review while waiting for checks without requesting a human decision. Paperclip created a generic Native completion review. A later successful report could not complete the task because that old card remained pending. **Expected behavior** Ordinary low-risk work completes after a valid done report, a successful run, and workspace finalization. Incomplete work stays with the agent. Explicit approval requests remain visible and must be resolved. **Steps to reproduce** 1. Complete a native run with needs_review and no attention requests. 2. Continue the task and submit a successful done report. 3. Observe that the old implementation leaves the task in review behind a generic confirmation card. ## What Changed - Require explicit attention requests to create native review cards. Route each request independently and preserve pending or declined decisions. - Withdraw only pending system cards with matching old decision, assessment, effect, contract, and prompt provenance. Preserve history and explicit or answered requests. - Reassess an affected current result without overwriting later task edits, runs, contracts, or workspace failures. - Return pending approval links and required actions through the completion tool. Reject contradictory done reports and empty review requests before accepting a result. - Allow one corrective continuation for incomplete results, then expose a recovery action. - Update status fixtures, database regressions, runner tests, and the completion contract documentation. ## Verification - `pnpm -r typecheck` passed after merging current master. - `pnpm build` passed after merging current master. - The combined branch passed 89 completion and Agent Chat tests. Other targeted tests passed: 170 external-chat and reconciliation tests; 50 runner-resume and control-plane tests; 13 arbiter tests; 7 chat delivery tests; 21 runner completion and runtime-context tests. - The full local test attempt exposed old review fixtures and a missing fake-provider binary. The fixtures are fixed and the helper is built. All affected suites pass in fresh reruns. The timing-sensitive Discord test also passed on rerun. - All latest-head CI checks passed, including build, typecheck, general and serialized tests, runner verification, browser tests, and canary dry run. Greptile is 5/5 with zero unresolved comments. ## Risks - Cleanup changes existing pending cards. It requires exact system provenance and only applies to low-risk agent-claim contracts. It does not delete history or dismiss explicit requests. - Status still commits after the turn and workspace finalization. Completion feedback reports current constraints and does not claim an early status commit. - Incomplete reports now request a bounded corrective run instead of an automatic approval. Repeated failures expose recovery. ## Model Used OpenAI Codex, GPT-6, with repository inspection, code execution, and test tools. The exact deployment identifier and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e14c61da7 |
fix: fence native startup against cancellation (#13316)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users can pause a task while its runner prepares to start. > - Cancellation must prevent preparation from creating new execution authority. > - Native runtime selection could run after cancellation and leave an unclaimed recovery coordinator. > - Saved user messages then waited for recovery that had no eligible worker. > - This pull request fences startup and lets explicit user continuation settle verified, unclaimed startup state. > - Tasks can continue after cleanup while keeping provider ownership and execution safeguards. ## Linked Issues or Issue Description Refs #13285 for startup controller leases and #13270 for explicit continuation and saved-message recovery. Related #13293 covers retained processes that actually started; this change covers cancellation before the native provider claim. Related #13315 covers legacy queued-message delivery. **What happened?** Pausing a task during startup could cancel its heartbeat before native runtime selection. Stale preparation then created an observed native coordinator on the cancelled run. The coordinator had no provider result or eligible recovery worker. A later Continue message stayed queued indefinitely. **Expected behavior** Cancellation fences native startup. After verified cleanup, a newer user message starts one fresh conversation turn. An unverified execution keeps its hold and a clear explanation. **Steps to reproduce** 1. Start a task with the native runner and delay startup preparation before runtime selection. 2. Pause the task, then release preparation. 3. Resume the task and send Continue. 4. Before this fix, native selection can persist after cancellation and block the saved message. 5. Repeat from persisted cancelled startup state after a server restart. **Paperclip version or commit** Reproduced against master at `586b5ec82` with isolated PostgreSQL regression fixtures. **Deployment mode** Self-hosted server built from source, with Paperclip Runner. ## What Changed - Serialize the cancellation fence and native runtime selection on the run row. Revalidate the startup controller lease. - Refresh the runtime before dispatching cancellation, and reject terminal or cancelled runs at the native provider claim. - Recognize never-claimed coordinators only after startup and environment cleanup are verified. Reject process, provider, owner, and conflicting launch evidence. - Settle that coordinator atomically with a new authenticated user turn. Preserve history, unknown outcomes, and attempt counts. - Reuse the saved-message worker after restart and retain pause, budget, approval, and ownership gates. - Add startup, restart, duplicate-admission, and negative-proof regressions. Document the rule. ## Verification - All 687 tests pass across the complete heartbeat recovery, explicit continuation, and native session executor suites on the rebased branch. - Six focused race regressions also pass: cancellation before and after native selection, Stop racing adapter registration, process termination during a database failure, and a run finishing during cancellation. - `pnpm -r typecheck` and `pnpm build` pass after rebasing on master at `ab15aff39`. - Greptile is 5/5 on `f3dea2ab27facdf0360b56172ebd3e5219e25538`, with no open review threads. Policy and security checks pass. - All 32 CI checks pass on the final head, including the full server/workspace test matrix, all three browser shards, runner verification, typechecks, build, and canary dry run. The two optional Storybook jobs are skipped. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34697945514). - Local full-suite limitation: the broad `pnpm test:run` attempt reported two failures outside the changed area after unusually long test durations (about 65 seconds for supporting skill-file saves and 933 seconds for setup-token login). Both cases passed isolated reruns, with no code changes. The broad local run was stopped after CI completed successfully; no clean full local-suite pass is claimed. ## Risks Cancellation and startup overlap. The run and coordinator locks provide the authority fence; cleanup and process evidence provide the containment proof. Historical runs without sufficient evidence remain blocked. A saved user message authorizes a fresh turn, not automatic replay. No schema migration or dependency change. ## Model Used OpenAI GPT-6 through Codex, using reasoning, repository inspection, code execution, and tests. This session does not expose the exact backend revision or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |