mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
92d4868e79b0451aeb0746681a7d93d1cc199dda
4535
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
92d4868e79 |
fix(server): isolate run errors and redact runtime capability headers (#13826)
## Thinking Path
> - Paperclip manages AI agents and their work.
> - Operators use Sentry to investigate failed runs and server errors.
> - Run reports attach a task ID, run ID, error code, and adapter
fingerprint.
> - The server skips Sentry's OpenTelemetry setup to preserve its
separate tracing and privacy settings.
> - Without an async context manager, a scope mutation can attach old
run data to later errors.
> - The HTTP logger also retains a runtime credential capability header.
> - This change isolates run metadata and redacts that header so
diagnostics identify failures without leaking credentials.
## Linked Issues or Issue Description
Refs #13446 and #13719.
**What happened?**
After a terminal run failure, an unrelated server exception can inherit
that run's tags, context, and fingerprint. Sentry then groups a database
error with an earlier adapter failure. The real SDK reproduces this with
the application's `skipOpenTelemetrySetup: true` setting. HTTP request
logs also retain the `x-paperclip-github-capability` header, which must
be treated as a credential.
**Expected behavior**
Run metadata belongs to the terminal run event. Later exceptions must
not inherit it. Every genuine error must still be captured. Runtime
capability headers must be redacted on success and failure logs.
**Steps to reproduce**
1. Initialize the optional Sentry SDK with the application's options and
an in-memory transport.
2. Capture a terminal run failure.
3. Capture an unrelated exception.
4. Inspect the second event. Before this fix, it contains the first
run's identity and fingerprint.
5. Send a request with a fixture runtime GitHub capability header.
Before this fix, HTTP logs retain the fixture value.
## What Changed
- Pass tags, context, and fingerprint directly to `captureException`
instead of mutating the ambient scope.
- Preserve the existing run fields, grouping keys, ordinary exception
capture, and privacy settings.
- Test two run identities interleaved with unrelated exceptions against
the real optional SDK.
- Update the capture contract tests and document event-local run
metadata.
- Redact the runtime GitHub capability header through the existing HTTP
logger policy. Test successful, denied, and failed requests.
- Add a dedicated GitHub-hosted CI check that installs the exact
optional SDK version declared in `server/package.json`. It fails if the
real-SDK regression would be skipped. The SDK stays outside the
workspace and production dependency graph.
## Verification
- The real-SDK regression failed before the fix because the unrelated
event contained `contexts.run_failure`.
- Five focused suites passed: 123 tests, including all optional SDK
tests. Suites: `run-failure-sentry-real-sdk.test.ts`,
`run-failure-sentry.test.ts`, `sentry.test.ts`,
`run-failure-report.test.ts`, and `http-log-redaction.test.ts`. A custom
in-memory transport prevented outbound Sentry delivery.
- All three new header-redaction cases failed before the policy fix and
passed afterward.
- The dedicated CI command passed locally with
`PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1` and the audited SDK available
through `NODE_PATH`.
- Server TypeScript check passed with a scratch configuration that
resolves this checkout's workspace packages. The existing dependency
links point to another checkout.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- Gitleaks and a separate private-data scan passed before push.
- Full local workspace typecheck, test, and build were not run. The
machine has less than 2 GiB free and those commands include Rust builds.
Full PR CI must pass before merge.
- The dedicated real-SDK GitHub check passed with 1 test executed and no
skips: https://github.com/paperclipai/paperclip/actions/runs/35774449002
- Greptile reviewed
canary/v2026.922.0-canary.11
|
||
|
|
a68f3d8e35 |
fix(runner-e2e): align Daytona image and provider pack provenance (#13814)
## Thinking Path > - Paperclip is an open source app people use to manage AI agents for work. > - The Daytona runner image provides the native runner and its provider package. > - The controller also sends a provider package to native Daytona cells when the image package does not match. > - A stale image and a package from another source revision caused a 1.8 GB upload before Claude could run. > - This pull request refreshes the reviewed lock checksum and documents how to reuse the exact package from an immutable image. > - The benefit is a reproducible setup path and clear evidence when image and package provenance do not match. ## Linked Issues or Issue Description **What happened?** A Daytona native Claude run used image source revision `45c99a0d06cbd5b04982b06b79de321149930ac5` with a controller provider package from revision `294853dc...`. Runtime verification rejected the image package and staged a large package upload before model execution. A hosted campaign also failed during image setup because the Dockerfile expected lock checksum `d7d96cf0...` while the resolved lockfile checksum was `4b796c312833ebf2be4c38228babc0292c76774fb40d43bd18b54bd6b205753d`. Related public work reviewed: [#12795](https://github.com/paperclipai/paperclip/pull/12795), [#12862](https://github.com/paperclipai/paperclip/pull/12862), and [#12887](https://github.com/paperclipai/paperclip/pull/12887). **Expected behavior** The Daytona image and controller provider package must come from the same verified build. A package extracted from the immutable image must pass the existing manifest, source revision, lockfile, binary, bridge, and artifact checks before a native Claude run starts. **Steps to reproduce** 1. Set `PAPERCLIP_E2E_DAYTONA_IMAGE` to the old immutable image digest. 2. Set `PAPERCLIP_RUNNER_REMOTE_PROVIDER_PACK_PATH` to a package built from a different source revision. 3. Run a native Claude Daytona cell. 4. Observe provider package verification failure followed by the large staging upload. 5. Build the image with the stale Dockerfile lock checksum and observe the checksum failure. **Paperclip version or commit** `3b8df3dcd6f99e99277faa45f3351989edb5c239`. **Deployment mode** Daytona native runner E2E. **Install method** Built from source. **Agent adapter(s) involved** Claude Code through the native ACPX runner. **Database mode** Not database-related. **Access context** Not applicable to the setup failure. ## What Changed - Refreshed `PAPERCLIP_RUNNER_LOCK_SHA256` in `docker/daytona-runner/Dockerfile` to the resolved lockfile checksum. - Added a local guide for extracting the provider package from an immutable verified image with Docker. - Documented the required provenance checks and the expected manifest-matched runtime log. - Documented that cold package upload coverage must remain separate from recovery coverage. - Pinned the preview-service test guest to the test runner’s Node executable and logged guest startup and bound ports for readiness diagnostics. ## Verification - 442 focused E2E tests passed. - Typecheck passed. - `pnpm build` passed. - Five provider-pack reuse tests passed. - Six Daytona image contract tests passed. - Fixed Daytona campaign [35740613581](https://github.com/paperclipai/paperclip/actions/runs/35740613581) passed. - The overall hiring campaign [35739993219](https://github.com/paperclipai/paperclip/actions/runs/35739993219) failed because of an unrelated Mini metadata failure; its Claude cell passed. - Full local `pnpm test:run` was attempted but did not complete. The isolated Postgres install was repaired and its 15-test probe passed. - After the fixture change, all seven preview reservation tests passed locally and CI server shard 6/12 passed on `20234f75f`. This removes login-shell Node resolution variance; the exact cause of the earlier CI-only timeout is not established. - CI run [35766034635](https://github.com/paperclipai/paperclip/actions/runs/35766034635) passed on `20234f75f`. All server, runner, browser, build, and typecheck gates passed. - The signoff browser case initially failed waiting for an approver run. All five signoff tests passed locally without changes; the one allowed CI retry passed all 19 shard tests. This is recorded as an intermittent failure, not a demonstrated product fix. - Greptile reviewed `20234f75f`: 5/5, no actionable findings. - [Follow-up report](https://pages.paperclip.ing/runner-daytona-hiring-20260922/) includes timings, evidence links, and the remaining hiring configuration failure. - Review the immutable image source revision and extracted `provider-pack.json` before another paid recovery run. ## Risks - The Dockerfile checksum gate intentionally fails when the resolved lockfile changes. A future dependency change must refresh the reviewed checksum with the image change. - The local extraction guide requires Docker and a pullable immutable image. - An image built from an older source revision can still fail runtime manifest verification. The guide does not bypass that check. - The change does not alter runner prompts, approval policy, or recovery behavior. ## Model Used OpenAI Codex using the primary GPT-6 backend; the exact backend deployment ID is not exposed. Repository analysis and code execution used tool access. Assistance also came from OpenAI gpt-5.6-luna. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.10 |
||
|
|
2788f20fc0 |
fix(ui): show ancestors in the task detail Tasks panel (#13823)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks form a hierarchy that explains why each piece of work exists. > - The streamlined Tasks panel shows child tasks and work created by the current task. > - It does not show ancestors, although the task response already includes them. > - This pull request adds linked ancestors above Subtasks in root-to-parent order. > - Users can now move up the task hierarchy from the same panel. ## Linked Issues or Issue Description Related implementation: #13241. A search found no duplicate ancestor-panel PR. **What happened?** The Tasks panel omitted the current task's ancestors. The streamlined header also hides hierarchy breadcrumbs. **Expected behavior** The Tasks panel should show the ancestor chain and let users open each ancestor. **Steps to reproduce** 1. Open a task with a parent and grandparent in the streamlined UI. 2. Open the Tasks tab in the side panel. 3. Observe that it shows child and created tasks but no ancestors. **Paperclip version or commit** Reproduced on master atcanary/v2026.922.0-canary.9 |
||
|
|
5f1100e3b3 |
refactor(server): remove retired operator UI snippet injection (#13789)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators can extend its interface through trusted plugin UI contributions. > - The server also accepts executable HTML through two legacy environment settings. > - This older path bypasses the plugin installation and lifecycle model. > - This pull request removes snippet injection from static and development pages. > - Operators must migrate existing integrations before upgrading. ## Linked Issues or Issue Description **What existing behavior does this improve?** Retire the operator HTML injection path from the server. Related public changes: #13168, #13245 and #13496 introduced the legacy settings; #13646 supplies the generic plugin host contract. **Current behavior** Managed instances append operator-supplied HTML or a decoded script body to every page. The same integration can use supported trusted plugin UI slots. **Proposed behavior** Ignore both retired settings. Serve normal branded static and development HTML, and keep plugin contributions unchanged. **Breaking changes** Installations that rely on `PAPERCLIP_CLOUD_UI_SNIPPET` or `PAPERCLIP_CLOUD_UI_SNIPPET_B64` must migrate before upgrading. The maintainer-owned staging and production deployments have completed the migration prerequisite. Other operators must migrate their integrations before adopting this change. ## What Changed - Remove the snippet injector and its static/dev rendering integration. - Remove injector-specific tests and retain a regression that old settings no longer change served HTML. - Replace setup instructions with a retirement and plugin-migration note. ## Verification - Rebased onto current master; `pnpm exec vitest run server/src/__tests__/static-index-html.test.ts server/src/__tests__/vite-html-renderer.test.ts`: 5 tests passed. - With repository-pinned Rust/Cargo installed, full `pnpm -r typecheck` and `pnpm build` pass after rebase. - Previous full local `pnpm test:run` encountered unrelated macOS runtime-cache rename `EACCES` errors and a missing AgentMail skill path; a focused reproduction confirmed 5 failures / 95 passes. That is not a passing full-suite result. The unchanged focused suites, build and typecheck were repeated after rebase; the full local suite was not repeated. All 54 refreshed GitHub checks/contexts passed on `bd62bff63f9d7980bfd10e54cb0102693d27ecfb`; fresh Greptile is 5/5 with no unresolved threads. - No UI component styling, database or API contract changed. - Maintainer approved the remaining rollout and cleanup. Staging snippet retirement and sleep/wake verification are complete. Production migration and removal of the legacy settings are complete for serving tenant instances. Remaining old warm inventory is excluded from new signups until configuration reconciliation completes. The maintainer authorized upgrading the remaining old deployments and clearing their pins. ## Risks - Removing the settings disables integrations that still depend on them; operators outside the completed maintainer rollout must migrate before upgrading. This is an intentional behavior change, documented at the existing setup-doc path. - Keep a previous image and its configuration for rollback. Existing browser tabs need a refresh to unload already-injected code. - Plugin UI remains trusted same-origin code. This does not add a security sandbox or change ordinary branding. ## Model Used - OpenAI GPT-6 (Codex; exact deployment variant and context-window size are not exposed in this session). Reasoning, repository inspection, local code execution and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused rendering tests; unrelated local full-suite failures are disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All refreshed checks pass on `bd62bff63f9d7980bfd10e54cb0102693d27ecfb` - [x] Fresh Greptile is 5/5 on the current head, with no unresolved findings - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3d78e3a4ec |
fix(runner): keep warm sessions alive with managed GitHub access (#13815)
## Thinking Path > - Paperclip manages AI agents and their work. > - The native Runner keeps a live provider process between task turns. > - Managed GitHub access used a token tied to one run. > - A new run forced Paperclip to replace that process to replace its token. > - This PR gives the session a stable credential transport and binds each operation to the active run. > - The agent can keep its process while Paperclip checks current identity and grants. ## Linked Issues or Issue Description Follow-up to #13738. Related credential-rotation work: #11770 and #8208 use process replacement for other adapter credentials; this change applies to managed GitHub access in the native Runner. **What happened?** A configured GitHub connection forced a warm native provider process to close at each new run. The saved conversation survived, but the live process did not. **Expected behavior** Keep the warm provider process. Resolve GitHub access for the current run when each command starts. Deny access while idle or after the run ends. **Steps to reproduce** 1. Configure managed GitHub access for a native Runner agent with a warm session. 2. Complete a turn, then send another message to the same task. 3. Observe the provider process close with the reason `warm native session configuration changed`. **Paperclip version or commit** Reproduced on master `8326e33ad`. Rebased onto `e3d8fb087` before submission. **Deployment mode** Local and remote native execution, including the sandbox callback bridge. ## What Changed - Move configured native GitHub transport and launcher ownership from the run to the provider session. - Bind the broker only after the executor acquires session ownership. Clear that binding when the run exits. - Keep the shared live-run, identity, grant, and trust-policy checks for each credential request. - Reject wrong scopes, idle requests, and credential responses that arrive after their run binding changes. - Retire transport and launcher files with the provider session. Keep anonymous commands available if bridge startup fails. - Add red/green executor tests, real subprocess and callback-bridge tests, and database checks. Update the runtime documentation. ## Verification - Before the fix, both new local and remote warm-session reuse tests failed. - After the fix, 435 targeted tests passed across the executor, broker, launcher, token, and database suites. - A real long-lived test process kept the same PID and original environment across two runs, including through the production callback bridge on local test processes. - Server typecheck and TypeScript compilation passed. - Full workspace typecheck and build passed. Server typecheck passed again after the review fix. - The fallback-logging regression failed before the fix; all 9 broker tests pass afterward. - The exact chat sidebar browser scenario passed locally. The initial CI timeout showed failed Vite module downloads; all eight browser shards pass on the latest commit. - All 53 latest-head checks passed, including the full CI test matrix and security checks (two unrelated conditional checks skipped). - The duplicate full local test run was stopped after CI passed; it is not claimed as a completed local pass. Targeted local tests, workspace typecheck/build, and the browser scenario passed. - Greptile reviewed the latest commit at 5/5 with no unresolved findings. - No fresh paid provider or Daytona campaign has run for this change. ## Risks - The broker now lives as long as the provider session. Tests cover idle denial, late cleanup, late responses, shutdown, and failed startup. - Its in-memory authority does not survive a controller restart. Existing checkpoint and process-recovery rules still apply. - Raw GitHub credentials remain confined to individual command processes. The session transport token cannot select a different task, agent, company, or run. - No database migration or public API change. ## Model Used OpenAI Codex, GPT-6, with reasoning, terminal tools, and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
74a9730acb |
fix: continue native agent chats after worker loss (#13813)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat uses native workers to run Claude and Codex conversations. > - A worker crash leaves a cleanup hold because its provider did not acknowledge suspension. > - A new user message must not reuse that unverified session or repeat old tool calls. > - The existing continuation path can preserve history and start a fresh session, but local cleanup ownership remained held. > - This pull request verifies the stopped local owners and releases only their cleanup hold for a new user turn. ## Linked Issues or Issue Description Refs #13775. **What happened?** After a native worker crashed, both providers retained cleanup quarantine. A saved plan survived, but the conversation could not produce another answer. **Expected behavior** Once the old worker and provider process groups have stopped, a new user message can continue in a fresh session with the saved work and prior action history. **Steps to reproduce** Run the opt-in `agent-chat-qualification` suite with case `worker-crash-retry` on native Codex and native Claude. The fixture saves a plan, kills the exact worker through a Linux pidfd, releases a local read-only brief, and sends a new message. ## What Changed - Verify the exact local worker stop receipt, provider identity receipts, released leases, and retained state before retiring a native cleanup hold. - Recheck process liveness and state before admission. Keep the old run and durable session files intact. - Use the existing explicit conversation continuation path. Generic Retry remains blocked for cleanup quarantine, including on the old failed-run marker after a successful continuation. - Extend the live oracle to require a successful fresh session, correct predecessor context, unchanged plan, one original message, and one answer containing a reference introduced after the crash. - Add physical-proof and database-backed admission tests. Document the precise qualification scope. ## Verification - Live Product E2E: **2/2 passed**, **2/2 cleanup passed**, with real native `gpt-5.6-sol` and `claude-sonnet-5`, Chromium, server, database, and public APIs. - Core recovery proof source: `3592b04c2bc76e23795fcdf964720e38a409dc4d`. Suite definition version 8: `9867367994d81a0c726956d91f2c7fddab6417a12f41b5cef3f7e62f3be417da`. - [Core recovery campaign](https://github.com/paperclipai/paperclip/actions/runs/35741746990) · [Public evidence report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35741746990-1/). - Both cells verify the real crash boundary, blocked generic Retry, unchanged saved plan, one original prompt, one fresh successor with predecessor context, and one run-attributed answer containing the post-crash reference. Billing coverage is partial because the crashed runs did not report complete usage; missing cost is not zero cost. - [Master baseline](https://github.com/paperclipai/paperclip/actions/runs/35737364443): both providers stopped at quarantine, with cleanup passing. - [Follow-up baseline without the fix](https://github.com/paperclipai/paperclip/actions/runs/35738638866): Codex produced a complete red result. Claude reached the same error, but its artifact upload was canceled. - [Complete Claude baseline with the same version-8 definition](https://github.com/paperclipai/paperclip/actions/runs/35740555177): red at cleanup quarantine, cleanup passed, source `6479a90c5754044356b39a9278b9a3e92ce8e55e`. - Earlier candidate attempts remain retained: [first](https://github.com/paperclipai/paperclip/actions/runs/35738449214) passed Codex and found a Claude fixture wait race; [second](https://github.com/paperclipai/paperclip/actions/runs/35740408065) exposed the normalized session-open receipt mismatch. Both corrections are in the final source. - Targeted server suites: 570 passed before the final two additional receipt regression cases. The physical-proof suite, including those cases, passed 46/46. Eval oracle and catalog: 40 passed. Server and Product E2E typechecks passed. - **All 54 PR checks passed on final head `db6f775df`**, including repository typecheck, build, tests, browser shards, and canary dry run. The canary job required one retry after its runner received a shutdown signal. On the earlier core proof head, two timing-sensitive tests passed in isolation and on a single CI retry. - Final UI regression checks: 5 passed; UI typecheck and token gates passed. Updated eval oracle/catalog: 40 passed; eval typecheck passed. - Final version-9 two-provider campaign: **2/2 passed, 2/2 cleanup passed**, including the browser assertion that the quarantined historical run never regains Try again. [Final campaign](https://github.com/paperclipai/paperclip/actions/runs/35747416013), source `db6f775dfff405e1514ec02fedb0450d42c7dad2`, definition hash `bf5abf1cc45cb6dad4e082fbf818b8fa0f4d8c282776a7e98762b22919eadab7`. [Final public evidence report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35747416013-1/). Final screenshots and retained state inspected for both providers; cost coverage remains partial. ## Risks - Missing or conflicting stop evidence keeps the conversation blocked. This change does not kill an unverified process. - This qualifies new user input after local worker loss. It does not enable automatic replay, exact-session recovery, remote crash recovery, or native onboarding defaults. - Old action outcomes remain part of the continuation. A process exit is not proof that an action did not happen. - No schema migration or production prompt change. ## Model Used OpenAI Codex, GPT-6. The runtime does not expose a more specific model identifier or context-window size. Used reasoning, repository inspection, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
110d176fc9 |
fix: preserve chat message bindings in review recovery (#13818)
## Thinking Path > - Paperclip manages AI agents and their work. > - External chat messages can start work on tasks that remain in review. > - Paperclip queues one recovery run when a task loses its review path. > - That recovery keeps the chat source but loses the admitted message IDs. > - The authorization check then rejects the recovery before execution starts. > - This change retains the message IDs so the existing check can verify current access. ## Linked Issues or Issue Description Refs #13809. **What happened?** A successful external-chat run can leave an active task in review without a maintained review path. Its automatic recovery then fails with `reviewed_chat_execution_binding_not_authorized`. The recovery context retains `chat:slack` but drops `wakeCommentIds`. **Expected behavior** An eligible recovery should retain its admitted message references and pass a new authorization check. Missing or revoked access must still prevent execution. **Steps to reproduce** 1. Finish a chat run whose task remains in review with no maintained review path. 2. Build the bounded review recovery from that run's context. 3. Dispatch the recovery through the reviewed-chat authorization check. The new regression tests fail before this patch. Related PR #13809 handles answered conversations that become idle. This patch handles recovery when the conversation remains active. ## What Changed - Retain the admitted message batch for external-chat review recovery. Derive the current comment reference from that batch. - Share the existing supported-provider selector between recovery and run-bound chat authorization. AgentMail remains on its separate email inbox path. - Keep the existing authorization check. Do not copy prior checkout, authorization, session, or prompt state. - Test Slack and Discord recovery, missing batches, revoked access, and changed task or company bindings. - Document the recovery authorization contract. ## Verification - The new tests reproduced the missing-message failure before the fix. - Six focused suites passed: 257 tests covering review recovery, reviewed-chat authorization, issue liveness, Slack lifecycle, comment-wake batching, and external-chat waits. PostgreSQL integration tests ran against a temporary local PostgreSQL database. - Focused test files: `review-path-recovery.test.ts`, `heartbeat-reviewed-chat-binding.integration.test.ts`, `heartbeat-issue-liveness-escalation.test.ts`, `slack-conversation-lifecycle.test.ts`, `heartbeat-comment-wake-batching.test.ts`, and `external-chat-wait.integration.test.ts`. - Server TypeScript check passed with scratch configuration that resolves this checkout's workspace packages. Existing dependency links point to another checkout; the default check reports stale shared-type errors. - `node scripts/check-module-boundaries.mjs` and `git diff --check` passed. - `git diff | gitleaks stdin --redact --no-banner` passed. The diff was also checked for private identifiers and user data. - [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35760757895) passed on the latest commit, including build, workspace typecheck, general and serialized tests, Rust checks, and all eight browser shards. The unchanged local-service readiness test and agent-chat page-load assertion passed when their failed shards were retried. The local-service suite also passed locally (6 tests). The first, superseded run lost a chat runner; its stuck browser job was cancelled to unblock the current run. - Full local workspace typecheck, tests, and build were not run. The machine has about 2 GiB free, and those commands include Rust builds. The full PR CI checks passed before merge. ## Risks Small context-construction change. Dispatch still checks current execution ownership and access for every admitted message. The recovery remains bounded to one attempt per consumed path. No schema changes. Existing failed runs are not retried by this patch. ## Model Used OpenAI GPT-6 (Codex), with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
213866fae0 |
chore(lockfile): refresh pnpm-lock.yaml (#13780)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.922.0-canary.8 |
||
|
|
83abfa46f6 |
feat(ui): enable Grok in Cloud agent setup (#13791)
## Thinking Path > - Paperclip manages AI agents and their execution settings. > - The new-agent flow selects an adapter before configuring credentials and a model. > - Cloud uses one adapter policy for the picker and direct setup links. > - That policy excludes Grok despite its existing adapter and xAI connection support. > - This change adds Grok to the Cloud policy. > - Cloud users can configure Grok with the existing subscription or API-key flow. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent creation on Cloud. **Current behavior** The Cloud picker offers Claude, Codex, and OpenCode. Direct Grok setup links also fail the shared adapter check. **Proposed behavior** Offer Grok alongside the existing choices. Use the existing xAI connection and managed sandbox setup. **Reason and benefit** Users can select an already-supported adapter through the Cloud creation flow. **Breaking changes** None. Loaded and enabled checks still apply. Other excluded adapters remain excluded. ## What Changed - Add `grok_local` to the shared Cloud creation policy. - Display four adapter choices in a 2×2 grid on desktop and mobile, and reuse the theme-aware provider mark on the connection step so Grok is visible in dark mode. - Verify picker navigation and both Grok authentication methods through sandbox setup, model testing, and agent creation. - Verify that probe and hire payloads carry the xAI connection binding without the entered API key. - Update the agent configuration specification. ## Verification - `cd ui && pnpm exec vitest run src/components/NewAgentDialog.test.tsx src/pages/NewAgent.test.tsx src/components/new-agent/AgentProviderConnection.test.tsx` — 69 tests passed. - Chromium checks against the production components and built stylesheet — four cards occupy two rows and two columns at 1280px and 390px; the connection step loads and displays the white Grok logo in dark mode and the black logo in light mode. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/ui build` — passed. - `pnpm check:token-gates` — passed. - `git diff origin/master...HEAD | gitleaks stdin --redact --no-banner` — no leaks found; manual diff review found no private identifiers or user data. - `pnpm -r typecheck` and `pnpm build` — blocked at the existing runner package because `cargo` is not installed locally. - `pnpm test:run` — started locally, then stopped after the equivalent CI suites passed. - Initial [PR CI](https://github.com/paperclipai/paperclip/actions/runs/35681351752) passed, including full typecheck/build, general and serialized tests, Rust checks, and all eight browser-test shards. One unchanged Cursor test timed out on the first attempt; its five-test file passed locally and the failed CI shard passed on retry. - The [latest CI run](https://github.com/paperclipai/paperclip/actions/runs/35683642812) passed build, typecheck, all general and serialized tests, Rust checks, and all eight browser-test shards. One unchanged local-service-supervisor test failed its HTTP readiness check on the first attempt; its six-test file passed locally, and the failed server shard passed on retry. - Live xAI login and model execution were not run; the setup tests mock provider calls. ## Risks Small UI policy change. The existing Grok adapter, authentication, and secret storage paths remain in use. No schema or control-plane change is required. Cloud must deploy a tenant-app release containing this change. Revert the policy entry to hide Grok from new-agent setup again. ## Model Used - OpenAI GPT-6 (Codex), with repository inspection, code editing, and shell-based verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.7 |
||
|
|
8725d6ce09 |
fix: make answered Slack conversations idle (#13809)
Settle published, successful Slack turns as Idle; resume the same conversation on an admitted message. Preserve unfinished work, delivery errors, and explicit dispositions. Verified through focused lifecycle/API/UI tests, full CI, and a real staging Slack conversation in the embedded browser. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.6 |
||
|
|
e3d8fb0876 |
feat: attest standard production images at full source commits (#13797)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators deploy its standard production container on several CPU architectures. > - Downstream image builders need to identify the exact source of their base image. > - A short commit tag does not provide signed source evidence. > - This pull request adds a full commit tag and signed image digest for canonical master pushes. > - Consumers can verify the source and compose from the immutable digest. ## Linked Issues or Issue Description **What existing behavior does this improve?** Publication of the standard multi-platform production image. **Current behavior** The Docker workflow publishes short commit tags and channel tags. It does not provide a signed standard-image contract tied to the complete master commit. **Proposed behavior** Canonical master pushes also publish `sha-<full-commit>` and attest the exact index digest after platform validation and an immutable-image orphan-reaping check. The signer certificate binds the source repository, commit, workflow and ref. Existing tags and the separate cloud producer remain available. **Reason and benefit** Downstream builders can prove the source of a standard base without adding their dependencies or repository details to the public workflow. No matching open issue or duplicate PR was found. ## What Changed - Add the canonical full-SHA tag without changing existing tag mappings. - Validate amd64 and arm64 descriptors and hash the exact registry response bytes and require its digest header to match. - Verify the immutable image and sign it with GitHub artifact attestations. - Run the contract tests in trusted PR verification and document the consumer contract. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/cloud-source-verification.test.mjs scripts/standard-image-contract.test.mjs`: 37 passed. - A read-only check against an existing published index returned its exact expected digest. - `actionlint -shellcheck='' .github/workflows/docker.yml .github/workflows/pr-trusted.yml`: passed. Normal ShellCheck reports only existing `ls` and word-splitting warnings. - `pnpm build`: passed locally with Cargo available. - `pnpm -r typecheck`: passed locally. - Full local Vitest was attempted: 8,260 passed, 14 failed, with 34 failing suites. The failures were missing embedded-PostgreSQL library aliases in this fresh install and existing macOS runtime-skill-cache rename errors. Native aliases are now restored. Rerunning the 33 affected database suites produced 577 passes and two unrelated AgentMail skill-root lookup failures (32 suites passed). The four directly failing database tests also pass independently. This is not a claim that the full local suite passed. - All final-head CI checks pass. One unrelated routine-route mock assertion passed on the single-shard retry; its 15 tests also pass locally. Greptile is 5/5 on this exact head, with no unresolved threads. - Actual signing requires a canonical master push. This draft PR does not publish trusted provenance. ## Risks The new attestation step requires OIDC and attestation write permissions in the merge job. Signing failure leaves the image available but without the new admission proof. Consumers must fail closed when proof is missing. Existing release tags, the legacy producer, and image retention remain unchanged. No database or application behavior changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, shell execution and test tools. The session does not expose a more specific model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.5 |
||
|
|
8326e33ada |
Fix oversized sandbox process launch payloads (#13793)
## Thinking Path > - Paperclip manages agent work and preserves context across retries. > - Sandbox ACP runs encode their command and environment into one launch value. > - A long continuation can make that encoded value exceed Linux's exec limit. > - The launch shell then exits before the agent can initialize. > - This PR transfers large command envelopes through a private temporary file. > - The agent receives the complete environment and can start normally. ## Linked Issues or Issue Description Refs #13777. **Bug description** Sandbox tasks with long retry context fail during ACP initialization with exit 127. The streamed bridge also drops the shell error that explains the failure. **Steps to reproduce** Launch the streamed sandbox process bridge on Linux with one valid 110,000-byte environment value. Its base64 command envelope exceeds the limit on one exec argument or environment string. Daytona reports `argument list too long: env`, then exit 127. **Expected behavior** The command envelope must not make a valid child environment too large to launch. A shell startup failure must retain its diagnostic in the run log. ## What Changed - Keep envelopes up to 64 KiB on the existing launch path. Upload larger envelopes in bounded chunks inside a mode-0700 session directory. Set the final payload file to mode 0600. - Read the file without following symlinks and delete it before spawning the child. Remove incomplete uploads on failure. Both streamed and polled bridges use the same envelope. - Preserve stderr when the launch shell fails before the wrapper emits a terminal event. Emit a fixed terminal error and shutdown acknowledgement if a payload cannot be read or parsed, without exposing its contents. - Cover large environments, file permissions, payload deletion, interrupted uploads, missing or malformed payloads, and startup diagnostics. Document the transfer and cleanup behavior. ## Verification - Both large-envelope regressions fail before the fix and pass after it. - Targeted bridge, ACP engine, real-spawn, and stdin-race checks pass: 370 tests. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - A gated live Daytona probe on the current sandbox image reproduced exit 127 with the old bridge. The fixed bridge launched the same command successfully. A second probe completed real Claude ACP initialization with a 110,000-byte context value. It did not run an agent task. All temporary sandboxes were deleted. - `pnpm -r typecheck` and `pnpm build` reach the unchanged Rust runner step and stop because this machine has no `cargo` executable. - Full CI passes on `ea816cf58d`: [run 35683853754](https://github.com/paperclipai/paperclip/actions/runs/35683853754). All 53 checks pass; two optional checks are skipped. The local full-suite run was stopped after equivalent CI suites passed; it has no final local result. - Greptile is 5/5 on `ea816cf58d`, with no unresolved review threads. The branch is mergeable. ## Risks Large envelopes require extra upload calls during startup. The temporary data stays inside the private session directory and is removed before child startup or during failure cleanup. Individual child environment values still obey the operating system's native limits. No migration or configuration change is required. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted checks; full-workspace limits described above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.4 nightly/v2026.922.0-nightly.0 |
||
|
|
c496acb570 |
Return client errors for known OAuth reconnect states (#13794)
Map missing OAuth refresh credentials and terminal reauthorization to HTTP 422. Preserve reconnect instructions and reporting of unexpected provider failures. All 363 focused tests and server typecheck pass. Required CI and review checks passed; Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1f3ff75d33 |
Recognize confirmed container loss when resuming Daytona leases
Daytona can retain a sandbox API record after its container disappears. Confirm the exact missing-container response with a bounded fresh read, then let the host apply its existing replacement and backup policy. Preserve unknown failures, recoverable errors, and identity mismatches. Verified 251 provider tests, 93 host lifecycle tests, plugin typecheck and build. Added 15 regressions. Read-only inspection confirmed the provider response and an existing verified backup without changing live state. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a7d3b17a97 |
Explain disabled Slack MCP app access during discovery
Recognize Slack's exact disabled-app response and return actionable setup instructions from catalog and health routes. Bound response parsing and keep unknown upstream errors reportable without exposing provider settings links. Verified 361 focused tests, server typecheck, and authenticated discovery. The three route regressions fail before this change and pass afterward. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.3 |
||
|
|
3c2bc4f546 |
ci: halve the isolated native Runner check by seeding it from the public master build cache (#13736)
## What Recurring CI health check (PAP-31): on the most recent fully-green PR run (`cb703ac`, run [35519542997](https://github.com/paperclipai/paperclip/actions/runs/35519542997)), the slowest check was **Compile isolated native Runner** at **452s** — ahead of the largest test shards (379s). Every run recompiled the full Rust dependency tree from zero, even though `docker.yml` already refreshes a **public** `mode=max` BuildKit cache (`ghcr.io/paperclipai/paperclip:buildcache-{amd64,arm64}`) on every master push, containing exactly these layers. This PR seeds **only the baseline build** with that registry cache, via anonymous pull. **Measured on this PR's own CI (which exercises the seeded path): the check completed in 123s, down from 452s — a 73% reduction, ~5.5 minutes saved per run.** ## Thinking Path Cost breakdown of the 452s from the job log: `cargo chef cook` dependency compile 214.7s, local cache export 43.1s, runner-core build 37.0s, metadata proof layer 36.8s, `cargo install cargo-chef` 36.4s, rebuild-verification build ~65s, setup/teardown ~20s. The dependency compile and toolchain layers are identical to what the production `docker.yml` build already caches publicly on every master push, so recompiling them here bought no signal — the check's real assertions live in the *verification* build, not the baseline. A first attempt used `actions/cache` plus a master `push` trigger, but the CI bot's GitHub App lacks `workflows` permission; the registry-cache approach is strictly better anyway (shared across PRs immediately, no 10GB Actions-cache quota pressure, no workflow change). ## What Changed - `scripts/check-docker-runner-cache.sh`: the baseline build now adds `--cache-from type=registry,ref=ghcr.io/paperclipai/paperclip:buildcache-{amd64|arm64}` (selected by host arch). `RUNNER_CHECK_SEED_CACHE` overrides the ref, or set it empty to force the old cold path. The script header documents the anonymous external read. - `.github/workflows/docker-runner-check.yml` (comment-only): the stale "no external cache" note now describes the anonymous GHCR seed and the verification build's local-cache-only isolation. This was pushed in a follow-up commit with workflow-edit permissions; the original CI-bot token could not touch workflow files. No Dockerfile stages or verification assertions changed. ## Verification - This PR's own `Compile isolated native Runner` check runs the seeded path (the script is in the workflow's trigger paths): **passed in 123s** vs the 452s baseline. - The rebuild-verification semantics are untouched: it still runs on a **fresh builder** importing **only the local cache exported by this run's baseline**, so it proves exactly what it proved before — that the runner image rebuilds reproducibly from this run's own exported layers. - Verified `ghcr.io/paperclipai/paperclip:buildcache-amd64` is anonymously readable (unauthenticated manifest pull succeeds), so the check gains no credential or secret dependency. ## Risks - **Stale or missing seed cache:** if the GHCR ref is unreachable, private, or garbage-collected, BuildKit logs a warning and falls back to the pre-PR cold compile — the check gets slower, never wrong. `RUNNER_CHECK_SEED_CACHE=""` restores the cold path explicitly. - **Cache trust:** the seed only accelerates the *baseline* build; the verification build still runs on a fresh builder against only this run's locally exported cache, so a stale or poisoned registry cache cannot make verification pass spuriously. The ref lives under `ghcr.io/paperclipai/*`, written only by repo CI on master pushes. ## Model Used Claude Fable 5 (`claude-fable-5`) via Paperclip agent **Bender (Fable)**, issue PAP-31. --------- Co-authored-by: Bender (Fable) <bender-fable@paperclip.local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ded156a904 |
Keep interaction continuations scoped to the target task
Separate an interaction's producer from explicit resume history. Do not import another task's comments or results. Filter newly captured foreign origins and recover older inherited origins only when the saved producer context and comment row prove their source. Keep missing context and company boundaries fail-closed. Verified 46 continuation tests, 201 related recovery tests, and server typecheck. Added 20 database regressions for provenance and scope guards. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3c70b3d008 |
docs(release): curate stable notes for the 2026.921.0-beta.1 promotion (#13785)
## Thinking Path > - The stable promotion reads `releases/beta/v<beta-version>.md` from `master` and publishes it as the GitHub Release body. > - Beta `2026.921.0-beta.1` (source `8f8a0ab7`, 77 commits since v2026.916.0) just published and starts its 3-day soak, headed for a stable around 2026-09-24. > - This PR rewrites the auto-generated draft into release voice during the soak, per the release checklist. ## Linked Issues or Issue Description Notes for the stable to be promoted from [`2026.921.0-beta.1`](https://github.com/paperclipai/paperclip/releases) — planned as `v2026.924.0`. **If the promotion date slips past 2026-09-24, bump the title and Released date before dispatching stable.** ## What Changed - Rewrote `releases/beta/v2026.921.0-beta.1.md` from the 77-entry scaffold into the standard release-notes structure: overview, Breaking Changes (legacy Composio broker retirement [#13758](https://github.com/paperclipai/paperclip/pull/13758); full-auto execution defaults [#13686](https://github.com/paperclipai/paperclip/pull/13686)/[#13693](https://github.com/paperclipai/paperclip/pull/13693)), Highlights, grouped Fixes, Improvements, and an Upgrade Guide (migrations `0280`–`0283`, `enableMcpAggregators` flag). - Folded out release-infra noise: CI-only PRs, test-only PRs, release-notes bookkeeping commits, and the hide/restore Google-connector pair ([#13551](https://github.com/paperclipai/paperclip/pull/13551)/[#13552](https://github.com/paperclipai/paperclip/pull/13552), net zero). - Noted that the composer fix ([#13562](https://github.com/paperclipai/paperclip/pull/13562)) already shipped early as stable 2026.916.1. ## Verification - Docs-only; every PR number cited was cross-checked against `git log dffc2b3ca..8f8a0ab7e`. - Migration range `0280`–`0283` verified against `packages/db/src/migrations` diff and attributed per commit. ## Risks None — documentation only. Content can be edited freely during the soak; the stable preflight only requires the file to exist on `master`. ## Model Used Anthropic Claude Fable 5 (`claude-fable-5`) via Claude Code, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (docs-only; no tests apply) - [x] I have added or updated tests where applicable (none apply) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending on this fresh PR) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
cf2742ac48 |
docs(release): canonicalize v2026.916.1 release notes (#13774)
## Thinking Path > - Stable `2026.916.1` just published from beta `2026.921.0-beta.0`; its curated notes live at `releases/beta/v2026.921.0-beta.0.md` on `master`. > - The release workflow's `canonicalize_stable_notes` job pushed this branch, which `git mv`s them to the canonical stable path. > - Merging restores the canonical `releases/` layout, matching every prior stable. ## Linked Issues or Issue Description Follow-up to the [v2026.916.1](https://github.com/paperclipai/paperclip/releases/tag/v2026.916.1) stable release; notes were curated in [#13766](https://github.com/paperclipai/paperclip/pull/13766). ## What Changed - `git mv releases/beta/v2026.921.0-beta.0.md releases/v2026.916.1.md` (workflow-generated, content unchanged). ## Verification - Rename only; the GitHub Release body was already published from this content. ## Risks None — documentation layout only. ## Model Used Anthropic Claude Fable 5 (`claude-fable-5`) via Claude Code, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (rename-only; no tests apply) - [x] I have added or updated tests where applicable (none apply) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending on this fresh PR) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending) - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>canary/v2026.922.0-canary.2 |
||
|
|
878734a061 |
fix(tools): treat OAuth sign-in challenges as client errors (#13786)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connected apps can require OAuth sign-in before they list their tools. > - Remote discovery recognizes this condition as `oauth_challenge`. > - Discovery and catalog refresh currently return it as HTTP 502. > - The server error handler reports that response as a crash. > - This pull request returns HTTP 422 for the explicit sign-in challenge. > - The operator keeps the sign-in instructions, while unexpected upstream failures remain reportable. ## Linked Issues or Issue Description **What happened?** Connecting a remote MCP app that answers with a recognized OAuth challenge returns HTTP 502. Reading an empty catalog or explicitly refreshing it does the same. The server error handler then sends the expected sign-in condition to error monitoring. **Expected behavior** A known sign-in requirement returns HTTP 422 with the existing `oauth_challenge` code, message, and setup/reconnect links. An unexplained upstream HTTP 400 or an unavailable upstream service still returns 502 and reaches error monitoring. **Steps to reproduce** 1. Configure a remote MCP app that returns HTTP 401 with a Bearer challenge. 2. Connect the app, read its empty catalog, or request a catalog refresh. 3. Observe HTTP 502 and a server error report before this change. **Paperclip version or commit** Reproduced on `6de50ba594b15efaa3eae6ed869cd39b3a436456`. **Deployment mode** Server with remote MCP connections. The regression coverage uses local PostgreSQL and mocked upstream HTTP responses. Related: #9750 addresses MCP initialization and session recovery. It does not change the classification of this recognized sign-in condition. Targeted searches found no duplicate classification PR. ## What Changed - Return 422 for `oauth_challenge` from discovery and from catalog health-error normalization. - Preserve the existing structured error and remediation links. - Test automatic empty-catalog reads and explicit refreshes. Verify that OAuth challenges produce no Sentry capture and that upstream 400/503 failures still do. - Update the direct-connect and blocked-redirect expectations and document the monitoring behavior. ## Verification - Before the fix, three sign-in route regressions fail with 502 instead of 422; all four upstream-error controls pass. - After the fix, all 339 tool-access and error-handler tests pass, including authorization and redirect protections. - These suites ran against disposable Homebrew PostgreSQL 16.14 through the existing test-constructor seam. The temporary setup and config remain outside the repository. CI uses the ordinary embedded PostgreSQL setup. - Direct server `tsc --noEmit` passes. - Full build and recursive typecheck were attempted; the Runner Rust step cannot run because `cargo` is absent on this machine. - Full `pnpm test:run`: 8,211 passed, 14 failed, 4,759 skipped. The 36 failed files match the existing embedded PostgreSQL startup/cleanup and macOS runtime-cache `EACCES` limitations. The changed database-backed service suite passed separately with local PostgreSQL. - Greptile: 5/5 with no unresolved comments. Seven CI workers received a simultaneous shutdown signal; the failed jobs are being retried through the normal workflow. Other completed checks passed. ## Risks Low risk. Clients now receive 422 instead of 502 for the explicit `oauth_challenge` condition. The code, message, and remediation links remain available. No permissions, credential handling, OAuth discovery rules, retry policy, or schema change. Other upstream failures retain their existing behavior. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, code editing, and test execution. The session does not expose an exact model snapshot or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (339 service and error-handler tests; full workspace limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.1 |
||
|
|
6de50ba594 |
fix(sentry): carry the deployment environment to the browser (#13784)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators can enable Sentry for the server and the signed-in browser. > - The server SDK reads `SENTRY_ENVIRONMENT` from the process environment. > - The browser receives its DSN through the session response, but receives no environment. > - A browser in staging therefore reports errors under the SDK's production default. > - This pull request passes the configured environment through the existing session and monitoring gate. > - Browser errors then identify the deployment environment while preserving the existing privacy settings. ## Linked Issues or Issue Description **What happened?** With `SENTRY_ENVIRONMENT=staging`, browser exceptions are tagged `production`. This can send errors to the wrong environment's alerts and makes deployment follow-up unreliable. **Expected behavior** The browser uses the server's configured Sentry environment. A reused image works in either staging or production. A signed-out browser still sends no events. **Steps to reproduce** 1. Configure a frontend Sentry DSN and `SENTRY_ENVIRONMENT=staging`. 2. Sign in and capture a browser exception. 3. Inspect the event environment. Before this change, it is `production`. **Paperclip version or commit** Reproduced on `a3749aac4680a901fa0fe1cc898907887abc9908` with the real browser SDK and a local test transport. **Deployment mode** Authenticated server and browser with optional Sentry monitoring enabled. No duplicate environment-attribution issue or pull request was found in the targeted GitHub search. ## What Changed - Add `sentryEnvironment` to the authenticated session response and shared schema. The optional field supports a newer browser reading an older server response. - Pass the environment to the browser SDK. An environment change restarts the client through its existing serialized lifecycle. - Cover environment attribution with a real SDK event, session authorization, unchanged-session refetches, environment changes, and legacy responses. - Document configuration and compatibility. Keep the loaded bundle's release identity and existing privacy filters. ## Verification - The regression test emits `production` for a requested staging environment before the fix. - Focused route, schema, browser lifecycle and real-SDK tests: 69 pass. - UI and shared-package typechecks, direct server `tsc --noEmit`, and token gates pass. - Full `pnpm build` and `pnpm -r typecheck` were attempted. Both stop at the Runner Rust step because `cargo` is absent on this machine. - Complete UI suite: 6,540 tests pass in 626 files. - Full `pnpm test:run`: 8,210 passed, 14 failed, 4,753 skipped; 36 files fail due to embedded PostgreSQL startup/cleanup and macOS runtime-cache `EACCES`. These match the existing local baseline; none touch the changed behavior. - Greptile: 5/5, no unresolved review threads. Linux CI has passed Build, Typecheck + Release Registry, and the completed test jobs so far. Remaining jobs are running or queued: the AWS runner provisioner is retrying EC2 CreateFleet `InternalError` responses. Full results will be recorded before merge. ## Risks Low risk. This adds one optional session field and changes Sentry attribution only. No migration or new monitoring opt-in is introduced. Missing settings keep the browser SDK default. Agent and unauthenticated requests still receive 401 without monitoring settings. Existing loaded browser bundles keep their old behavior until refreshed. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, code editing, and test execution. The session does not expose an exact model snapshot or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass (focused and full UI suites pass; full-root environment failures documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5842185e4f |
fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat needs reliable native execution before native runners become the onboarding default. > - Existing stories covered idle reassignment and controller restart, but not an executing worker handoff or worker process loss. > - Status answer tests also need to reject stale claims and invented facts. > - This pull request adds six opt-in full-stack cells with independent state assertions and retained evidence. > - The probes exposed a misleading Retry across server projection and recovery-banner paths; the fix reports the blocked recovery honestly. > - The tests preserve failures without changing recovery policy, production prompts, or onboarding defaults. ## Linked Issues or Issue Description Refs: #13762. Related: #13765 (Retry targets the latest failed attempt), #13753 (task context ownership), #13746 (native recovery work). ## What Changed - Add active reassignment with saved draft and plan preservation, old-worker cancellation, and successor completion checks. - Preserve recovery-needed projection when native cleanup fails before its coordinator exists, refuse a generic retry that would immediately fail again, and replace the recovery banner's misleading Retry with Inspect run. - Add verified local worker process loss with a required successful continuation; retain a failing qualification result when recovery is unavailable, while independently verifying the UI/API refuse doomed retries. - Add two-turn factual answer checks for current blockers, stale claims, inactive backlog work, and unknown facts. Retain prose for separate semantic review. - Add positive and negative oracle calibration and document fault isolation, cleanup, billing, and qualification limits. ## Verification - Eval TypeScript check passes. - All 442 eval support tests pass locally. The 89 focused server tests and server typecheck pass. Six recovery-banner UI tests and token gates pass. - Initial new-cell campaign: https://github.com/paperclipai/paperclip/actions/runs/35657128077. All six results are retained; four failed on fixture-contract issues and two exposed real worker cleanup quarantine. - All 26 existing native onboarding cells: https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26 passed on mastercanary/v2026.921.0-canary.10 |
||
|
|
c221589b69 |
fix: close sandbox process proxies after remote exit (#13777)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - Sandbox agents use a local proxy to exchange ACP messages with a remote process. > - ACP keeps the proxy input stream open while it waits for a reply. > - The proxy received a remote exit but kept that input stream open. > - This left the proxy alive and could block later task retries. > - This pull request closes the input handle after a terminal remote event and lets final output drain. ## Linked Issues or Issue Description **What happened?** A sandbox process could exit while its local proxy stayed alive. During startup, this could appear as a handshake timeout. A later task retry could then stop because the previous process was still alive. **Expected behavior** The proxy should exit with the remote process status, even when ACP has not closed its input stream. Final output and diagnostics should arrive before the proxy closes. **Steps to reproduce** 1. Start a sandbox process-session bridge with a child that writes output and exits. 2. Start the generated local proxy and keep its input stream open, as ACP does during startup. 3. Observe that the proxy stays alive after the remote process exits. **Paperclip version or commit** Reproduced on master at `846336e5a0`. Related: #13272 covers legacy sandbox startup recovery. #13765 covers the task retry UI. This change fixes the proxy process lifecycle. ## What Changed - Close the proxy input handle on remote exit or error, and preserve the exit status. - Stop forwarding input after the terminal event. - Wait for process close in the test helper so output is fully drained. - Test successful exit, failed exit, and spawn failure with input held open in both output modes. Check complete delivery of 208 KiB of final output. ## Verification - Before the fix, all four remote-exit regression cases timed out. - After the fix, all six terminal-event cases pass. - Targeted runtime suite: 362 tests pass across the sandbox bridge, stdin queue races, real ACP spawn, and ACP engine suites. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - Full workspace tests hit local platform failures: embedded PostgreSQL fails during initialization, and runtime skill-cache tests fail with `EACCES` while renaming read-only directories on macOS. The affected server code is unchanged in this PR. Those suites pass in CI. - `pnpm -r typecheck` and `pnpm build` stop at the existing Rust runner step because this machine has no `cargo` executable. The CI typecheck and build gates pass. - [CI run 35658635907](https://github.com/paperclipai/paperclip/actions/runs/35658635907) passes all required gates. The first browser shard run passed all 12 tests but hit a GitHub 403 during report upload; the rerun passed, including upload. - Greptile scored commit `9c3c89148e` 5/5 with no review threads. The branch is mergeable. ## Risks The proxy now closes its input immediately after a terminal remote event. Tests cover final output delivery and preserve nonzero exit codes. Existing orphan processes still require cleanup or a server restart. Live sandbox startup needs validation after deployment; this change addresses the confirmed proxy hang and does not establish why an earlier remote process stopped. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted checks; full-workspace limits described above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (existing behavior restored; no documentation change needed) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a3749aac46 |
fix(deps): share Lezer node properties across editor languages
fix(deps): share Lezer node properties across editor languages The installed graph gave syntax highlighters @lezer/common 1.5.1 and language parsers 1.5.2. Their independent NodeProp counters collided, so highlighting ordinary code read unrelated metadata as tags and crashed with tags-is-not-iterable. Override @lezer/common to one compatible version in both manifests. Extend the installed-graph check and exercise Python, JavaScript, HTML and SQL highlighting through the editor dependencies. All four examples failed before the override and pass with it. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.921.0-canary.9 |
||
|
|
846336e5a0 |
test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path > - Paperclip lets people manage agents through ongoing conversations. > - Chat users can change instructions while a provider is already working. > - Existing chat evals wait for each turn to settle before the next message. > - They cannot prove delivery during active work or the saved effect of a correction. > - Existing fixtures also enable Agent Chat through the API rather than the settings UI. > - This PR adds bounded browser workflows and checks their persisted outcomes. ## Linked Issues or Issue Description Refs #13741, #13752, #13750. **What happened?** The chat suites cover planning, delegation, status, and recovery. They lack active-turn follow-ups and the experimental settings lifecycle. A sequential conversation can pass even if messages sent during work are lost. **Expected behavior** A follow-up submitted during a provider turn survives and affects the final reply. A changed launch day appears in the saved plan. Disabling Agent Chat rejects new messages while preserving history; re-enabling resumes the same conversation. **Steps to reproduce** Run the explicit `agent-chat-stories` suite. It selects three local cases for each native Claude and Codex profile. An ordinary provider command waits for a fixture brief file so the browser can send the follow-up at an observed active-run boundary. ## What Changed - Add six opt-in Product E2E cells for settings, active follow-ups, and plan corrections. - Drive experimental settings through the UI and verify disabled sends are rejected by the public API. - Use a bounded file wait in the actual isolated agent workspace, with provider-written readiness and an undisclosed brief reference. - Grade persisted user messages, final replies, native run outcomes, and exact saved plan fields. - Accept active-turn steering or one queued successor; reject lost input, duplicate input, and stale outputs. - Allow one steered run or two sequential runs throughout the shared harness, while preserving exact counts for other cases. - Require a single marker-bearing response attributed to the final provider run. - Unload the development browser client before restarting the server, avoiding reconnect/navigation races without weakening the post-restart memory check. - Add browser regressions for restart isolation and asynchronously saved settings switches. - Document prepared-agent setup, native onboarding limits, and the separate API-tool rollout gate. ## Verification - Eval TypeScript check passed. - Eval support suite: 436 tests passed in 39 files. - New oracle calibration: six tests passed, including plausible invalid outcomes. - Browser support regressions: seven tests passed; the restart regression was observed failing before the fix. - Catalog discovery selects exactly six local native cases and leaves default paid selection unchanged. - [Consolidated existing native chat report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/): master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a browser navigation timeout across restart; the page request returned 200 and the chat rendered. - [Nine targeted restart/replay cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/) passed on `1fe2fe275`, including the original failure, across native Claude/Codex and local/Daytona; all cleanup passed. - [Initial six-story campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817) retained all six failures: asynchronous switch assertions, unavailable fixture paths, and rich-text escaping in raw command comparisons. The corrected fixtures preserve the same behavioral assertions. - [Six-story campaign v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/) on `8232773a0`: 4/6 passed (both settings cases and both Claude interruptions). Codex could not see the host-temp fixture outside its workspace; this failed before follow-up delivery was exercised. - [Four affected interruption cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/) all passed, including cleanup, on definition v3 / `ad6ac0545`. Files live inside the actual agent workspace and the observed run workspace is verified. Both providers saved Friday in the real plan with the undisclosed brief reference; follow-ups persisted while the original run was active. Together with both unchanged settings cases from v2, all six new scenario variants have passing live evidence. - Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful checks, two intentional skips, zero pending/failing checks; mergeable and clean. Fresh Greptile 5/5, zero unresolved findings. - Full typecheck, tests, build, and browser CI passed remotely. One earlier head encountered a signoff-policy browser timing failure; the final head passed that shard. - Local pnpm wrapper could not fetch its version/signature metadata in the restricted environment; local eval checks used the installed Node executables. Repo-wide validation was completed by GitHub Actions. ## Risks These are eval-only changes. The file wait is a timing fixture in the isolated agent workspace, not a production runner hook. Native Codex host-filesystem isolation stays unchanged. It has a two-minute limit and is released in `finally`. The prepared-agent settings case is not full native onboarding: the wizard currently offers legacy adapters. The disabled-entry assertion uses full document navigation, which clears the prior React Query cache; preserved history is checked through the public API and re-enabled chat. No production prompt, rollout default, adapter behavior, or credential policy changes. Active-task reassignment and worker-crash recovery remain outside these new cases. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f8a0ab7ef |
docs(release): curate stable notes for v2026.916.1 (#13766)
## Thinking Path > - Paperclip publishes stable releases from curated notes: the stable promotion reads `releases/beta/v<beta-version>.md` from `master` and publishes it as the GitHub Release body. > - Beta `2026.921.0-beta.0` was just published from a candidate branch carrying the 2026.916.0 source plus the cherry-picked composer fix #13562, headed for patch stable `2026.916.1`. > - Its `draft_stable_notes` job pushed the auto-generated scaffold; this PR rewrites it into release-notes voice so the stable promotion's preflight finds curated notes. ## Linked Issues or Issue Description Notes for the upcoming `2026.916.1` patch stable. The fix being shipped is [#13562](https://github.com/paperclipai/paperclip/pull/13562) (fixes #13561, the desktop chat composer send button starting out disabled). ## What Changed - Rewrote `releases/beta/v2026.921.0-beta.0.md` from the auto-generated commit list into the patch-release format used by `releases/v2026.831.1.md`: what the patch is, the user-facing composer fix, the rode-along document-conflict classification repair, and an upgrade guide (no migrations, no configuration changes). ## Verification - Docs-only change; no code paths affected. - Format matches the prior patch release notes (`releases/v2026.831.1.md`). - The stable promotion will resolve this file from `master` for beta `2026.921.0-beta.0` and canonicalize it to `releases/v2026.916.1.md` after publishing. ## Risks None — documentation only. If the wording needs adjusting during review, the stable dispatch simply waits until this is merged. ## Model Used Anthropic Claude Fable 5 (`claude-fable-5`) via Claude Code, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (docs-only; no tests apply) - [x] I have added or updated tests where applicable (none apply) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending on this fresh PR) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>v2026.1001.0 canary/v2026.921.0-canary.8 beta/v2026.921.0-beta.1 nightly/v2026.921.0-nightly.1 |
||
|
|
df66219780 |
fix(ui): retry the latest failed task attempt (#13765)
## Thinking Path > - Paperclip manages work done by AI agents. > - The task thread lets an operator retry a failed run. > - Legacy runs with transcript output did not get a failure marker. > - The thread could therefore offer Try again for an older failure. > - The server correctly reused that failure's existing retry, even when it had already failed. > - This change keeps the latest failure actionable and reports stopped retry responses to the operator. ## Linked Issues or Issue Description **What happened?** Try again could return success without starting work. An initial setup failure had an empty transcript. Its later retry produced output and failed. Only the initial failure had a retry marker, so the button kept requesting the initial failure's already-failed successor. **Expected behavior** Try again targets the latest failed attempt. An already-stopped retry response shows an error and refreshes the task's run state. **Steps to reproduce** 1. Start a legacy adapter task that fails before producing transcript output. 2. Retry it. Let this attempt produce output and a final failure comment before it fails. 3. Click Try again in the task thread. 4. Before this fix, the click targets the original failure and replays the stopped successor. **Paperclip version or commit** Reproduced against `1483bb8bcf`; the regression is also present on the branch base `8813a50105`. **Deployment mode** Authenticated server with a legacy adapter. The bug is in the shared task UI and retry API client. Related: #11650 adds a different recovery-notice action. This change fixes failed-run markers and retry response handling. Searches found no duplicate of this failure case. ## What Changed - Render legacy failure markers even when the run has a transcript or final comment. - Keep later cancelled automatic retries from replacing the failed run's retry action. - Reject already-stopped retry responses in the API client so existing error feedback appears. - Refresh task run queries after both successful and failed retry requests. - Add regression coverage for failed and timed-out attempts, execution gates with output, and retry response states. ## Verification - Before the fix, the new regression tests failed: two selected the original failure, and four accepted a stopped successor as success. - Targeted task-thread, retry API, marker, and issue-page tests: 279 passed. - `pnpm --filter @paperclipai/ui typecheck`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `pnpm check:token-gates`: passed. - Full UI suite: 6,529 tests passed across 626 files. - [CI run 35650385023](https://github.com/paperclipai/paperclip/actions/runs/35650385023): all 53 checks passed, including full workspace build, typecheck, unit/integration suites, and browser tests. The redundant local full-workspace test run was stopped after CI passed; it is not counted as a completed local pass. - Greptile: 5/5 on `608ee58c99`, with no review threads or unresolved comments. The branch is mergeable. - `pnpm -r typecheck` and `pnpm build` were attempted. Both stop at the Runner's Rust checks because this host has no `cargo`. The full workspace checks passed in CI. ## Risks Low risk. This changes UI presentation and response handling only. The server's exact-retry idempotency, authorization, execution ownership, and recovery gates remain in place. No schema changes or live task mutations. Existing documentation describes this retry action; the fix restores that behavior. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted checks; full-workspace limits described above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (existing behavior restored; no documentation change needed) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8813a50105 |
feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path > - Paperclip manages agent work as tasks and runs. > - GitHub chat brings repository conversations into those tasks. > - A review bot needs the assigned agent, its authority, and governed provider tools. > - The existing channel connection did not supply that review workflow or a complete setup journey. > - This pull request adds GitHub App setup, account access, event prompts, task-bound review tools, and exact-commit checks. > - Operators can inspect each review through the same task, run, and activity systems. ## Linked Issues or Issue Description **Subsystem affected** GitHub chat, governed connection tools, task execution, shared/database contracts, and connector setup UI. **Problem or motivation** Operators need a GitHub review bot that runs their assigned Paperclip agent. Mentions and PR events must preserve task ownership and requester authority. Provider publication must use the bot App identity and enforce the configured permissions. **Proposed solution** Extend the existing GitHub chat connector with resumable App onboarding, linked-member and sponsored-guest access, editable event prompts, and governed review operations. Validate structured assessments on the server and compute a stable Paperclip Review check for the exact head commit. **Alternatives considered** A separate review scheduler would duplicate Paperclip execution and permissions. Reusing personal GitHub credentials would change the bot identity and credential boundary. **Roadmap alignment** This extends the existing Connected Apps and governed-tool infrastructure. The project owner requested and approved this design. Related PR #8645 imports external Codex review feedback; this change runs an assigned Paperclip agent and publishes its results through the existing chat connector. ## What Changed - Include the current Paperclip instance origin in the copied setup prompt. Storybook uses its configured Paperclip origin; callback parameters and URL credentials are excluded. - Add a Claude/Codex copy button in the real setup and Storybook opening step. Its detailed prompt asks four setup questions and guides embedded-browser setup, verification, and optional required checks. Clipboard failure exposes selectable instructions. - Add a tutorial that explains why App installation, review scheduling, and required checks are separate choices. - Add manifest registration, an existing-App path, separate installation and repository selection, repository refresh, and explicit account confirmation. - Add low-trust agent guidance, effective capability verification, member selection, and explicit restricted guests with a sponsor. - Add configurable PR events, prompts, repository overrides, rating thresholds, and separate formal-review permissions. - Give the assigned agent governed App tools to read PRs, comment, begin an assessment, submit findings, and optionally submit a formal review. - Bind review history, root PR events, and inline replies to ordinary tasks. Deduplicate deliveries/findings and reject stale publication. - Link check Details to the underlying task on the current trusted hostname, or to Reviews before task creation. - Add schema migration 0283, API contracts, production UI, and 49 interactive Storybook states. - Repair local lease recovery. Keep the Cloud Dockerfile identical to master; no provider-pack layer or runtime-default environment variable is added. - Retry only rolled-back wake-admission transactions after transient endpoint-lock contention. A deterministic held-lock regression proves one accepted wake. ## Verification - Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero diff against master. Final workspace typecheck and build passed. The new PostgreSQL migration regression passed and preserves existing relation and constraint identities after replay. - Greptile reviewed this exact head at 5/5. There are zero unresolved review threads and no merge conflicts. - All current-head checks are green: 54 passed and two conditional Storybook jobs skipped. This includes complete server/workspace test suites, build, typechecks, policy checks, Runner suites, browser suites, and security status. One timing-sensitive callback-ordering test passed in isolation and its CI shard passed one retry. The duplicate local full-suite run was stopped after CI completed; it is not counted as a local full-suite pass. - Before the final Slack rebase and migration renumbering, 186 focused GitHub tests, 14 native bootstrap cases, token gates, and Storybook build passed. The final rebase retained the new Slack communication guidance. - The embedded-browser setup test copied the full detailed prompt, including the configured Paperclip instance URL. Desktop and narrow layouts were checked. Component tests cover successful copying and clipboard failure with selectable text and retry. - Live local and hosted GitHub acceptance evidence refers to application revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks exercised issue mentions, automatic PR reviews, inline findings, repeated mentions, task continuation, and failing-to-passing checks after a push. The Storybook agent generated, built, and browser-rendered pages; missing acceptance text failed, matching text passed, and broken JSX produced an incomplete result. - Live cases also covered independently disabled push events, prompt injection, duplicate signed deliveries, rapid pushes, stale-result rejection, finding deduplication, and restart recovery. Formal reviews were denied while disabled and published only after explicit enablement. Check Details links pointed to the underlying task on the trusted hostname. - Those hosted native Claude runs used the provider-pack layer now removed from this PR. They do not prove native Claude works on the standard Cloud image. A replacement hosted native Codex run is not yet verified: the disposable QA tenant has only an Anthropic AI connection. No new staging or production deployment was made for the packaging removal. - Required-check merge enforcement could not be tested because the private disposable repository's GitHub plan rejected the rules configuration. Published success/failure/incomplete check states were verified directly. ## Risks - Latest master allocated migration 0282 to Slack. The GitHub migration is regenerated as 0283 with replay-safe table/index/constraint creation; a PostgreSQL regression verifies existing relations and constraints are preserved. Existing preview tenants remain subject to the fleet migration-history compatibility preflight; no bypass is introduced. - Migration 0283 adds company-scoped configuration, registration, review, and publication records. Existing connections retain their behavior until reviews/tools are enabled. - Signed webhooks and expiring registration state remain required. Hosted installations also need the companion narrow Cloud gateway exemptions. - Agent assessments can be incomplete or wrong. The server enforces coverage/result structure, current-head publication, rating policy, and separate formal-review permission; it does not replace code-review judgment. - No Cloud image packaging changes are included. Remote native ACPX/Claude and OpenCode retain their existing operator-supplied provider-pack prerequisite. Native Codex and Codex with managed MCP tools do not require that pack. Earlier staging deployment evidence refers to its stated revision, not this packaging-removal head. Production rollout and merging remain outside this change. ## Model Used OpenAI GPT-6 through Codex, with repository, code execution, API, and embedded-browser tools. The exact serving model ID and context-window size were not exposed by the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.921.0-canary.7 |
||
|
|
d9b3a5653e |
feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people use the same tasks and agent tools from external conversations. > - Agents need communication guidance that fits the conversation medium. > - That guidance belongs in the original task context, without repeated instructions on each turn. > - Connection owners also need clear settings and a consistent way to remove a connection. > - This pull request adds initial Slack guidance, optional connection instructions, and chat connection menus. > - The benefit is clearer Slack replies with the existing Paperclip workflow and permissions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent replies in Slack and chat connection management in the Apps catalog. **Current behavior** Slack tasks do not carry a saved communication profile. The catalog shows a separate Manage button and does not offer removal on every chat connection row. **Proposed behavior** Save Slack guidance when a new conversation creates a task. Restore that original guidance when a model session is rebuilt. Do not append it to ordinary follow-ups. Expose optional additional instructions in Slack Settings. Put Manage and Remove connection in a three-dot menu for all chat providers. Keep Finish setup visible for drafts. **Reason and benefit** Small answers fit in Slack. Substantial deliverables use ordinary document or artifact tools with a useful Slack summary. Connection settings apply to new tasks and cannot change permissions. Users can remove both active and unfinished chat connections from the catalog. **Breaking changes** Two additive database columns store endpoint preferences and the initial conversation snapshot. Existing endpoints default to empty preferences. Existing conversations keep their original behavior. Non-Slack guidance is unchanged. Related public context: https://github.com/paperclipai/paperclip/pull/13741 improves native chat recovery. This change adds communication context to those existing execution paths. A search found no duplicate communication-guidance PR. ## What Changed - Add a provider-guidance registry, enabled for Slack first. - Persist optional endpoint communication instructions and capture an immutable snapshot when a conversation creates a task. - Resolve guidance from the verified company-scoped connection. Restore it for fresh native and legacy sessions without per-turn reminders, extra model calls, or extra context queries. - Add the Slack Settings field, validation, audit coverage, and Storybook save/error states. - Add Manage and Remove connection menus for all seven chat providers. Keep the draft setup button. Require removal confirmation and allow retry after failure. - Add regression coverage, an active/draft menu story, and connector documentation. ## Verification All CI checks are green for |
||
|
|
b82661b561 |
refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path > - Paperclip manages agents and their access to external tools. > - Connectors expose these tools through a governed MCP gateway. > - PR #13755 added a direct Composio MCP connection behind the experimental MCP aggregators flag. > - The old project API-key broker still created toolkit child connections and showed a separate Services tab. > - Keeping both paths leaves obsolete setup and session code in the product. > - This change removes the broker and preserves direct MCP setup, credentials, permissions, and execution. > - Saved legacy records fail closed and remain available for explicit removal. ## Linked Issues or Issue Description Related: #13755. This retirement supersedes the legacy-path fixes proposed in #12630, #12632, #12634, and #12906. It does not close those PRs. **What existing behavior does this improve?** Composio connector setup, management, and runtime dispatch. **Current behavior** Composio offers both direct MCP and a project API-key broker. The broker mints sessions and creates one child connection per toolkit. **Proposed behavior** Offer only direct MCP. Remove the toolkit Services UI, REST routes, API client, and session broker. Block saved legacy parent and child records from discovery, execution, health checks, reconnect, and OAuth. Preserve their records and credentials until the operator removes each connection. **Reason and benefit** The direct MCP connector becomes the single supported Composio workflow. Provider accounts remain managed in Composio. ## What Changed - Remove the API-key catalog method and its generated-source definition. - Delete Composio broker clients, session creation, account synchronization, child lifecycle, and toolkit routes. - Remove the Services tab, service rows, child provenance, and cascade-removal controls. Keep Vercel provenance intact. - Retain a shared retirement guard for stored legacy records. Show Retired status and replacement/removal guidance in the connection list and details; hide obsolete runtime controls. - Preserve the experimental MCP aggregators flag and direct MCP infrastructure. - Replace broker fixtures with retirement tests and extend direct Composio catalog/reconnect coverage. ## Verification - Focused shared, server, and UI tests passed with one worker. Server retirement tests use a name filter; no full local test suite was run, as requested. - Server and UI TypeScript checks passed. - Token gates and UI build passed. - Real browser: opened the saved Composio connection, refreshed all 11 tools, and ran the provider's read-only GitHub account-list operation through the standard Test dialog as an agent. The provider returned success using the existing OAuth credentials. - See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live evidence. - Storybook build passed. A fresh real agent used `COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway audit records confirm both calls succeeded. - Browser retirement check: a credential-free legacy fixture showed the guidance, opened the direct MCP replacement flow, and was removed through the standard confirmation. - Focused regressions for the experimental settings copy and exact OpenAPI route coverage passed. All latest-head CI checks passed (54 successful, two intentionally skipped); Greptile scored 5/5 with no unresolved review threads. The PR has no merge conflicts. ## Risks This intentionally breaks the old Composio project API-key and child-connection workflow. Existing legacy records cannot run, even if their stored status is active. Operators must create a new direct MCP connection and choose access rules; credentials and grants are not migrated. Remove each old record separately to delete its credentials. No schema migration or data deletion runs automatically. Direct MCP connections keep their existing grants and secrets. ## Model Used OpenAI GPT-6 via Codex, with reasoning, code execution, and browser tools. The exact runtime variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.921.0-canary.6 |
||
|
|
1483bb8bcf |
fix: pass plugin workers to issue tree resume (#13757)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue tree controls can resume tasks and wake their assigned agents. > - A sandbox-backed run needs the shared plugin worker manager to acquire its lease. > - The issue tree route constructed a heartbeat service without that manager. > - A ready sandbox provider therefore appeared offline and the resumed run failed setup. > - This pull request passes the existing manager to the route and distinguishes missing wiring from a stopped worker. ## Linked Issues or Issue Description Related implementation: #10262 identified this wiring defect in several dispatch paths. This PR applies the issue-tree resume fix to current master and adds coverage in the current route and runtime suites. The other dispatch paths remain in the scope of that earlier PR. **What happened?** Resuming a paused issue subtree can fail to acquire a sandbox lease even while its provider plugin is ready and its worker is running. The error says the worker is not running because this route's heartbeat service never received the process's worker manager. **Expected behavior** Resumed work uses the same worker manager as normal issue dispatch. Missing wiring and an actually stopped worker produce distinct diagnostic errors. **Steps to reproduce** 1. Configure an agent with a plugin-backed sandbox environment and start the provider worker. 2. Pause a subtree assigned to that agent. 3. Release the hold with `metadata.wakeAgents: true`. 4. The old route constructs an unwired heartbeat service and the resumed run fails during setup. **Paperclip version or commit** Reproduced by source inspection and regression coverage against `9d19f98b50`. **Deployment mode** Server with a plugin-backed sandbox environment. ## What Changed - Pass the process's plugin worker manager from `createApp` through issue tree controls to heartbeat dispatch. - Check for a missing manager before reporting the sandbox worker as stopped. - Exercise the resume request with a shared manager and cover both runtime failure cases. Model a stopped worker explicitly in the existing infrastructure-retry fixture. Existing board and company access checks remain covered. ## Verification - Targeted Vitest route/runtime/recovery suites: 17 passed; 384 native database tests skipped because embedded Postgres cannot start on this host. - Direct server `tsc --noEmit` passed. - `pnpm -r typecheck`, server typecheck wrapper, and `pnpm build` were attempted. Their runner dependency requires `cargo`, which is absent on this host. CI must pass these checks before merge. - Full `pnpm test:run` was attempted. The general server stage reported 8,146 passed, 4,747 skipped, and 14 failed tests across 35 failed suites. Failures were embedded Postgres startup errors (with related teardown errors) and 10 cache-directory rename failures on this macOS host. The local run stopped there; Linux CI must pass the complete suite before merge. - All CI gates are green, including the native database recovery suite, workspace typecheck/build, runner checks, and browser tests. Greptile reviewed the current commit at 5/5 with no unresolved threads. - No live provider operations were performed by these tests. ## Risks Low risk: dependency forwarding and error classification only. The route retains board authorization, company boundaries, hold semantics, and existing cancellation/replay guards. The change does not replace a sandbox or discard a retained lease. No schema or API shape changes. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tooling, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] This fixes existing behavior and does not add planned core features - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the bug issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no private ticket or instance data - [x] Targeted local tests pass; full-suite and toolchain limits are recorded above - [x] I have added or updated tests where applicable - [x] No documentation changes are needed for dependency forwarding - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.921.0-canary.5 |
||
|
|
e8c8ba3c19 |
feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its tool gateway applies company access rules and approval controls to connected apps. > - MCP aggregators expose many apps through one provider endpoint. > - Each aggregator needs its own credential, catalog, grants, and lifecycle in Paperclip. > - This pull request adds independent Zapier, Arcade, Composio Connect, and Executor setup with a common Access → Connect layout. > - A default-off MCP aggregators flag lets operators opt in while we complete provider acceptance tests. > - Agents use the normal Paperclip permissions, Test screen, and gateway after setup. ## Linked Issues or Issue Description **Subsystem affected** Apps, connection setup, shared contracts, and the remote MCP gateway. **Problem or motivation** Aggregator endpoints need clear provider setup and correct MCP sessions. Generic setup does not explain each provider's authentication or broad execution tools. Provider approval must preserve the original execution instead of replaying a write. **Proposed solution** Add four separate connectors behind Settings → Experimental → MCP aggregators. Start with human and agent access, then connect the endpoint and read its tools. Enable tools by default. Use the existing Permissions and Test screens after setup. Keep legacy Composio API-key and child connections intact. **Alternatives considered** A shared connection for all providers would mix credentials and access rules. Separate provider-specific permission and test screens would duplicate existing controls. Vercel Connect is outside this change. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps capability and the Connected Apps roadmap area. This work was requested and reviewed by the maintainer. Related work: #11894, #12630, #12632, #12634, and #12906 concern the legacy Composio broker. #13102 also covers remote MCP pagination. This change preserves the broker path and adds initialized sessions, response matching, and provider resume handling alongside pagination. ## What Changed - Add branded setup and interactive Storybooks for Zapier, Arcade, Composio Connect, and Executor. Use the existing access controls and normal action tests. Do not request a connection name or action choices during setup. - Add the default-off `enableMcpAggregators` flag to settings, managed feature metadata, the catalog, and setup guards. Hidden connections keep running. Legacy Composio connections remain unchanged. - Reuse the vault, grants, policy, and catalog models. Support OAuth discovery, bearer tokens, custom headers, and credential-bearing URLs. Add no database tables or migrations. - Initialize and retain Streamable HTTP sessions by connection and effective credentials. Read paginated catalogs and match streaming responses to request IDs. - Classify unfamiliar aggregator tools as writes despite upstream read-only hints; only exact reviewed read capabilities enter the read-only allowlist. Legacy Composio child behavior is preserved. - Preserve provider authorization links and execution IDs. Support Executor approve/resume, decline, and cancel without automatic replay of uncertain writes. - Preserve Off and Ask first choices during refresh and reconnect. Allow new tools and retire removed tools. Keep agent access updates atomic and preserve an empty agent selection. - Document connector UX rules, provider branding sources, and live acceptance results. - Stabilize the existing Sentry release fixture after its repeated CI failure by reusing one module mock; production Sentry behavior is unchanged. ## Verification - Final head `d11781970`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/35633534900) passed, including broad typecheck, test shards, build, and E2E. All 54 checks pass; 2 optional checks are skipped. Greptile is 5/5, Security Scan passes, and all review threads are resolved. - Passed 27 focused connector Vitest checks and 18 connector-only Storybook browser checks before the flag change. All 85 stories rendered at desktop and narrow widths. - Passed 5 connector lifecycle/server checks and 7 selected flag checks after adding the flag. The latter cover settings, managed defaults, cached catalog visibility, and all four setup routes. - Review fixes passed 13 risk/handoff/lifecycle checks, dedicated session-expiration and transport regressions, 13 selected connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A real Composio connection-list call also succeeded through the refreshed UI on `9ab115f71`. - UI and server TypeScript checks passed. UI build, Storybook build, token gates, and diff whitespace checks passed during implementation. - Real browser and real Paperclip agent tests passed for Arcade, Composio, and Executor. Tested action permissions, denied agent access, reconnect, disconnect, and isolation. Tested Arcade catalog additions/removal and Executor provider approve/resume, decline, and cancel. - Zapier live acceptance is incomplete. Its dedicated provider server is configured, but its credential-copy dialog returned an empty clipboard through browser automation. No live Zapier action is claimed. - The three isolated Sentry release cases pass after the CI fixture fix. - Local verification is deliberately narrow at the maintainer's request. The full local suite, recursive typecheck, and repository-wide build were not run. CI provides the broader checks. ## Risks - Shared MCP transport changes affect other remote MCP servers. Protocol fixtures cover initialized sessions, streaming response matching, pagination, and isolation. - Broad execution tools remain broad permissions. The provider governs actions inside those tools. - Provider handoff links are retained briefly in memory. After a server restart, a one-time link may require reopening the provider dashboard. Paperclip does not replay the original call. - Zapier remains unproven live. Custom-header imports and self-hosted endpoints have fixture coverage rather than a separate live account for every variant. - Turning the experimental flag off hides setup; it does not revoke existing credentials or stop existing connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, shell execution, and browser automation. The exact runtime model ID and context-window size are not exposed in this session. A separate Anthropic-backed Paperclip agent performed live gateway acceptance tasks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3790ca2f13 |
fix(runner): repair approval and Stop races and eval infrastructure (#13750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Runner tasks must continue after approval and stop when the user presses Stop. > - Live evals found races at approval delivery and provider startup. > - Browser readiness and CI setup errors also hid the actual task results. > - This pull request fixes those races and the related test infrastructure. > - Regression tests and saved live reports show which cases now pass. ## Linked Issues or Issue Description Companion eval definitions PR: https://github.com/paperclipai/paperclip-evals/pull/25 (AgentCore paused and provider/environment infrastructure). Related: #13741 now supplies the late-startup Stop fence and warm-attachment recovery; this PR retains that fence and extends startup tracking and regression coverage to both native backend paths. #13539 introduced queued approvals during active runs. #13738 fixes child assignment, task replies, and warm process continuity and is already in the base. #13291 concerns automatic continuation of interrupted legacy sandbox runs; this PR fixes native startup cancellation and does not change that recovery policy. **What happened?** An accepted service approval could wait after its source run stopped. Stop could return success before the provider handle existed. Work could then start after Stop, or a cancelled run could be recorded as failed. Some E2E tests also failed on unloaded browser content or irrelevant reply wording. Runner CI could fail before model work because of dependency or sandbox setup. **Expected behavior** Deliver each settled approval once after its source run stops. Do not start work after an acknowledged Stop. Preserve the audited cancellation. Test the intended product behavior with a ready browser and verified runtime dependencies. **Steps to reproduce** 1. Approve a service request while its source run is active. Let the run finish. Check that its result starts one continuation. 2. Delay provider startup. Press Stop before its handle is available. Check cancellation, then submit `/new`. 3. Run the browser, warm-workspace, and Stop-and-redirect cases from the linked report. **Paperclip version or commit** The branch includes master at `9d19f98b5`. The report records the original source for each focused attempt. **Deployment mode** Isolated local development instances and disposable Daytona sandboxes. ## What Changed - Deliver settled tool-action results for the exact company and source run during final cleanup. Keep the existing idempotent receipt and periodic recovery sweep. - Wait for startup to hand off its provider handle before acknowledging Stop. Reject first-turn admission after cancellation. Preserve a matching audited pending or acknowledged cancellation. - Wait for mounted task history and connector controls in browser tests. Record failure evidence. Grade workspace contents and process continuity separately from exact reply wording. Require each warm-turn marker once and in order, allowing surrounding prose. - Stop-and-redirect now checks that the source file exists and work is active before Stop. - Resolve target dependency locks in an uncredentialed CI job. Verify the lock artifact hash. Keep orchestration and publication on the trusted workflow revision. - Materialize the pinned OpenCode executable and configure the exact Codex executable's user-namespace profile before provider credentials are available. - Compress Daytona directory uploads with gzip. Preserve files, executable modes, symlinks, empty directories, and confinement checks. - Classify file-transfer RPC deadlines as infrastructure. Keep unrelated runner RPC failures visible. ## Verification - [Focused live report with screenshots and original attempts](https://pages.paperclip.ing/runner-reliability-20260921/): 14 of 15 selected Product E2E cases pass across the recorded revisions. Claude and Codex Stop → `/new`, Claude service approval, delegation, both hiring/reuse cases, and native Daytona warm continuity pass. - Two credentialed Runner smoke cases pass. These are not full protocol coverage. - E2E harness after the master merge: 429 tests pass. E2E and server TypeScript checks pass. - Daytona plugin: 239 tests pass, 6 skipped. Plugin TypeScript build passes. The compression test fails against the old code and passes with the change. - Runner backend/runtime regression group: 161 tests pass. Cancellation/startup selection: 26 tests pass. Approval delivery: 34 real-database tests pass. - Workflow security: 7 tests pass. Both edited workflows pass actionlint. Runner TypeScript and Rust builds pass. - After merging master, all 389 native executor tests pass, including both native backend paths and late startup after the Stop deadline. - Post-merge `pnpm -r typecheck` and `pnpm build` pass. The monolithic local `pnpm test:run` was interrupted to integrate master and is inconclusive. The [hosted CI test partitions](https://github.com/paperclipai/paperclip/actions/runs/35620461738) pass on `50a3e43822bcba1e0d07b1b45b0be91cbf9312da`. An unchanged sandbox callback schema test initially received HTTP 503. It passed five isolated local runs, its full local test file, and one failed-job CI retry. No assertion was weakened. ## Risks - Stop can wait for the bounded startup handoff. If it cannot settle, the existing pending-recovery state remains instead of a false acknowledgement. - Immediate approval delivery must remain idempotent across cleanup and recovery sweeps. Tests cover duplicate delivery and company/run boundaries. - The workflow changes still need hosted Linux verification. They retain the trusted workflow and credential boundaries. - Gzip reduces the observed provider upload from about 1.8 GB to 663 MB. It does not yet fix the remaining Claude Daytona transfer timeout. That recovery test never reached Claude, so recovery remains unverified. Use a matching image with the verified provider package preinstalled for the next recovery test; retain cold-upload coverage separately. - The report preserves diagnostic runs with missing source metadata and marks them as such. It does not claim a new full-suite pass. - This PR adds no new prompt policy or historical status reconciliation. ## Model Used OpenAI GPT-6 through Codex performed the primary implementation and review. The exact primary backend model ID is not exposed in this session. OpenAI `gpt-5.6-luna` assisted with bounded infrastructure work and verification. The agents used repository tools, code execution, and browser tests. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 <noreply@openai.com> |
||
|
|
0103cb2292 |
test: use neutral references in chat hiring evals (#13752)
## Thinking Path > - Paperclip helps people manage agents and their work. > - The agent chat hiring eval checks delegation, saved output, and worker reuse. > - It asks a worker to include a unique reference in each document. > - The open wording let the worker choose a credential-style label. > - The existing redactor removed that reference and failed the coordination check. > - This PR specifies a neutral reference line while keeping the same grading checks. ## Linked Issues or Issue Description Refs #13741. **What happened?** The Codex hiring eval saved a checklist with `Tracking token: [REDACTED]`. The runner treats the chosen label as credential syntax. The previous prompt only asked for the identifier and did not choose its label. This tests a redaction boundary unrelated to hiring and reuse. **Expected behavior** The coordination fixture requests ordinary business content with a neutral reference label. The grader still requires the exact identifier in the saved output. **Steps to reproduce** Run `agent-chat-hardening.runner-codex.local.hire-delegate-reuse` with the previous fixture. The retained failing attempt is in [campaign 35617045456](https://github.com/paperclipai/paperclip/actions/runs/35617045456). ## What Changed - Request `Reference: ...` in the checklist and review documents. - Increment the hardening suite definition to version 5 and record the reference format. - Document the fixture boundary. Production redaction and all grading checks stay unchanged. ## Verification - `pnpm test:e2e:runner:typecheck` passed. - `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files. - The catalog lists exactly the Codex and Claude local hiring cells for the selected case. - [Live campaign 35620731321](https://github.com/paperclipai/paperclip/actions/runs/35620731321) passed both selected local cells on `d8d7afe21`: native Codex (`gpt-5.6-sol`) and Claude (`claude-sonnet-5`). Both passed on attempt 1, including cleanup. [Published eval report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35620731321-1/). - Inspected saved evidence: each provider preserved both exact reference lines, used the same hired worker for checklist and review, and completed the final status query. Definition version 5 and hash `2c1c9f754c049c3e7cafd6c8d0a137e7ca06d546ce1e7b07ad5b7823f0f28db9` distinguish these results from the prior prompt. - [PR CI 35620741101](https://github.com/paperclipai/paperclip/actions/runs/35620741101): full build, type checks, and test partitions passed. One unrelated Telegram integration test initially failed because two random fixture company IDs produced the same seven-character issue prefix. The failed shard passed on one retry; no test code was changed. All latest-head merge checks are green. - Greptile reviewed `d8d7afe21` at 5/5 with no review threads. ## Risks The live models can still fail the coordination workflow. This change does not qualify or change credential-redaction policy. Earlier failed attempts remain part of the evidence; the new fixture has a distinct definition version. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.921.0-canary.4 |
||
|
|
9d19f98b50 |
fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent chat uses native runner sessions to plan, delegate, and track that work. > - A user can press Stop while the native session is still starting. > - The server can acknowledge that Stop without dispatching it, then let the session submit a turn. > - This leaves chat recovery waiting for an execution that the user expected to stop. > - This PR waits for the startup handle, dispatches cancellation, and prevents a late startup from submitting a turn. > - New full-stack evals check the resulting records and outputs across Claude and Codex. > - Those evals also exposed missing ACPX readiness fields, unbounded polling, and an old-run identity check that rejected valid warm handoffs. ## Linked Issues or Issue Description **What happened?** Stop during native startup could record an acknowledged cancellation with `dispatched: false`. The provider could then begin work. A subsequent `/new` stayed queued. A remote Claude follow-up also exhausted the command journal while probing warm-session readiness: ACPX never returned the readiness fields required by the shared transport. Once readiness worked, attachment incorrectly compared the next run descriptor against the old run ID. The 25 ms polling loop could issue 4,800 commands during its two-minute wait, beyond the 500-command bound. The existing chat eval treated lifecycle logs as proof of an active provider turn, so it did not distinguish startup cancellation from active-turn cancellation. **Expected behavior** A Stop during startup must reach the pending session. A late session must not submit a prompt after Stop. Recovery must retain control when startup exceeds the bounded wait. Chat evals must check saved task state, document contents, worker identity, account binding, and duplicate effects. **Steps to reproduce** 1. Start a native Claude or Codex chat turn. 2. Press Stop after process startup is requested but before the provider turn starts. 3. Send `/new`, then send a fresh message. 4. On the affected base, cancellation can be acknowledged without dispatch and the reset stays queued. **Paperclip version or commit** The live Claude baseline reproduced this on `29d6b3509`. The branch also includes master commit `0f5fafe16`. Related work: #13678, #13686, #13693, #13291, #13738. A separate runner reliability branch also contains a startup-wait fix. Its overlap must be reconciled before merging; this branch additionally prevents prompt submission after a late startup. ## What Changed - Wait for a pending native startup before acknowledging a run-scoped Stop. Preserve the existing recovery error when that wait expires. - Keep a Stop guard on startup. Cancel a late handle before it can submit a provider turn. - Add regression tests for normal handle publication and publication after the Stop deadline. - Back off blocked warm-attachment probes. Keep the fast two-snapshot barrier, fail closed, and record changed blockers. - Add red/green tests for delayed readiness, persistent blockers, alternating readiness, and readiness near the deadline. - Publish ACPX readiness and blockers. Preserve the old authority’s event acknowledgement barrier; only settled sessions can proceed to attachment. - Bind warm ACPX descriptors to the validated next authority while retaining old-run event correlation until activation. Preserve session identity and provider profile checks. - Exercise two consecutive run rotations through a qualified fake sidecar, verifying checkpointing, provider identity, pre-activation rejection, and new-run work admission. - Separate startup and active-turn cancellation checkpoints in the browser eval. - Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells across Claude and Codex. - Cover hiring and reuse through managed AI accounts, source-based review, current blocked-task status, request replay after a lost HTTP acknowledgement, server restart continuity, and Stop/reset continuity. - Use ordinary production agent instructions. Enable API tools only for the two coordination cases that need them. - Calibrate the matchers with invalid records and outputs. Require remembered context after restart and a structured status snapshot that distinguishes the current blocker from history and task status from active execution. Compare the public issue mutation contract and relationships during read-only reporting. Preserve before/after source records in failed eval evidence. - Fix the lost-ack browser harness and verify it against a real HTTP server. Check the chat composer after restart instead of waiting for an unrelated document lifecycle event. - Document the scope and limits of each case. ## Verification - The startup regression failed on the unfixed executor and passed after the fix. - `pnpm test:e2e:runner:typecheck` passed. - `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files. - `pnpm exec vitest run server/src/services/native-runtime/native-session-executor.test.ts` passed: 385 tests. - [Baseline live campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868): Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed. Codex delegation was blocked by provider capacity. - [Eval-only startup campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786): both providers failed as expected. Both persisted `dispatched: false` and left `/new` queued. - [First fixed campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706) on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both providers. Failed cases exposed eval harness defects and remote continuity failures. All attempts remain available. - [Original workflows and stronger memory checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649) on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and startup Stop passed for both providers; Codex remote restart passed. Claude remote restart exposed the missing readiness contract. Two Codex planning cells hit provider capacity. - [Unchanged-model retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548): Codex planning and backlog creation both passed. - [18-cell campaign with ACPX readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963) on `6a98ef743`: 16/18 passed, including all local/remote Stop and committed-send cases. Claude remote continuity exposed the next-authority check, now fixed. Codex hiring produced its checklist, but the runner redacted the requested marker after it appeared as “Tracking token: …”. That content-redaction policy is unchanged and remains an explicit limitation. - [Structured status grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725) on `50448c228`: both providers passed on their first attempt, including cleanup. - [Complete read-only state grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011) on `551e13892`: both providers passed. - [Final ACPX handoff and hiring retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456) on `cbd637587`: all three Claude Daytona cases passed (restart continuity, active Stop/reset, and lost-ack replay). Codex hiring reproduced the content-redaction failure: the saved checklist contained `Tracking token: [REDACTED]` instead of the required business marker. All four cases completed cleanup successfully. [Published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/). The only subsequent commit adds the qualified-sidecar integration test; production code is identical to this live proof. - `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without paid models. - Runner TypeScript typecheck passed. All 5 warm-readiness tests pass; two failed with the prior fixed-rate loop, and the late-readiness test failed before the pacing correction. - ACPX readiness and warm-identity regressions each failed before their fixes. All 292 runner-core Rust library tests passed. The qualified-sidecar integration test passes. Rust formatting is checked. - Status-grader regressions for misleading historical mentions and previously unchecked mutations each failed before tightening the oracle and pass now. - [Latest-head CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307) passed on `a4093c8f1`: full build, type checks, test partitions, browser E2E, and native runner checks. Two unrelated tests initially failed (Sentry fixture release attribution and local-service fixture readiness); both passed locally together (35 passed, 5 optional SDK tests skipped) and on the failed-job retry. No changes were made to those tests. - Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed and all review threads are resolved. - The paid live suite is not fully green: the reproducible content-redaction case remains red. This is separate from the passing PR merge checks. No production content-redaction, prompt, model, or completion-policy change is included. - Managed-account hiring and review cases explicitly enable API tools; these do not qualify default new-user onboarding. ## Risks - Stop can wait up to 30 seconds for startup, then use the existing pending-recovery path. This does not prove that remote cleanup has finished. - Blocked warm readiness adds up to 750 ms between later probes with the two-minute remote budget, or about 32 ms with the default five-second budget. Ready sessions retain the short second barrier. - Paid evals can fail because of provider capacity or agent decisions. Each failure needs evidence-based classification. - The HTTP request replay case checks comment idempotency and duplicate effects. It does not prove replay safety for an ambiguous provider tool call. - The new suite is opt-in. It does not increase the default paid campaign. - No production prompts or model selection change. Review-handoff behavior and content-redaction policy remain separate product decisions. The latter can remove harmless business content that looks like credential syntax; the failing attempt is retained. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.921.0-canary.3 |
||
|
|
e9bc2efdaa |
docs: turn DeepWiki into a contributor handbook (#13745)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Contributors need to understand its control plane and find the code that owns each behavior. > - The public DeepWiki has many pages, but it lacks a guided path through task execution and system boundaries. > - A refresh alone cannot define that reading path or the evidence each page must show. > - This pull request adds a 30-page contributor handbook configuration with source anchors and clear page boundaries. > - The handbook helps contributors trace work, investigate failures, and choose the right code and tests. ## Linked Issues or Issue Description **Issue type** Unclear or confusing documentation; outdated generated documentation. **Where is the issue?** [Paperclip DeepWiki](https://deepwiki.com/paperclipai/paperclip). The repository did not have a `.devin/wiki.json` configuration. **What's wrong?** The wiki needs a contributor reading path. Readers need to connect product terms to implementation, follow a task through both execution paths, and find the state and tests that explain failures. Index freshness also depends on regeneration. **Suggested fix** Define six chapters and 30 pages in `.devin/wiki.json`. Give each page source anchors, required questions, and coverage boundaries. Use shared notes for evidence standards, diagrams, terminology, and freshness rules. Keep existing generation effort settings. Searched public issues and PRs for DeepWiki, wiki, and contributor handbook work. No duplicate or related DeepWiki change was found. Checked `ROADMAP.md`; this is a documentation configuration change. ## What Changed - Add only `.devin/wiki.json`, using the [documented DeepWiki configuration](https://docs.devin.ai/work-with-devin/deepwiki#steering-deepwiki). - Define 30 pages under Understand Paperclip, Contribute to Paperclip, Orchestrate Work, Run Agents, Extend and Integrate, and Govern and Diagnose. - Add 68 generation notes with 244 verified repository source paths. Require source and test citations, useful parent pages, and explicit ownership boundaries. - Require a task tour with separate direct-adapter and experimental Runner paths, run and task state diagrams, and a connection authorization flow. - Separate run logs, operator-configured observability, and first-party telemetry. Require defaults and feature gates to be resolved from the indexed code. ## Verification - Passed JSON and supported-field validation. Checked unique titles, valid parents, an acyclic hierarchy, the exact page tree, note limits, and all 244 source paths. - Reviewed coverage against the six reader questions in the plan. Confirmed each page has source anchors, test anchors, and coverage boundaries. - Passed `git diff --check`. - Passed `pnpm -r typecheck` and `pnpm build` on the implementation base, `c65fc9e3c81c41aafe421aa90a00514b84343285`. - Ran `pnpm test:run`. The broad run hit a 15-second timeout in `server/src/services/native-runtime/remote-deliverable-file.test.ts`, in “reads verified remote bytes without touching controller paths”. Stopped the broad run after the failure. Full-suite verification remains incomplete. - Reran that test file in isolation: all 30 tests passed in 6.45 seconds. - Rebased onto current `master`, `57fd8b70d`. Repeated configuration, source-path, and diff validation after the rebase. No application tests were added for this configuration-only change. - After merge, request DeepWiki regeneration. Confirm its indexed commit contains the configuration. Then check the generated task tour and one page from each chapter against implementation and test citations. Regeneration and generated-page review are pending. ## Risks - Generated text can still contain errors. The instructions require citations and honest treatment of conflicting evidence, but the generated pages need review. - The new page structure can change navigation and page links. - Source paths and implementation can change between generations. Notes require replacement anchors, current defaults, and clear labels for experimental behavior. - The public index stays stale until someone requests regeneration. This PR adds no refresh service and changes no application behavior. ## Model Used OpenAI Codex, based on GPT-6. The exact runtime model identifier, context window size, and reasoning setting are not exposed in this session. Used repository inspection, web research, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.921.0-canary.2 |
||
|
|
57fd8b70d2 |
feat: add agent avatar download to Slack setup and settings (#13740)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Slack connections let a team talk to those agents in Slack. > - Agents now have a saved avatar, but Slack setup did not offer that image. > - A matching avatar helps a team recognize its agent. > - This pull request adds an optional avatar step and a download in connector Settings. > - Users download a PNG and upload it directly in Slack with clear instructions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Slack connector onboarding and its Settings page. **Current behavior** Setup does not offer the assigned agent's avatar or explain how to upload it in Slack. **Proposed behavior** After Slack connection verification, users can download a 512 × 512 PNG of their agent's saved avatar. They can upload it in Slack, confirm, or skip. Settings keeps the download and upload instructions available after onboarding. **Reason and benefit** The same avatar helps people recognize the agent across Paperclip and Slack. Users who skip the optional step can return to it in Settings. **Breaking changes** None. No schema, authentication, Slack scope, or provider API change. Completed connections keep their existing completion state. Searched existing Slack avatar and Cliptoon PRs; no matching implementation was found. ## What Changed - Add an optional avatar step before personal Slack account linking. Keep the numbered sidebar and shared footer. - Resolve the selected agent's saved appearance for the preview and PNG download. - Add the same download and expandable upload instructions to connector Settings. - Remember uploaded or skipped per company and endpoint in browser storage. Treat uploaded as user confirmation, not provider verification. - Reject failed or non-PNG download responses and allow retry. - Reuse the production avatar components in onboarding and Settings stories. - Test wizard progression, resume, Settings, download recovery, storage isolation, and terminated assigned agents. - Exercise real PNG downloads in the Slack browser flow and keep default app names consistent with app creation. - Fetch the assigned agent directly so its saved avatar remains available after termination. ## Verification - Focused chat suites: 48 passed; the two affected suites passed again after the final naming fix (32 tests). - Slack browser E2E passed through setup, avatar download, account linking, and Settings download. PNG signature and 512 × 512 dimensions verified. - UI token gates passed. - Browser: downloaded the real 512 × 512 PNG; checked confirmation, return, mobile layout, and Settings instructions. - Full workspace typecheck, application build, and production Storybook build passed. - All latest-head CI checks passed (54 passed, 2 skipped), including all browser, chat, general, and serialized test groups. The unrelated Sentry test failed once and passed on the single CI rerun; its suite also passed locally. - Local full-suite attempt encountered a rapid Slack callback ordering failure under concurrent build load; that test passed in isolation, and all three chat shards passed in CI. The remaining local run was not used as the merge gate. - Review the Connections / Slack / Add avatar and Avatar in Settings stories. ## Risks - Slack upload is manual. Confirmation does not claim to verify the Slack icon. - Optional step progress is browser-local. Clearing storage or changing browsers can show it again. Setup still works when storage is unavailable. - The existing avatar API remains the image source. Download failures show a retry message. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fd071748ee |
fix: stop repeated notifications for finalized run failures (#13739)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native run reconciliation repairs saved execution outcomes after
interruptions.
> - The sweep also visits runs whose final results are already
committed.
> - An unchanged failed run still received a new status delivery ID on
every sweep.
> - The browser treated each delivery as a new failure after its short
duplicate window expired.
> - This pull request makes unchanged projections a no-op and suppresses
repeated or historical run toasts.
> - Operators receive fresh failure alerts without repeated alerts for
old work.
## Linked Issues or Issue Description
**What happened?**
An old failed run repeatedly produced failure toasts while the browser
remained open. The task could already be cancelled. Reconciliation
rewrote the same failed outcome and queued another status broadcast.
**Expected behavior**
An unchanged committed run must not queue a new status notification.
Repeated deliveries must still refresh cached state without another
toast.
**Steps to reproduce**
1. Finalize a native run with a failed result and successful workspace
finalization.
2. Deliver its pending execution status and cancel its task.
3. Replay finalization and status delivery on each periodic sweep.
4. Observe another failure broadcast for every sweep before this fix.
**Paperclip version or commit**
Reproduced against source commit
canary/v2026.921.0-canary.1
|
||
|
|
0f5fafe16b |
fix(runner): preserve task replies and warm process continuity (#13738)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - The Runner connects provider sessions to task state, replies, and delegated work. > - Full-stack tests found lost final replies, rejected helper calls that stopped the parent, and unnecessary process restarts. > - A completed child could also receive a new assignment wake that the scheduler then cancelled. > - This change fixes those boundaries and gives agents clearer teammate instructions. > - The tests retain strict completion and process-continuity requirements. ## Linked Issues or Issue Description **What happened?** A generated attachment comment could suppress an agent's final reply. A known Codex helper could stop its parent when it requested a Paperclip tool. Native Daytona processes restarted between turns because Paperclip minted an unused GitHub broker token. Reassigning a completed child queued a run that immediately cancelled. Revision instructions also allowed agents to do work assigned to a named teammate themselves. **Expected behavior** Keep the final reply. Reject helper tool requests without borrowing parent authority or stopping the parent. Keep an unconfigured sandbox process alive between turns. Treat assignment-only changes to completed tasks as metadata changes. Preserve explicit teammate assignments during revisions. **Steps to reproduce** Run the retained Runner E2E cases for file handoff, teammate reuse, Daytona warm continuity, and Legacy Claude interview/plan acceptance. The focused regression tests reproduce the reply, helper, process-lifetime, and assignment-wake defects without provider calls. **Paperclip version or commit** The live lifetime and completion campaign used `db3857807`. This PR replays the changes on master `c65fc9e3c`. See Verification for the limits of that evidence. **Deployment mode** Isolated local instances and native Runner sessions in Daytona sandboxes. Related work: #13546 handles a different queued-run issue after an issue-lock compare-and-set failure. This PR prevents the unnecessary assignment wake earlier. #13410 covers retained user services; this PR covers the provider process. No duplicate fix was found. ## What Changed - Exclude generated deliverable-binding comments from final-reply deduplication. Preserve the attachment and explicit user-facing replies. - Reject Paperclip tool and input requests from known Codex helper threads without terminating the parent. Keep unknown-thread rejection intact. - Explain how to hire or reuse a persistent teammate and preserve named delegation on revisions. Update generated protocol fixtures. - Use stable, token-free GitHub wrappers for unconfigured native sandboxes. Preserve credential isolation, configured-account rotation, and cleanup after partial staging failures. - Do not queue assignment-only wakes for done or cancelled tasks. Keep explicit reopening behavior. - Make warm-continuity fixtures create real review cards. Read the persisted final response selected by production presentation logic. Missing selected evidence still fails. ## Verification - Before rebase: 560 focused route, native-executor, and launcher tests passed. The new regressions were reproduced before their fixes. - Live E2E: Legacy Claude interview/plan acceptance passed 3/3 repetitions. Daytona warm continuity passed 2/3 full repetitions. Each successful run retained one process and provider session for all three turns. - The remaining Daytona repetition stopped after a same-URL browser reload left the page blank. Both completed turns retained the same process. Its failed verdict remains unchanged; this PR does not claim the blank-page cause is fixed. - Reports: https://pages.paperclip.ing/runner-e2e-lifetime-race-20260920/investigation.html and https://pages.paperclip.ing/runner-e2e-behavior-followups-20260919-results/investigation.html - Post-rebase `pnpm build` and `pnpm -r typecheck` passed. All 414 Runner E2E harness unit tests and its typecheck passed. Codex protocol tests: 88 passed, 2 ignored. - Latest-head CI: 55 successful checks and 2 intentional skips. Greptile: 5/5 with no review threads. The unchanged workspace exposure tests hit a fixed-port collision on the first CI attempt; their local suite passed (25 tests, 3 platform skips), and the CI shard passed on one retry. - The duplicate local `pnpm test:run` was stopped after the full hosted general and serialized test shards passed. It did not finish locally and is not counted as a local full-suite pass. ## Risks Configured GitHub accounts retain run-scoped credential rotation and can still restart warm processes. That limitation requires a separate design. Known provider helpers cannot use Paperclip coordination tools directly; they must return findings to the parent. The delegation prompt is an instruction, not an enforced guarantee; Codex Mini hiring/reuse failures remain open. No schema or workflow changes are included. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and parallel coding agents. The exact deployment suffix and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub issues and shared reports) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; full hosted test shards also pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
29d6b35096 |
chore(lockfile): refresh pnpm-lock.yaml (#13645)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.921.0-canary.0 |
||
|
|
c65fc9e3c8 |
fix: recover authentication and browser connection failures (#13724)
fix: recover authentication and browser connection failures Include connect timeouts in the bounded retry policy for idempotent actor synchronization. Handle WebSocket constructor failures through existing reconnect paths and preserve HTTP polling while realtime is unavailable. Refresh visible company queries until the socket recovers and clear all fallback timers on hiding or unmount. Verify 172 focused tests, server/UI typechecks, UI build, and design token gates. Full workspace build/typecheck require the unavailable Rust toolchain; the full test run is tracked separately. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9f30eb10dd |
fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path > - Paperclip manages AI agents and keeps their work attached to tasks. > - Chat connectors carry user messages and agent replies between a provider and those tasks. > - Each extra startup and context reset delays a reply. > - Managed account metadata was lost during adapter decoding, so compatible follow-ups started fresh. > - This branch fixes the reset and measures the remaining preparation, execution, and delivery costs. > - The changes must preserve account isolation, authorization, durable output, and recovery ownership. ## Linked Issues or Issue Description Refs #13699. The related service lifecycle work in #13410 and #13408 is separate; this branch focuses on task-bound chat response latency. **What happened?** Managed AI follow-ups started new provider sessions even after their configuration fingerprint stayed stable. The Codex codec removes unknown fields. The resume check then read the removed credential identity and treated it as a credential change. **Expected behavior** Compatible follow-ups resume the correct provider session. Changes to credentials, responsible users, permissions, or task configuration retain their reset behavior. **Steps to reproduce** 1. Use a Slack connector with a managed AI connection. 2. Send a message, then send a same-thread follow-up. 3. Inspect the configuration reset reason and the provider session identity. **Paperclip version or commit** Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`. **Deployment mode** Cloud staging with a native Codex runner. ## What Changed - Read saved credential identity before adapter decoding discards it. - Remove the internal credential identity from adapter-facing session params. - Test the real Codex codec and missing, changed, or unmanaged identity cases. - Preserve configured warm Codex runners and flush refreshed credentials after every turn. - Fence detached or closing session handles from successor credential ownership. - Stage current Codex launch credentials after restoring durable session history, uploading launch assets only once. - Reuse a runner binary already in the retained sandbox only when its SHA-256 matches the controller-owned artifact; still verify required capabilities before launch. - Lock the task before the run when saving results, preventing deadlocks with task updates. - Scope reusable projectless sandboxes to the company, environment, task, agent, and runtime configuration; verify Daytona sentinels for that scope. - Admit a new authorized chat message after a fully committed failed run and verified process cleanup. - Send compact deltas for verified plain-text Slack continuations. Match the actual prior run and current comment identity/body; exclude edited historical comments and prior agent output, preserve genuine brief edits and the full bootstrap fallback. - Keep attachments, omitted input, questions, approvals, recovery, and other providers on their existing framing. - Document managed session compatibility, credential lifecycle, and compact continuation boundaries. ## Verification - Workspace/session coverage: 156 tests passed. - Native session and credential ownership coverage: 390 tests passed, including exact artifact reuse, mismatches, failed probes, timeouts, and explicit artifact overrides. - Explicit continuation and durable chat authorization coverage: 172 tests passed. - Session resume and launch preparation coverage: 416 tests passed. - Result persistence coverage: 15 tests passed. The new concurrency test reproduced a PostgreSQL deadlock before the lock-order fix. - Environment lifecycle coverage: 92 tests passed, including projectless reuse and task/agent isolation at both selection and atomic handoff. - Daytona plugin coverage: 237 tests passed; 6 gated tests skipped. Standalone plugin build passed. - Compact Slack continuation and native resume coverage: 69 tests passed, including full-bootstrap retention, matching message authors/bodies, current-delivery selection, rejection of duplicate identities and historical comments, brief edits, and attachment/recovery fallbacks. - Final frozen-head `pnpm test:run` on repository-supported Node 26: 668 suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped. The sole failure was a local `socket hang up` in `issue-recovery-actions.test.ts`, not an authorization assertion mismatch. All 57 tests in that suite passed three fresh reruns, and the suite passed latest-head CI. The full local invocation is therefore not claimed green. - An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup failure; that 11-test catalog suite passes on Node 26 and in CI. No test behavior or timeout was relaxed. - Full local typecheck and build passed. Latest-head CI is green; Greptile is 5/5 with no unresolved review threads. - Two real Slack baseline replies took 25.1 and 24.6 seconds (24.9-second mean). Three same-thread signed probes on this head took 23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample and a modest wall-clock improvement, not a large or statistically established speedup. - In that same thread, uncached provider input fell from 8,514 tokens before compact input to 694–765 tokens afterward. The current delivery uses a 362-character delta; the full 19–21k-character bootstrap remains available for failed resume. Verified runner artifact preparation fell from about 1.2 seconds to 0.6 seconds. - A fresh thread created a separate task, sandbox, and provider session with full bootstrap (24.4 seconds). Its follow-up reused its own sandbox/session and compact input (28.3 seconds, including 16 seconds of model execution). Model variability and process startup remain substantial. - A signed duplicate webhook produced exactly one user comment, one successful run, and one final Slack reply. Slack's API independently confirmed the actual replies and a public task URL without an internal or pool hostname. - Earlier signed probes verified recovery after a failed run and reuse across a server deployment. The final idle test observed Daytona report the sandbox as stopped, then delivered a new reply in 19.9 seconds using the same sandbox/provider-session identity and compact input. Slack’s API confirmed that reply. - Live probes use signed synthetic inbound webhooks and real outbound Slack delivery, read back through Slack’s API. The final browser recheck found the Mac locked and the Slack tab blocked by another extension, so this is not claimed as full UI E2E proof. - This is a review branch. Do not merge until the maintainer reviews it. ## Risks - Incorrect session reuse could mix account or task context. Missing or changed identities continue to reset, and existing authorization checks remain in place. - Warm mode remains opt-in. Remote warm mode requires a reusable sandbox lease. Retained processes keep credentials until they close, so idle expiry and ownership fences are required. - A fresh user message may continue after a committed provider failure. Approval, current authorization, process termination, and prior-result checks remain required. - Projectless sandbox reuse is task- and agent-scoped. Missing or mismatched ownership cannot replace an existing lease; existing workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox per task/agent, so provider auto-stop and deletion policies still determine idle compute and storage costs. Fleet defaults are unchanged. - Compact prompts apply only after proven resume and a matching prior-run delta. Missing or specialized context falls back to full input; fresh sessions always receive the full bootstrap. - No schema or migration changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, tool use, and test execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>nightly/v2026.921.0-nightly.0 canary/v2026.920.0-canary.2 |
||
|
|
600e552d7b |
fix: attribute Sentry errors to the loaded source release (#13719)
Attribute optional server and browser Sentry events to their source build. Use validated build commits for Docker and source/npm artifacts, preserve explicit server release overrides, and keep cached browser bundles tied to the commit they loaded. Verify 127 focused tests, server/UI typechecks, Docker and source build stamps, all 53 CI checks, and Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.920.0-canary.1 |
||
|
|
2a99de80ec |
fix: supply public task links and preserve managed AI sessions (#13699)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors deliver agent replies to external conversations.
> - Agents need a public task link when a user asks to open the task.
> - The prompt and task tools lacked that link, so an agent could invent
an internal address.
> - Managed AI credential directories also changed the session
fingerprint on each run.
> - This change supplies public task URLs and excludes only those
temporary directory values from the fingerprint.
> - Follow-up messages can reuse compatible sessions while real
configuration changes still reset them.
## Linked Issues or Issue Description
Refs #13680 and #13694 for the related Cloud-origin fixes. No duplicate
open PR was found.
**What happened?**
An external chat reply could contain an invented internal task URL. The
publication filter then removed the link. Follow-up runs also lost their
saved provider session because each managed credential home used a
different temporary path.
**Expected behavior**
Agents receive the current public board URL for a task. Temporary
credential directories do not reset an otherwise compatible session.
Account, credential, model, permission, and custom environment changes
still invalidate it.
**Steps to reproduce**
1. Use a chat connector with a managed AI connection.
2. Ask for the current task link.
3. Send a follow-up message with the same agent configuration.
4. Inspect the task URL and the session reset reason.
**Paperclip version or commit**
Reproduced on the source at
nightly/v2026.920.0-nightly.0
canary/v2026.920.0-canary.0
|
||
|
|
b193077582 |
fix: report Slack callback health correctly behind Cloud proxies (#13694)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack connections turn messages into governed agent runs and return replies to Slack. > - Remote runs must replace an incompatible sandbox runner with the controller's packaged binary. > - The fallback used a package-relative path that does not match the vendored server layout. > - Slack callback health also compared the internal proxy address with the public callback address. > - Master now contains the runner fallback fix; this pull request fixes callback health and extends missing-runner regression coverage. ## Linked Issues or Issue Description Refs #13677, #13680, #13691, and #13686. **What happened?** A packaged controller stopped a Slack-triggered remote run with `runner_remote_artifact_unavailable` when the sandbox runner needed replacement. Working Slack callbacks also showed a stale URL warning behind the Cloud gateway. **Expected behavior** The controller stages its packaged runner when needed. Callback health uses the observed public address and still detects real address changes. **Steps to reproduce** 1. Run a packaged server with a sandbox that has an older runner or no runner. 2. Send a Slack mention to an agent that uses that sandbox. 3. Route signed Slack callbacks through a claimed Cloud gateway that rewrites the upstream host. 4. Check the run and the Slack callback health panel. **Paperclip version or commit** Reproduced on master at `aeef493f4a7603b7b1254421b80fb00212982390`. The fix branch also includes #13691. **Deployment mode** Packaged server with a Cloud gateway and a Daytona sandbox. ## What Changed - Extend the controller-owned runner fallback tests with a missing sandbox binary case; retain the fix now merged in #13686. - Prefer dedicated gateway diagnostic headers that survive provider rewrites of standard forwarded headers. Use validated host hints only for callback-health evidence on claimed Cloud instances after provider acceptance. Preserve request bodies, routing, authentication, and configured callback URLs. - Cover current, stale, and missing sandbox runners, all Slack callback surfaces, rejected callbacks, malformed proxy hints, real host and port changes, and the existing self-hosted behavior. - Document the artifact lookup and callback-health boundaries. ## Verification - Native session executor and binary resolver suites: 379 tests passed after merging current master (`45c99a0d0`). - Targeted callback integration suite with disposable PostgreSQL: 4 tests passed before rebase. - Full `pnpm -r typecheck` and `pnpm build` passed after merging current master. The full local suite passed 12,668 tests; one suite failed to start its disposable PostgreSQL. Rerunning that suite alone passed all 31 tests. - The built server resolver selected the executable under `server/dist/vendor/paperclip-runner/bin/`. - Final callback regression: all 4 targeted integration tests pass, covering provider header rewrites, default ports, uppercase/trailing-dot hosts, and ignored self-hosted hints. - Live staging proof of the runner fix: a previously failed Slack thread recovered, a new mention received its requested response, and the account-connect command succeeded. Unsigned callbacks returned 401. - Live browser and Slack acceptance passed: generated callback URLs, account linking, a new mention, an interactive question and answer, and all three callback-health indicators. The connector was activated through the onboarding UI. - All 54 latest-head checks pass, two optional checks are skipped, and Greptile is 5/5 with no unresolved threads. ## Risks - Runner fallback must select a binary for the remote platform. This preserves explicit remote artifact overrides and the existing capability checks. - Proxy headers are not identity proof. They are used only for diagnostics on claimed Cloud instances after the provider accepts the request. Self-hosted instances ignore them. Wrong public hosts and ports still warn. - No database migration or new public API contract. ## Model Used OpenAI GPT-6 via Codex. Used reasoning, repository tools, code execution, and browser/native-app testing. Exact model variant and context window were not exposed by the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
45c99a0d06 |
fix(adapters): default legacy harnesses and connected tools to full auto (#13693)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Legacy adapters launch provider CLIs and expose connected tools. > - Existing defaults did not consistently grant full automatic permission. > - Remote Claude used a fixed tool list that omitted MCP tools and future tools. > - Direct Codex launches and OpenCode configuration also used narrower defaults. > - This change gives all these paths the same full-auto default as native runners. > - Explicit restrictive settings continue to work. ## Linked Issues or Issue Description Refs #13686. This PR is stacked on that native-runner and task-reassignment PR. Merge #13686 first. Related: #831 (constructed Claude agents), #1935 (adapter-switching permission defaults). ## What Changed - Use actual Claude permission bypass for local and remote runs and probes. Remove the fixed tool list so MCP and future provider tools are included. - Identify actual managed sandbox targets to Claude with `IS_SANDBOX=1`. Do not mark ordinary host execution as a sandbox. - Default direct Codex execution to approval and sandbox bypass, matching agent creation. Preserve explicit false, CLI profiles, sandbox modes, approval policy, and network restrictions. - Set OpenCode's full-auto runtime permission to `allow` for every tool and connection. Preserve the existing explicit opt-out. - Default Gemini probes to the same YOLO mode as execution. Make the legacy ACP `default` alias use `approve-all` for fresh and resumed sessions. - Add default, opt-out, remote, probe, connected-tool, and resume regression tests. Update adapter configuration documentation. - Other adapter paths already request full automatic permission or have no provider approval gate. ## Verification - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server and legacy ACP tests passed, including defaults, explicit opt-outs, remote launches, connected tools, and fresh/resumed sessions. - Greptile reviewed current head `8ca135eaffcf9cfdba6f1368e896a781a0891d50` at **5/5**. The security reviewer acknowledged the documented full-auto requirement. Acknowledged discussions are resolved. - Current head has **54 passing checks**. [PR checks](https://github.com/paperclipai/paperclip/pull/13693/checks). The process-adapter signoff browser shard passed on one retry after its first attempt exceeded a three-second issue-run wait. - **Six native Claude/Codex real-provider cases passed on their first attempt, with cleanup passing**, against the combined branch: plans, reassignment, and backlog creation/status. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). This does not claim a real-provider run of every legacy adapter. - The live-tested revision is `a37881c824dcd7170380fc4b788732fc743e5da7`. The current head differs only in the corrected heartbeat test expectation; application code is identical. - The campaign result-enforcement job passed. The separate report publisher failed during frozen dependency installation because the trusted workflow's patched-dependency configuration does not match its lockfile. Passing case evidence remains downloadable from the workflow. - Full-suite coverage comes from CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. ## Risks - Missing permission settings now grant all provider operations, including connected tools. OpenCode full-auto also overrides ambient provider permission rules. An explicit Paperclip permission opt-out preserves restrictive behavior. - Claude refuses full bypass as root outside an identified sandbox. Ordinary host deployments must run Claude as a non-root user. Managed sandbox launches include the required marker. - These defaults do not grant additional Paperclip roles, connections, or company access. Existing controller authorization and governance still apply. - This PR depends on #13686. Retarget it to master after that PR merges. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
04546c82d5 |
fix(runner): reconnect Daytona sessions after controller restart (#13691)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner can execute a task inside a Daytona sandbox. > - The sandbox can keep running when the Paperclip controller restarts. > - Recovery treated sandbox process IDs as local process IDs and selected the wrong recovery path. > - Live verification also found races between startup, shutdown, and queued task cleanup. > - This pull request verifies the existing remote owner and orders those transitions. > - Users can continue the same task and provider session after a controller restart. ## Linked Issues or Issue Description **What happened?** The Daytona `recover-controller` cases failed with `runner_state_identity_mismatch`. Remote process IDs can be absent on the controller or collide with unrelated local processes. Recovery then looked for remote state in the local runner directory. Later turns could also start before the previous executor released its sandbox resources. **Expected behavior** Reconnect to the original sandbox and authenticated runner. Preserve the task, provider session, and queued comments. Reject a replacement sandbox or mismatched identity. Do not start another provider during reattachment. **Steps to reproduce** Run the `everyday-workflows` `recover-controller` case for `runner-codex` or `runner-acpx-claude` in Daytona. The browser creates a Python tool, requests a revision, restarts the controller during execution, and queues another revision. It then downloads and tests the final ZIP. Related: #13682 is the preceding operational fix. #13291 addresses legacy sandbox conversation recovery, a different execution path. #13666 includes broader run-capacity work; this change guards cleanup of an existing native task executor. ## What Changed - Add remote runner recovery without interpreting sandbox PIDs on the controller. - Verify the original provider lease, remote workspace, durable state, process marker, and authenticated PRP authority before adoption. - Compare the process marker with live Linux boot identity and start ticks to reject PID reuse. Read virtual proc files through the guaranteed Node runtime; unavailable proof blocks adoption without blocking a fresh launch. - Make the E2E supervisor own the actual server process so forced restart cannot leave a late database closer behind. - Scope the chat delivery lease test to its own fixture instead of draining other tests’ pending deliveries. - Preserve provider-attempt counts and recorded evidence during reattachment. - Serialize an idle-session checkpoint with admission of the next native turn. - Wait for an in-progress startup to acknowledge restart detachment. Fail after a bounded deadline if it cannot. - Keep a queued comment waiting until the previous native task executor releases its resources. Allow unrelated tasks to continue. - Update the Daytona image's resolved lock digest to match current dependency manifests. - Add classifier, ownership, process, startup, checkpoint, and queued-admission regression tests. Document recovery behavior. ## Verification - 415 focused tests passed across native execution, restart recovery, workspace synchronization, queued admission, and real-process restart tests. The final Node-based fingerprint change passed all 375 native-session tests. - Runner harness unit tests: 394 passed. Chat integration shard 2: 335 passed after fixture isolation. - The exact fingerprint command succeeded twice in a disposable Daytona sandbox and returned the same identity; the sandbox was deleted. - 11 real-process restart integration tests passed, including absent and colliding remote PIDs. - Repository typecheck and final build passed. Broad local checks found machine-dependent database startup and timing failures; focused retries passed. The final-revision PR pipeline is green. One unrelated browser shard hit a five-second blank-page timeout on the first run and passed its targeted retry. - Final-revision local headed browser E2E: `everyday-workflows.runner-acpx-claude.daytona.recover-controller` passed on attempt 1 in 4.7 minutes, **40/40 checks**. Manual browser inspection confirmed Done, all three ZIPs, and delivery of the queued follow-up. All three runs succeeded using the same provider session. The harness downloaded and independently tested the final artifact. - Final-revision Daytona campaign: https://github.com/paperclipai/paperclip/actions/runs/35463999611 — **Codex passed first attempt (4.8 minutes); ACPX Claude passed first attempt (6.1 minutes)**. Campaign aggregation/publication is finishing; both test jobs succeeded. - Greptile reviewed `beb08d8493b3286f5bb988dead369ff8c96a395d`: **5/5**, no open findings. - Staging browser verification is pending selection of a disposable staging instance and removal of a Chrome extension UI block. ## Risks - Recovery now depends on the original sandbox remaining available. A replacement or mismatched identity still blocks adoption. - Shutdown waits up to 30 seconds for a native startup to reach a safe detach point. An unfinished startup returns a clear failure instead of a false detach receipt. - Queued native work on the same task waits for cleanup. Unrelated tasks remain eligible. - The image digest update rebuilds the Daytona runtime image. No database migration or public API change is included. ## Model Used OpenAI Codex, GPT-6, with repository inspection, code execution, and browser tools. The runtime does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.919.0-canary.10 |
||
|
|
aeef493f4a |
chore(db): keep only the newest 5 drizzle snapshots and stop shipping them (#13687)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - `@paperclipai/db` owns the Drizzle schema and the migration history > - Drizzle writes a full copy of the schema as a snapshot for each generated migration. Each snapshot is now about 1.3 MB. > - The `meta/` folder is about 103 MB. That is about half of each checkout and each worktree. The build also copies it into `dist`, so the published `@paperclipai/db` package is 112.8 MB unpacked. > - `drizzle-kit generate` reads only the newest snapshot. The runtime migrator reads only the `.sql` files and `_journal.json`. > - This pull request keeps the newest 5 snapshots and removes snapshots from `dist`. > - The benefit is a checkout that is about 100 MB smaller, and a published package that is about 1.5 MB instead of 113 MB. ## Linked Issues or Issue Description Refs #11240, #11254, #12333 (earlier snapshot work: diff collapse, binary diffs, drift repair) **What existing behavior does this improve?** The size of the Drizzle migration snapshots in the repository and in the published `@paperclipai/db` package. **Subsystem affected** `packages/db`: migrations and the build. **Current behavior** `packages/db/src/migrations/meta/` holds 141 snapshots (102.8 MB). The size grows faster than the number of migrations, because each snapshot is a full copy of the schema. `build` runs `cp -r src/migrations dist/migrations`. `@paperclipai/db@2026.916.0` contains 147 snapshot files. It is 112.8 MB unpacked and 5.3 MB as a tarball. **Proposed behavior** Keep the newest 5 snapshots. `generate` deletes older snapshots after it runs. `dist` gets only `*.sql` and `meta/_journal.json`. **Reason and benefit** - In drizzle-kit 0.31.10, `generate` sorts `meta/*` and diffs against the last snapshot only (`bin.cjs`, `preparePrevSnapshot`). The [generate docs](https://orm.drizzle.team/docs/drizzle-kit-generate) also say it compares against "the most recent" snapshot. - Gaps in the snapshot history already work. 138 of the 280 migrations never had a snapshot, because they were written by hand before `doc/DATABASE.md` required `generate`. - We keep 5 snapshots instead of 1. This lets a developer undo the latest generated migration, and it keeps the `prevId` chain for recent branches. - Git keeps the history cheaply. The 631 snapshot versions use only 2.2 MB of the pack, because git stores each version as a delta of the previous one. The cost is in the checked-out files, not the clone download. Therefore this change does not use git-lfs and does not rewrite history. Old snapshots stay available with `git show <rev>:<path>`. ## What Changed - `packages/db/package.json`: a new `prune:snapshots` script keeps the newest 5 `*_snapshot.json` files. It is `ls | sort -r | tail -n +6 | xargs rm -f`, which works with the BSD tools on macOS and the GNU tools on Linux. `generate` runs this script after `drizzle-kit generate`. - `packages/db/package.json`: `build` copies only `src/migrations/*.sql` and `meta/_journal.json` into `dist/migrations`. - Deleted 137 older snapshots. `0277`–`0281` remain. - `chat-identity-migration-reconciliation.test.ts`: removed the walk over snapshots `0254`–`0268`. Those files do not change after merge, and the walk would fail after pruning. The journal-order assertions in the same test remain. `migration-snapshot-drift.test.ts` still makes sure that the newest snapshot matches the schema. - `doc/DATABASE.md`: documented the retention rule. - The snapshots were already marked `linguist-generated=true -diff -merge` by `packages/db/.gitattributes` (#11240). No change there. `git check-attr` confirms it. ## Verification - `pnpm --filter @paperclipai/db exec vitest run`: 43 files, 160 tests pass. - `pnpm --filter @paperclipai/db build`: `dist/migrations` contains 280 `.sql` files and `meta/_journal.json`. It is 1.5 MB, compared with about 105 MB before. - `prune:snapshots` was run on macOS (BSD) and in `debian:stable-slim` (GNU findutils 4.10). With 7 fixture snapshots, it keeps the newest 5. When 5 or fewer are present, it deletes nothing and exits 0 on both. ## Risks - Low risk. Runtime migration does not read snapshots. The only commands that read older snapshots are `drizzle-kit check` and `drizzle-kit drop`. No script or CI job calls them, and they still have the newest 5 snapshots. - A branch that is open now can still add its own snapshot. If the branch conflicts, the rule is the same as today: renumber the migration and run `generate` again. - When we upgrade to drizzle-kit v1 (folder per migration), check whether its new cross-branch "commutativity" checks need a longer snapshot history. ## Model Used - Claude Opus 5 (`claude-opus-5`) in Claude Code, with tool use (shell, file edits, web fetch). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
da257c3069 |
Warn when routine webhook URLs may not be publicly reachable (#13684)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Routines can start that work when another app sends a webhook. > - Local and private URLs often cannot receive events from public services. > - HTTPS alone does not make a Tailscale address public. > - This pull request explains these limits during setup and editing. > - Users can still finish setup for senders on their own network. ## Linked Issues or Issue Description Refs #13637. Webhook setup needs a clear warning when the generated URL appears local, private, or unencrypted. The warning must explain how to make the endpoint reachable without blocking private-network use. ## What Changed - Add a shared warning banner to the Connect, Check connection, and Edit webhook views. - Distinguish localhost, private network addresses and domains, HTTP, and Tailscale hostnames. - Explain the difference between Tailscale Serve and Funnel. Link to the Paperclip HTTPS guide. - Add five full-page Storybook examples, design guide examples, and documentation. - Add URL classification tests and a regression test that finishes setup despite the warning. ## Verification - Passed 41 focused URL and trigger-flow tests. - Passed workspace typecheck, workspace build, token gates, and Storybook build. - Browser-tested the Tailscale story through Check connection, Finish setup, and Edit webhook. The warning stays visible and does not block setup. - Open Product / Routines / Webhooks stories 11–15 to review the warning states. - All 54 PR checks passed, including the full test matrix and eight browser shards; two optional Storybook jobs were skipped by workflow policy. - The server supervisor readiness test timed out once in CI, then passed on rerun and locally (6 tests). - The duplicate local full-suite run was stopped after the complete CI matrix passed. - Greptile: 5/5 on the current commit, with no unresolved review comments. ## Risks - URL checks are hints. They do not test DNS, firewall rules, or actual reachability. - A Tailscale hostname can serve either private Serve traffic or public Funnel traffic. The warning explains this uncertainty and permits both. - No API, schema, authentication, or webhook delivery behavior changes. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |