mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
docs/connector-launch-kit-inputs
622
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
42a077385e |
docs(connections): require launch kit inputs and shared brand provenance
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d56a8a39d6 | Merge remote-tracking branch 'origin/master' into HEAD | ||
|
|
c46092f202 |
fix(skills-catalog): validate server-supplied discovery URLs before fetching
The authoring workflow told the reader to curl two URLs an unauthenticated
MCP endpoint chooses: resource_metadata out of the WWW-Authenticate
challenge, and the issuer out of the document that URL returns. A hostile
endpoint could point either at 169.254.169.254, at a loopback service, or
at a public hostname whose DNS record is 127.0.0.1, and a bare curl would
go there with the author's network position.
references/safe-discovery.md adds safe_curl, and step 2 now uses it for
both fetches. It is an allowlist, not a denylist, and it fails closed:
https only, no userinfo, port 443 only, and every resolved address checked
against loopback, private, CGNAT, link-local, unique-local, multicast,
benchmarking, documentation and reserved space. IPv4-mapped, NAT64 and 6to4
forms are unwrapped and the embedded IPv4 address is checked under the IPv4
rules. One bad answer in a round-robin set refuses the whole host.
Two details that decide whether this works rather than looks like it does.
The fetch pins --resolve to exactly the addresses that were validated, so a
second DNS answer cannot move the request between the check and the
connection. And no redirect is ever followed: a 3xx is reported and
refused, and going there means re-running safe_curl on the Location so it
is validated on its own merits.
Executed against adversarial metadata rather than a URL list: a hostile
challenge and two hostile protected-resource documents, parsed the way step
2 says to parse them. localtest.me is the case a string check cannot catch
-- an ordinary public hostname with a real public record pointing at
loopback. Twenty refusal cases, one per class; the redirect refusal against
a live 301; and the documented flow still returns Enterpret's 401 challenge
and protected-resource document unchanged. The two scripts in the document
were extracted and diffed against the tested copies: byte-identical.
Also the four Enterpret acceptance findings that belong in this package:
- Every pinned count in catalog-contract.md was stale. APP_STORE_DEFINITIONS
went 46 -> 56 and SELF_SERVE_MCP_CANDIDATES 43 -> 48 since the file was
written, so following it wrote a failing assertion. The counts are now a
command that reads them out of the checkout in front of you.
- Every line-number citation in both skills is gone, in favour of a file
plus a symbol to grep. Four playbook citations landed on unrelated text;
every claim behind them survives and was re-read at
|
||
|
|
18dac1e1ef |
feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path > - Paperclip manages AI agents and their work. > - Connections give agents governed access to external tools. > - Agents need durable memory across tasks and execution environments. > - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory tools. > - This pull request adds their setup flows behind an experimental toggle. > - It also delivers assigned MCP tools through the native remote Codex runner. > - Operators can connect a provider once and use the same governed tools locally or in Daytona. ## Linked Issues or Issue Description **Problem or motivation** The Apps catalog lacks a complete set of memory providers. Remote native Codex agents also need access to assigned managed MCP tools without receiving provider credentials. **Proposed solution** Add five memory connectors behind the disabled-by-default Experimental memory connectors setting. Use the existing connection setup and permissions UI. Default their tools to Allowed. Preserve company boundaries, operator permission changes, provider scopes, and audit attribution. **Alternatives considered** Direct provider credentials in each sandbox would duplicate setup and bypass the managed gateway. The remote runner instead uses its existing protocol channel to call the gateway on the server. **Roadmap alignment** This maintainer-requested experiment supports the Memory / Knowledge and Connected Apps roadmap areas. It adds provider connections without introducing a separate memory UI. Related connector authoring documentation is tracked in #13692; no duplicate memory-provider implementation was found. ## What Changed - Add provider definitions, official branding, and the experimental setting for all five providers. - Use OAuth for Zep and Supermemory, API credentials for Mem0 and Honcho, and a bundled Cloud API bridge for Cognee with no runtime downloads or subprocesses. - Default memory tools to Allowed and classify destructive actions explicitly. - Fix personal remote credential resolution and propagate provider tool errors. - Relay assigned managed MCP tools to remote native Codex through the runner protocol. Recheck current authority for each call and rotate stale tool contracts. - Add 21 Storybook states and complete OAuth walkthrough fixtures. - Document provider research, sanitized tool inventory, and live local and Daytona proof. ## Verification - Passed workspace typecheck: `pnpm -r typecheck`. - Passed production build: `pnpm build`. - Full local `pnpm test:run`: 13,273 passed, with failures from process/readiness timeouts under parallel load. Reran all 16 affected suites with one worker: 456 passed, leaving two macOS `/var` versus `/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`. No test failures remain unverified. Latest-head remote CI passes all 54 checks (two optional Storybook jobs skipped). Greptile is 5/5 with no unresolved findings. - Latest Cognee gateway regression: 71 passed, including public deployment without a runtime host and immediate recovery after a provider error. Bundled bridge tests: 25 passed. - Browser setup and real agent tasks exercised all five providers. Mem0, Cognee, Zep, and Honcho have successful store/retrieve proof. - Supermemory now has scoped read/write consent. Local storage and real Daytona write, document read, and semantic recall passed; indexing completion was verified before claiming success. - All five providers were exercised through a real Daytona sandbox and its native runner MCP relay. Provider credentials remained on the server. The final bundled Cognee bridge also passed a fresh Daytona store/recall run. All disposable sandboxes were removed and verified absent after testing. - Storybook is rebuilt and contains the experimental toggle, catalog, setup, permissions, and error states. The Zep and Supermemory access steps advance correctly. ## Risks - Provider OAuth scopes and plan limits remain independent of Paperclip tool permissions. An Allowed tool can still be rejected by the provider. - Providers can queue memory indexing; save acceptance does not prove that semantic recall is ready. - Remote tool contracts must stay synchronized with current connection authority. Regression tests cover revocation and stale contracts. - This change has no database migration. Existing connections remain usable when the experimental catalog toggle is disabled. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository edits, code execution, and browser automation. The exact context window size is not exposed in this session. Live acceptance agents used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f3a214fe77 |
Fix relative symlinks in secondary sandbox repositories (#13953)
## Thinking Path > - Paperclip manages agents and their task workspaces. > - A task can use several independent Git repositories. > - Sandbox staging copies secondary repositories from temporary clones. > - The copy changed relative symlinks into absolute host paths. > - Those links broke skill discovery and made workspace restore fail. > - This change preserves link targets while keeping extraction checks intact. ## Linked Issues or Issue Description **What happened?** Staging a secondary repository rewrites a link such as `.claude/skills/demo -> ../../skills/demo` to an absolute path in a temporary Git clone. That clone is then removed. The link is broken in the sandbox, and Daytona refuses the outbound archive during workspace restore. An agent can finish its turn but still have its run fail during restore. **Expected behavior** Repository links keep their original targets after staging. Links within a repository remain usable, and changes return to the local checkout. Unsafe outbound archive links still fail before extraction. **Steps to reproduce** 1. Create a project with a primary repository and a secondary repository. 2. Commit a relative skill directory link in the secondary repository. 3. Stage the workspace for sandbox execution and inspect the copied link. 4. Restore that repository through Daytona. Before this fix, the link points at a removed host temporary directory and restore rejects it. **Paperclip version/commit** Reproduced on `0f8750627f11d855552abce9a837d7f3b67c9ddf` in the multi-repository sandbox path. Related: #13442 introduced multi-repository provisioning. #13882 adds native Grok but keeps legacy adapters; #12991 addresses Grok instruction isolation and leaves skill staging unchanged. Searches of open/closed PRs and open issues found no direct fix for this copy behavior. ## What Changed - Set `verbatimSymlinks: true` when copying secondary Git clones. This [Node option](https://nodejs.org/api/fs.html#fspromisescpsrc-dest-options) preserves the stored link target instead of resolving it against the temporary source. - Test directory, file, chained and dangling links after temporary-clone cleanup. Assert the copied Git checkout remains clean. - Cover skill-link reads, edits and restore in fresh, warm-adoption and durable-seed workspace modes. - Extend Daytona checks for valid relative directory links and rejected absolute targets. - Document the staging behavior. ## Verification - Four regression cases fail without the source fix: one clone test and three staging modes. - `pnpm exec vitest run packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts`: 301 passed. - `pnpm -r typecheck` and `pnpm build`: passed. - `pnpm test:run` was started locally, then stopped after the full Linux CI suite passed. It did not finish locally; this is not a full local-suite pass. - Full PR CI: 53 checks passed, two conditional skips, on `07b30ff298baab327a4e60708ade899935688a3f`. Three server shards were interrupted by runner shutdowns; the unchanged mobile repository test timed out waiting for a disabled Save changes button. One same-commit failed-job rerun passed. Original attempts remain in [run 36036538369](https://github.com/paperclipai/paperclip/actions/runs/36036538369). - Greptile: 5/5 on the same head, no review threads or actionable findings. - Diff scanned for secrets and private identifiers; no matches. ## Risks Low risk: the production change is one copy option. It preserves symlinks instead of following or materializing their targets. Daytona extraction guards, workspace exclusions, authentication and database behavior do not change. This prevents corruption in newly staged snapshots. It does not rewrite an already corrupted warm workspace or durable seed; those need fresh staging from the source checkout. No live provider run or customer-task replay was performed. The tests use real Git, filesystem and tar operations with mocked provider transport. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f8750627f |
fix(codex): preserve typed ACP quota classification and reset time (#13945)
## Thinking Path > - Paperclip runs agents and recovers failed tasks. > - Codex ACP can report usage exhaustion as a typed terminal failure. > - The shared engine removes provider text before it stores the result. > - Codex did not use the existing terminal classifier hook, so a quota failure became a generic turn failure. > - This change classifies explicit usage exhaustion before that text is removed. > - Recovery can then use quota backoff and the supported reset clock. ## Linked Issues or Issue Description **What happened?** A Codex ACP `limit` failure with explicit usage-exhaustion text produced `acpx_turn_failed`. Recovery could schedule the ordinary short retries because the quota classification and reset time were lost. **What did you expect to happen?** Keep the failure visible and use the existing provider-quota wait. Preserve a supported reset timestamp without storing provider text. **Steps to reproduce** Return a terminal ACP session failure with category `limit` and title `You've hit your usage limit for GPT-5. Switch to another model now, or try again at 4:30 PM (America/Chicago).` The real child-process regression tests exercise both pinned ACPX versions in persistent and oneshot modes. **Paperclip version** Base commit: `32573876d4`. Related work: #13651 and #13831 provide the shared hook and Claude classification. #11854 includes quota handling as part of optional credential rotation, but reads the already-sanitized result; this patch handles the typed terminal boundary without adding rotation. #13549 reads recovery text after this boundary and cannot recover discarded provider text. #9011 concerns the Codex CLI backoff. This change leaves those other mechanisms in place. ## What Changed - Register a Codex terminal-failure classifier with the existing ACP engine hook. - Recognize explicit usage exhaustion only in a typed `limit` failure. - Reuse the Codex reset-time parser and existing provider-quota recovery fields. - Leave context, turn, rate, budget, storage-capacity, and unknown failures on their existing paths. - Test real ACP children, both dependency patches, privacy, and recovery classification. Document the boundary. ## Verification - 88 focused tests passed across Codex ACP, parsing, and server recovery classification. - An initial cross-adapter run passed 53 tests, including the existing Claude quota suite. - Removing only the classifier registration makes five new integration tests fail; restoring it passes all 19 new tests. - `pnpm -r typecheck` and `pnpm build` passed. - Full Linux CI passed on `b10d60e002`, including all unit/integration shards, browser shards, typecheck/build, runner checks, and release/package gates. - The duplicate local `pnpm test:run` reported three skill-cache failures and two runner-suite failures. It is not claimed as a full local pass. - The three skill-cache failures reproduce in a clean worktree at the unchanged base commit (3 failed, 107 passed, 3 skipped across the skills and runner files). The cache rename reports `EACCES` on macOS. - An isolated runner-suite run reports two embedded PostgreSQL startup failures before its assertions (37 tests pass). No runner source is changed; the Linux CI runner and server suites passed. - Greptile reviewed the current head at 5/5 with no actionable findings or unresolved threads. ## Risks - Explicit usage-exhaustion failures now wait for quota recovery instead of short generic retries. - Unknown wording retains the existing behavior. The generic historical terminal-limit message alone cannot establish quota exhaustion. - Reset parsing keeps the existing Codex clock formats. A missing or unsupported reset uses the existing quota backoff. - No schema, credential, UI, dependency, or deployment changes. Provider text stays in memory and is absent from results and logs. ## Model Used OpenAI GPT-6 via Codex. Exact runtime model variant and context-window size were not exposed. Used reasoning, repository inspection, editing, and local tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b0155a681a |
feat(slack): connect Paperclip conversations and scheduled messages (#13920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack conversations use the same tasks and agents as the Paperclip board. > - A board reply must reach that Slack conversation and let the agent continue the work. > - An assigned agent also needs its Slack tools during normal tasks and scheduled routines. > - Both paths must keep the linked user's authority, delivery rules, and conversation history. > - This pull request adds those paths and reduces setup friction for Slack bots. ## Linked Issues or Issue Description **Subsystem affected** Server orchestration, Slack connector tools, shared contracts, and chat setup UI. **Problem or motivation** Replies entered in Paperclip did not provide a complete round trip to the linked Slack thread. Slack and board wakeups could select different model sessions for the same task. Agents also lacked their assigned Slack tools outside Slack-origin work, which prevented a routine from sending its responsible user a briefing. Inviting a bot could leave the new channel disabled. **Proposed solution** Mirror human board messages with author attribution and route the agent result to the same thread. Use the same session key across both entry points. Supply Slack tools to the connection's assigned agent in normal tasks and routines, using the current responsible user's verified link. Enable newly invited channels while preserving explicit disabled choices. Add a browser-agent setup prompt to the Slack wizard. **Alternatives considered** A separate Slack scheduler or task dispatcher would duplicate existing Paperclip workflows. Reusing the connection owner's identity would grant the wrong authority. Replaying old channel history could start unintended work. This change uses ordinary task wakeups, routine dispatch, and Slack's original invitation mention event instead. **Roadmap alignment** Extends the shipped Scheduled Routines and governed Apps capabilities. It does not add a separate task lifecycle. Related work: #13828 and #13809. Related test stabilization: #13877. The existing plugin Slack-control proposals are separate from this built-in connector change. ## What Changed - Queue human Paperclip messages for the original Slack thread with display-name attribution and stable delivery identities. Require the author’s current linked Slack identity and recheck access before delivering messages or agent replies. - Apply pause, dependency, cancellation, and closed-workspace guards before explicit Board sends request work and again when the durable outbox dispatches it. - Route agent results back to Slack and preserve model-session continuity, including replies that reopen completed tasks. - Resolve assigned Slack connections for normal agent tasks and routines. Recheck the responsible user's link, membership, and permissions at execution. - Add `slack_open_dm` for the responsible user's bot DM and request the `im:write` scope. - Enable newly discovered invited channels. Keep explicit OFF choices and normal admission and deduplication rules. - Add a copyable Slack setup prompt for a computer-use agent, with Storybook coverage. Share the prompt-button component with GitHub. - Update Slack tool documentation and runtime instructions. - Stabilize the mobile project browser test by waiting for the final canonical route before editing, preserving all persistence assertions. ## Verification - Live staging: invited the bot after the first mention. The channel became enabled and the bot answered that original mention. - Live staging: a normal Paperclip reply appeared in Slack with author attribution. The agent completed the calculation and replied once in the original thread and in Paperclip. - Live staging: a codeword entered in Slack was recalled from Paperclip. A following Slack calculation used the result from the Paperclip turn. Run metadata confirmed the same model session for both entry points. - Live staging: a scheduled routine used `slack_open_dm` and `slack_post_message` to deliver one DM. The existing app was reinstalled with `im:write`. The test routine was paused after verification. - Before the master merge: 397 focused feature tests passed. The continuity fix passed all 76 issue comment/update route tests and six focused route/integration cases. Typecheck, build, and token gates passed. - Review fixes: 47 focused integration cases passed, covering link revocation/replacement, private membership removal, guarded outbox dispatch, concurrent workers, lost scheduler responses, a real one-connection pool, exact reply provenance, and attachment retries. All 99 issue-comment route tests and the Slack catalog browser test passed. - Full local typecheck, production build, token gates, and module-boundary checks passed. The full local test command passed 25,915 tests before a 15-second timeout in `issue-thread-interaction-routes.test.ts`; that entire suite passed on isolated rerun (81 tests). Remaining serialized coverage is provided by the current-head CI shards. - An unchanged Cursor adapter test hit its 10-second limit in CI; all five tests in that file passed on a local rerun in 3.11 seconds, and the failed CI shard passed on its single retry. - The preview-server readiness test passed a local rerun (28 tests). The mobile-project readiness fix passed three repetitions of both browser tests (6/6). - Final commit `64ac0d9897f4353375996f1b1b38e5040bdeb0a0`: all CI gates passed, including all eight browser shards, all server/chat suites, typecheck, build, runner checks, and security checks. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - Human messages on a Slack-linked task now publish to its Slack thread. The task banner states this behavior. Incoming Slack messages and internal agent bookkeeping must not echo back. - Normal tasks and routines can now use the assigned bot. Authority remains bound to the current responsible user's link; it does not fall back to the connection owner. Revocation, private-context limits, and queued-write checks still apply. - Existing Slack apps need `im:write` and a reinstall to open DMs. Other existing capabilities remain available without that scope. - New invited channels default to enabled. Explicit disabled choices remain disabled. Channels created by bot tools still require a person to enable responses. - No database migration or new provider credentials are required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, GitHub CLI, and browser tools. The runtime does not expose a more specific model build identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c341588bdb |
fix(ui): add Cloud invitations to the Members page (#13922)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People manage collaborators from the Members page. > - Cloud manages invitations outside the tenant's local invitation system. > - The local Invites tab is hidden on Cloud, so this page has no way to invite a person. > - This pull request adds an Invite people action for the current Cloud stack's owner or admin. > - The action opens the existing Cloud People settings for that stack. ## Linked Issues or Issue Description **What happened?** A Cloud owner opens Organization Settings → Members and finds no invitation action. The tenant-local Invites tab is hidden, and the page does not link to Cloud's invitation flow. **Expected behavior** Cloud owners and admins can start an invitation from Members. **Steps to reproduce** 1. Sign in to a Cloud-managed instance as the current stack's owner or admin. 2. Open Organization Settings → Members with `company.invites` hidden. 3. Look for an invitation action beside the page heading. **Paperclip version or commit** `7b7c4d4172d6aac14919e2682b702ae87bc17653`. **Deployment mode** Cloud-managed, authenticated. Related search: #2388 proposes broader member-management UI. This change only connects the existing Members page to Cloud invitations. No duplicate Cloud invitation action PR was found. ## What Changed - Add **Invite people** beside the Members heading for the current Cloud stack's owner/admin. - Read the role from the authenticated Cloud portfolio. Ownership of another stack does not enable the action. - Navigate to the current stack's People settings on the configured Cloud origin. Keep the local Invites tab hidden when configured. - Cover allowed roles, denied roles, loading, failed refresh, missing configuration, current-stack selection, and self-hosted behavior. Document the navigation contract. ## Verification - Focused Members and Cloud link tests: 23 passed. - Full UI suite: 6,669 passed across 634 files. - UI typecheck and `pnpm check:token-gates`: passed. - `pnpm build`: passed. - `pnpm -r typecheck`: passed. - All PR CI checks passed, including the full general, serialized, browser, and runner test jobs. The duplicate local repository-wide `pnpm test:run` was stopped after CI completed; it is not reported as a local pass. - Manual acceptance after tenant rollout: an owner/admin opens Members, selects **Invite people**, and reaches the same stack's Cloud People settings. A member does not see the action. ## Risks - The action needs a tenant app update before it appears on an existing stack. - The portfolio request must identify the current stack and its role. The action stays hidden when that information is unavailable or the request fails. - Cloud rechecks invitation authorization at the destination. No schema, API, or invitation-acceptance behavior changes. ## Model Used - OpenAI Codex, GPT-6, with reasoning, code execution, and repository tools. The session does not expose a more specific model identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8ee8f1fd6e |
ci: retire recurring public cloud image builds (#13827)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Core publishes standard images and source verification for downstream services. > - Managed services can now compose private images from the signed standard image. > - Core still builds a second public cloud image on every master push and release. > - That duplicate producer consumes build capacity and retains an obsolete readiness contract. > - This pull request retires recurring cloud publication while preserving the standard producer and rollback artifacts. ## Linked Issues or Issue Description Refs #13797 and #13789. Related: #12856 changes image dependency packaging; it does not retire this producer. **What existing behavior does this improve?** Core's recurring Docker publication and Cloud readiness workflow. **Current behavior** Master pushes call the legacy cloud publisher from Cloud readiness. Release tags and manual Docker runs call it too. Canary promotion also requires the legacy image. **Proposed behavior** Publish standard Core images and retain `Cloud source verified v1`. Let downstream services build their managed image. Keep explicit commit previews and existing images available. ## What Changed - Remove `docker-cloud.yml`, its master and release callers, and its unused cache selector. - Remove the legacy image/migrator wait and `Cloud deployable v1` job. Keep the full source verification workflow and exact source-proof name. - Make canary promotion inspect and promote the standard image only. - Preserve signed standard-image publication, direct migrator publication, and explicit `release.yml` previews. The preview path still uses the Dockerfile `cloud` target. - Update workflow, preview, build-stamp, and packaging tests. Exercise the promotion shell with mocked registry commands, including missing-image and missing-tag cases. - Document frozen legacy aliases, consumer requirements, preview compatibility, and rollback retention. ## Verification - All 377 workflow tests pass: `node --test .github/scripts/tests/*.test.mjs`. - All 129 release-registry tests pass: `pnpm test:release-registry`. - Focused source-proof, standard-image, preview, and workflow tests pass: 256 tests. - Focused image packaging/build-stamp tests pass: 16 tests. - Actionlint passes on all three changed workflow files. `git diff --check` passes. - Full local `pnpm build` and `pnpm -r typecheck` pass. - The policy follow-up updates an old assertion that required the removed readiness job. All 37 source-proof/release-workflow tests pass locally. - Full local `pnpm test:run` did not complete successfully while the Mac ran out of disk space. No full-suite pass is claimed. Removed 1.2 GiB of generated Cargo output from this isolated worktree with `cargo clean`. GitHub CI passed on the final head: 52 successful checks and 2 optional skips. - Fresh Greptile review for `4f5fe1951f0bd7f7739cf6655d395ff78f1ed944`: **5/5**, successful current-head check, zero review threads. - September 23 refresh: the unchanged PR head merges cleanly with current master `db8f8fe5b73a2697684a30261b0d306a9c631aba`. In an isolated temporary worktree, all 377 workflow tests and 29 release/preview tests pass on the combined tree. `git diff --cached --check` passes. - Refreshed Actionlint workflow validation passes with ShellCheck disabled. Full Actionlint reports the same 10 existing ShellCheck diagnostics as master, with no added diagnostics. No source changes or new PR commits were needed. - The full local build/typecheck and current-head Linux CI results above remain the verification for the unchanged PR head. They were not rerun for this metadata-only refresh. No image publication or tenant deployment was initiated for this refresh. ## Risks **Deployment prerequisite satisfied (September 23):** The combined cleanup release is deployed to staging and production, and production Support is verified. Active managed-fleet automation uses standard-image composition. Explicit immutable previews remain supported by the retained preview publisher. This PR is ready for maintainer review; keep auto-merge disabled and wait for explicit merge authorization. - A consumer still selecting `Cloud deployable v1` will stop advancing at the last legacy-ready commit. Confirm active automatic consumers use the standard-image composition contract before merge. - Legacy cloud release-channel aliases stop advancing. Standard self-hosted aliases continue. - This PR deletes no registry images, cache tags, migrators, credentials, or runner infrastructure. Existing immutable releases remain usable for rollback. - Explicit legacy previews remain for commit-specific operator deployments. Retiring that compatibility path requires a separate consumer migration. - These changes affect CI publication, not database schema or application behavior. ## Model Used OpenAI Codex, GPT-6. The runtime does not expose a more specific model identifier or context-window size. Used repository inspection, reasoning, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
066a4e8019 |
doc(connections): Enterpret acceptance test against the generic MCP path
PAP-18519. The recorded answer to the question both connector skills ask first -- does this vendor need a catalog entry at all -- for the vendor Michael named as the acceptance case. Reproduces Enterpret's public metadata from this runtime (no drift on any stated fact, three additive observations), then exercises Paperclip's client-resolution ladder against a loopback mirror of that metadata, on a disposable self-hosted instance started for the pass and torn down after it. No Enterpret credential exists, none was requested, no client was registered there, no consent given, no tool called. The three questions the ticket asked, answered with executed output: - Discovery resolves completely from the 401 challenge alone, in two metadata fetches. The catalog definition should ship serverUrl only; a hard-coded endpoint pair would become authoritative and skip discovery for good. - The ladder falls through to DCR, but the public-HTTPS condition the playbook names is not what decides it -- this authorization server never advertises a Client ID Metadata Document, so DCR is the only rung on any deployment. Asserted for both callback origins. - Revoke is enforced entirely Paperclip-side. Zero revocation calls were made because there is nothing to call: revocation_endpoint is read nowhere in server/src and has no slot in the persisted oauth config. Scenario 8 cannot read verified for this provider. The catalog-entry recommendation is yes, and the reason is governance rather than connectivity. With no annotations to infer from, all ten documented tools classify as read -- including run_graph_query and execute_cypher_query, which execute operator-supplied queries against the vendor's customer-feedback graph. The wizard shows the operator ten reads and no changes. The nine runbook scenarios are marked per deployment. Every row naming Enterpret itself is not run, for one reason: no authorized credential. Cloud is untested throughout. The mirror server is embedded in the doc rather than added as a script, so the skills package keeps its markdown_only trustLevel. Extracted and booted from the published text to confirm it runs as written. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
fff410dfe7 |
fix(ui): retain bounded context for browser render errors (#13904)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its browser error boundaries recover from failed renders and report the exception when error monitoring is enabled. > - The boundaries already receive the React component stack, but discard it before reporting the error. > - A minified DOM insertion error can therefore lack enough context to identify the affected component. > - Browser translation can replace text nodes that React still uses as insertion anchors. > - This pull request preserves bounded component names and browser state for the next failure. > - Maintainers can locate the failed component without collecting page content or customer URLs. ## Linked Issues or Issue Description Related: #13719 attributes browser errors to the loaded release. #13784 supplies the browser environment. This change adds context to error-boundary reports. **What happened?** A browser error boundary reports a DOM `NotFoundError` with only a minified JavaScript stack. The React component trace is logged to the console, which production builds remove. Browser translation is a plausible cause, but the report cannot identify the component or confirm the translation marker. **Expected behavior** An error-boundary report includes a bounded component trace and limited browser state. It excludes component props, text, HTML, element IDs, arbitrary CSS classes, page URLs, and query strings. **Steps to reproduce** 1. Render a component with conditional content before a text node. 2. Replace that text node with a translation element outside React. 3. Enable the preceding conditional content. React tries to insert before the detached text node and throws `NotFoundError`. 4. The boundary shows its recovery UI, but the old report loses the component trace. **Paperclip version or commit** Based on `6681c71b40`. The regression test reproduces the DOM mutation with the installed React version. **Deployment mode** Built browser UI with optional Sentry monitoring enabled for the signed-in session. ## What Changed - Pass the component stack and boundary kind from both error boundaries. - Keep at most 40 component names from at most 16 KiB of stack input. Drop locations and unrecognized lines. - Snapshot document readiness, visibility, and the browser translation root-class marker before the asynchronous reporting queue runs. - Attach diagnostics to that event only. Preserve the existing monitoring gate, sign-out behavior, and original exception if diagnostics fail. - Preserve function names in production bundles. Test the actual Vite production pipeline. - Document the fields, privacy limits, and translation-marker limitations. ## Verification - Focused diagnostics, boundary, real Sentry SDK, and production-build tests: 49 passed. - The translation DOM-mutation test reproduces `NotFoundError`, verifies the failed component trace, and keeps the recovery UI usable. - The real SDK test checks emitted events, private fixture exclusion, event isolation, and no capture after sign-out. Its transport stays in-process. - `pnpm check:token-gates`: passed. - [Greptile review](https://github.com/paperclipai/paperclip/pull/13904#issuecomment-5804220318): 5/5 on `92a2a23d9c`, with no review threads. - `pnpm -r typecheck` and `pnpm build`: passed. - Full CI unit and integration test matrix: passed, including all server, chat, workspace, runner, and serialized suites. - Local `pnpm test:run` was started, then stopped after the equivalent full CI matrix passed. The unsharded local run was not completed and is not claimed as a pass. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/project-repositories.spec.ts --repeat-each=3 --trace=on --reporter=line`: six tests passed against the production build. - Initial CI browser shard 8 timed out after the repository form replaced an enabled Save button with a disabled one before the click. The retained page snapshot shows the original repository selection. No error-boundary fallback appeared. [CI attempt 2](https://github.com/paperclipai/paperclip/actions/runs/35929990870/attempts/2) passed the failed jobs on the same commit. The full PR check matrix is green. This confirms an intermittent failure, but does not establish the cause of the first failure. - Full browser builds with and without name preservation passed. The initial JavaScript chunk grows from 1,571,919 to 1,668,453 gzip bytes (+6.1%). Total JavaScript across all chunks grows by 219,610 gzip bytes (+5.3%). ## Risks - This adds diagnostics. It reproduces a translation failure mode but does not identify or repair the specific application component from a past report. - The root-class marker is a hint. Other translation tools may omit it, and its presence does not prove causation. - Name preservation increases bundle size as measured above. No source maps are published by this change. - The error is still reported. DOM operations, browser translation, and recovery behavior are unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code editing, terminal tools, and test execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6681c71b40 |
fix(ci): stabilize chat startup and close the initial live-update gap (#13895)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Browser tests check chat state across navigation, reload, and agent runs. > - Runtime tests check that a service is ready before Paperclip publishes its address. > - CI for #13891 failed in these paths, then passed attempt 3 with the same code. > - The browser traces stopped during development asset startup. A separate sidebar assertion used an unstable focus path through the rich editor. > - This PR gives browser tests fresh built assets and a direct keyboard path to the star button. It adds service-worker reload coverage under CPU throttling. > - A later browser failure exposed a real reload race: comments can change between the first query and the first live subscription. The UI now refreshes active queries when that subscription opens. > - Runtime fixtures now have separate registry state and better failure evidence. The readiness deadline remains unchanged. > - A later CI run exposed a wall-clock backoff assertion and three authorization cases sharing one test lifecycle. The PR anchors the assertion to transport time and separates the cases. ## Linked Issues or Issue Description Refs #13891. Related runtime ownership and cleanup work: #11791, #11389, #11278. Evidence: [original CI run, attempt 3](https://github.com/paperclipai/paperclip/actions/runs/35910335089/attempts/3). Attempt 3 passed both affected shards without code changes. This PR is separate from the wake-payload change, which has since merged. | Failure | Diagnosis and classification | | --- | --- | | Sidebar blank page and retry text missing after reload in CI | Both traces show a blank document before React startup. Vite connects, but the failing page makes no application API requests. The service worker forwards the unbundled development module graph. The retry response is already stored and its process-adapter run succeeded before reload. This places the failure in browser bootstrap, not wake payload or reply persistence. The exact reason the development module graph stopped is not established by the retained trace. The harness now serves built assets, and the new test covers startup and controlled reload at 4x CPU throttling. | | Star opacity remains zero on macOS | Reproduced on unchanged master `f55759942b`. The old test clicked the rich editor, focused the star, then used Tab and Shift+Tab. An instrumented baseline run captured the sequence: Tab moved from the star to the next sidebar link, then the editor bundle called `focus()` on its contenteditable before Shift+Tab. That key reached the composer, where it is a work-mode shortcut. This confirms a pending editor selection update stole focus; it was not the browser skipping sidebar buttons. The assertion did not prove the star retained focus. The test now moves the pointer away, focuses the preceding sidebar link, presses Tab once, and asserts both actual focus and opacity. This is test synchronization and keyboard traversal, not a demonstrated CSS defect. | | Runtime readiness exceeds 10 seconds | The old error only says `fetch failed`. There is no child startup output in that failure, so it cannot distinguish slow process startup from a refused or stalled probe. It passed unchanged locally and took 3.416 seconds in attempt 3. Resource contention is plausible but unproved. This PR does not claim a proven historical runtime root cause: it isolates fixture registry/log files, checks ports before spawn and after stop, checks live backends at publication, and preserves transport errors, probe count, elapsed time, and fixture startup timestamps for the next occurrence. | | Later CI: recovery reply disappears after reload | Product synchronization bug, distinct from the blank-page bootstrap failure. Reproduced locally with a trace: the comments request started at `21:56:07.443`, the server saved the reply at `.520`, and the first live subscription started at `.541`. The reload fetched a successful run and its complete log, but missed the comment event. The provider refreshed after reconnects only. It now refreshes active queries on the first connection too. A deterministic regression test fails before the fix. No browser assertion was changed. | | Later CI: Slack backoff and authorization tests | The 30-second backoff assertion required more than 25 seconds to remain when it read the saved action. CI spent 10.311 seconds in the test, exceeding that five-second allowance. A local six-second read delay reproduces the failure; the new transport-anchored lower and upper bounds pass the same fault injection. The neighbouring authorization test ran three independent fixtures in one test and timed out at 15 seconds. Each case now has its own fixture cleanup and the normal per-test deadline, so earlier cases do not remain active during later global worker sweeps. No specific production slow call was established. | | Self-hosted runner loses communication or shuts down | The original lost-communication failure has no assertion. On final-head [attempt 1](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/1), Build, Runner Vitest 1/2, and chat 2/3 ran on three separate fleet instances. All received a runner shutdown signal at `22:04:55 UTC`, within 35 milliseconds, then cancellation. Server shard 6/12 received the same shutdown signal one minute later. All four were Spot `m7i-flex.xlarge` instances in `us-east-1a`. No test assertion or build error preceded those stops. This is infrastructure interruption; the reason the fleet stopped the runners is not available in job logs. | ## What Changed - Build the browser fixture UI into the static server's preferred directory, `server/ui-dist`, and disable Vite middleware. Ignore these generated assets. Start the source CLI directly from the repository root, as required by the CLI invocation safety contract. - Keep all existing chat assertions. Add first takeover and three service-worker-controlled reloads under CPU throttling. Assert that the page uses built module assets and that opening it creates no chat task. - Use forward keyboard traversal from the Zeta link to its star. Check focus before checking the reveal style. - Give each runtime exposure test a temporary Paperclip home and restore environment state after process cleanup. - Check that reserved ports are free before spawn, serve the fixture response before exposure, and become free after stop. - Include the nested fetch error, probe count, and elapsed time in readiness failures. Add a unit test for this diagnostic contract. - Refresh active queries when the first live connection opens, covering events missed during initial page loading. Keep reconnect toast suppression unchanged. - Measure Slack retry timing from the transport attempt and recovery completion. This checks the full provider-requested delay without spending a small wall-clock allowance on unrelated processing. - Run each Slack authorization-revocation scenario as a separate test, with cleanup between cases. All assertions remain. - Document the browser fixture's build and serving mode. No configured assertion deadline, readiness deadline, retry count, or skip was added. ## Verification - Final-head [Linux CI, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/2): **green**. All 52 check runs passed; two conditional checks were skipped. The legacy Snyk status also passed. No pending or failed checks remain. - `pnpm -r typecheck` passed. - `pnpm build` passed again after the final UI fix. - Focused runtime suites: 38 passed, 3 existing platform skips. - CLI invocation safety suite: 39 passed. - Slack timing negative control: inserting a six-second delay before reading the saved action fails the old assertion. All four revised authorization/backoff cases pass with that same delay. The diagnostic delay is not committed. - Live update suites: 92 passed, including the new regression. UI typecheck and token gates passed. - Recovery browser negative control: the unchanged tests reproduced the missing reply (1 failed, 9 passed). After the first-connection fix, all six recovery paths passed twice (12 passed). Browser assertions and deadlines are unchanged. - Server typecheck passed after the Slack test adjustment. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-chat --shard-index 0 --shard-count 3`: 342 passed. The other 682 tests belong to the remaining shards; collection verified exact coverage. - Final direct-source CLI launch: all five chat session tests passed. - `PAPERCLIP_E2E_PORT=32993 PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/agent-chat-sessions.spec.ts --repeat-each=3`: 15 passed, including nine controlled reloads at 4x CPU throttling. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-server-without-chat --shard-index 10 --shard-count 12`: 55 files passed, 867 tests passed, 21 existing skips. An earlier run could not initialize PostgreSQL because this Mac exhausted its System V shared-memory slots. After reclaiming the orphaned segment from this task's stopped browser server, the full shard passed. - On final commit `1f1fafc08d`, [browser 4/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400837947) passed all 16 tests; [browser 8/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838155) passed all 23 tests; [server 11/12](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838328) passed 887 tests with one existing skip. The readiness lifecycle case took 1.842 seconds. Chat 1/3 also passed. All four jobs interrupted by runner shutdowns passed unchanged on their single rerun. - Greptile reviewed `1f1fafc08d`: **5/5**, with no unresolved comments. - Full local `pnpm test:run` was attempted and stopped after confirming failures outside this patch: a sibling `skills` directory shadows bundled Slack/AgentMail skills; macOS rejects rename of read-only skill-cache directories (`EACCES`, also reproduced in an isolated filesystem probe); and the host exhausts PostgreSQL System V shared-memory slots. Focused reruns confirmed these limits. The complete Linux CI run is the repository-wide verification; the full local run is not green. - Baseline evidence: the original sidebar test failed on unchanged master; a development-mode run at 4x CPU throttling passed six selected cases, so CPU pressure alone did not reproduce the CI bootstrap stall. ## Risks - Default browser tests now exercise the shipped static UI. They no longer implicitly cover Vite middleware or HMR; use the development server for those checks. - The new browser startup test uses Chromium CDP, matching the only configured browser project. - Opening a live subscription now causes one active-query refresh to close the initial event gap. This adds startup API reads but no recurring poll. - Runtime behavior and deadlines are unchanged except for error details. The historical readiness stall remains unconfirmed; a green rerun alone cannot establish its cause. - Local verification runs on macOS. The runtime lifecycle checks also passed on Linux CI. ## Model Used - OpenAI GPT-6 through Codex. The session identifies the model family as GPT-6; an exact served model ID and context window size are not exposed. Used reasoning, repository inspection, shell tools, code editing, and test execution. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #123` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted suites passed; full local host limits are listed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24429024e7 |
feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps gives agents governed access to external tools through stored credentials. > - Routines start work when an external service sends an event. > - Fireflies provides meeting transcripts and summaries through an official hosted MCP server. > - This PR adds that connection and accepts signed meeting events through the shared app webhook flow. > - Agents can review completed meetings with the same permissions and audit records as other work. ## Linked Issues or Issue Description **Problem or motivation** Operators need agents to read Fireflies meetings and start follow-up work when a summary is ready. The Apps catalog lacks Fireflies. The shared app webhook flow needs to accept its signed deliveries. **Proposed solution** Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer API key. Extend the existing Another app or script flow with signed webhook support. Verify the raw-body signature and pass the JSON payload as external data. Select Meeting Summarized in Fireflies. Deduplicate identical signed deliveries, including setup deliveries. **Alternatives considered** A separate REST connector would duplicate the governed MCP path. Polling, legacy V1 payloads, and automatic provider-side webhook registration are outside this change. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps and Scheduled Routines surfaces. It adds a provider to those systems. It does not introduce a second integration framework. **Additional context** A GitHub search found no existing Fireflies issues or PRs. Provider references and verification limits are in `doc/connections/FIREFLIES.md`. ## What Changed - Add the official Fireflies catalog definition, generated registry, provider evidence, and branded artwork. - Reuse Access → Connect, dynamic discovery, Permissions, vault storage, policy, and audit behavior. - Classify Fireflies sharing, movement, and access revocation as writes. - Preserve Off and Ask first restrictions during OAuth reauthorization and API-key replacement. New actions retain normal defaults. - Add `app_webhook` authentication to the shared Another app or script flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier `fireflies_hmac` triggers and revision snapshots for compatibility. Existing text columns need no migration. - Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact request body. Preserve generic event payloads and deduplicate identical signed requests. - Keep the routine wizard generic. Show one webhook URL and secret in Another app or script. Keep all new app webhook event names provider-neutral. Keep provider setup instructions in the connector documentation. - Pass generic webhook JSON to the task in an explicit external-data block, capped at 16,384 characters. Keep strict meeting validation for existing legacy Fireflies triggers. ## Verification - Feature implementation commit `0882dc8a1`: all 54 CI checks passed; two conditional Storybook checks skipped. This includes full tests, typecheck, build, browser E2E, canary dry run, and security checks. Greptile rated this commit 5/5; all review threads are resolved. - Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed on the final code. Targeted connector, gateway, webhook, revision, and UI suites passed during implementation. After the provider-neutral follow-up, all 84 app-webhook and routine-service tests passed; the final payload-to-task assertion also passed in the 72-test routine suite and a clean-config rerun. - The long local `pnpm test:run` invocation started before the final edits and was stopped after the final-commit CI suites passed. It reported one generic webhook test failure while those files were changing; that test and the entire routine suite passed on the final source, including a clean-config reproduction. The interrupted local run is not counted as a full-suite pass. - In the embedded browser, completed official OAuth consent and discovered 20 live actions. Real meeting listing, transcript retrieval, and summary/action-item retrieval succeeded as the selected agent. Turning a live read Off blocked its test; catalog refresh preserved the restriction. - Embedded-browser Another app or script setup, back/save/resume, narrow layout, and a signed synthetic Fireflies delivery succeeded. The UI reported authentication passed without creating a task. Fixtures cover signature tampering, malformed requests, ordinary app event names, duplicate/setup deliveries, rotation, revisions, pause/archive, and company isolation. - Existing MCP browser suite: 8 passed and 2 provider-dependent cases skipped. Branding checks passed; connector artwork and webhook setup were checked at desktop/mobile widths and in light/dark modes. - An unauthenticated POST to a correctly formatted public webhook URL reached the staging tenant verifier through the existing Cloud gateway. - A real Fireflies webhook delivery remains unverified. A staging callback is available for the operator walkthrough. Live API-key authorization, credential expiry, and a new meeting's summary completion were not tested against the provider. Fixtures cover these protocol and lifecycle paths where applicable. - Storybook follow-up `c54174faa`: 27 production-component stories cover every UI change, with a source-to-story map in the connector documentation. Static Storybook build, UI typecheck, token gates, and Playwright checks for all stories and the mobile footer pass. All PR checks passed for this Storybook follow-up; Greptile reviewed `c54174faa` at 5/5. ## Risks - Fireflies may change its hosted MCP tools or OAuth behavior. Tool discovery stays dynamic. Experimental search/fetch tools are not required. - Public webhook setup requires HTTPS and a separate signing secret. Fireflies normally emits events for meetings owned by the configuring account. - Reauthorization touches shared MCP permission code. Regression tests cover existing restrictions, new actions, connection removal, and other gateway callers. - Webhook receipt grants no tool access. The routine agent still needs an authorized Fireflies connection. - New generic triggers rely on provider event subscriptions. Without a sender-supplied idempotency key, changed request bytes count as a new event. Existing legacy Fireflies triggers retain summary-only filtering and per-meeting deduplication. ## Model Used OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing, code execution, and embedded-browser testing. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b8ec588f3 |
Stop duplicating wake context in adapter environments (#13891)
## Thinking Path
> - Paperclip manages agent work and preserves task context.
> - Built-in adapters already include wake context in the agent prompt.
> - They also copy the full wake JSON into a process environment
variable.
> - A large environment entry can prevent the agent from starting with
`spawn E2BIG`.
> - This change removes the duplicate environment entry and uses the
existing prompt delivery.
> - The agent keeps its context without extra file transport or new
history limits.
## Linked Issues or Issue Description
Refs #13144, #13860, #13872, #13793.
Large wake payloads can exceed the operating system limit for one
environment entry. The launch-envelope fix in #13793 handles the outer
transport but leaves that child environment entry intact.
Credit to @nickyleach for the prompt-only approach in #13144. This PR
applies that part on current master. It does not include that PR's
30-item history limits or recovery-history endpoint. Those behavior
changes can be reviewed separately from the process launch fix.
## What Changed
- Stop exporting `PAPERCLIP_WAKE_PAYLOAD_JSON` in the shared ACP engine
and all ten built-in adapter writers.
- Ignore configured values of the retired variable so saved adapter
settings cannot restore the oversized entry. Also drop inherited copies
in Hermes, which builds its environment directly.
- Keep scalar runtime variables, existing prompt rendering, continuation
history, resume deltas, gateway bodies, and Hermes JSON template
variables.
- Document the prompt delivery contract and the migration for custom
instructions that read the retired variable.
- Test large local and sandbox child-process launches, fresh and resumed
ACP turns, SDK delivery, and configured-variable filtering.
## Verification
- `pnpm -r typecheck` passed.
- Focused adapter utility, ACP, Codex child-process, and Cursor Cloud
suites: 340 tests passed.
- Hermes execution and prompt tests: 18 tests passed using its package
Vitest configuration.
- The child-process tests deliver over 128 KB of context through stdin
and check the complete text. The ACP test retains 50 complete messages
and 50 completed actions, then checks the resumed delta.
- `pnpm build` passed.
- `pnpm test:run` was attempted, then stopped after it reproduced ten
macOS runtime-skill-cache permission failures (also reproduced on
unchanged master) and one HTTPS backfill test failure. The HTTPS test
passed when rerun unchanged on this branch and master. The complete
local suite was not completed; Linux CI provides the full-suite gate.
- CI is green on
|
||
|
|
f55759942b |
fix(runner): accept compatible Codex releases from 0.149.0 (#13853)
Keep the minimum fixed at 0.149.0 until a deliberate maintainer change. Accept stable versions below 0.157.0 separately from the reproducible install pin. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b41ccf097f |
fix(apps): configure MCP aggregators from inline task cards (#13879)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agents request connections through cards in task threads. > - MCP aggregators need provider-specific URLs and authentication. > - The task dialog used the generic setup form and omitted these fields. > - This pull request uses the same provider setup controller in tasks and Apps. > - Users can configure a connection without leaving the task. ## Linked Issues or Issue Description Related: #13755, #13855. **What happened?** An inline Executor request opened a very wide dialog with an empty credential step. Connect failed because the MCP URL was missing. The other aggregator cards also bypassed their provider setup. **Expected behavior** Each card shows its provider instructions, URL field, and authentication options in a bounded dialog. Completing setup grants access only to the requesting agent. **Steps to reproduce** 1. Enable experimental MCP aggregators. 2. Have an agent request Zapier, Arcade, Composio, or Executor from a task. 3. Open the card and continue past Access. ## What Changed - Route page and task setup through the same provider controller. - Bound the task dialog width and preserve the requesting agent's access scope. - Support existing accounts, saved drafts, URL/token setup, and task-bound OAuth. - Keep a sign-in link available when the browser cannot open a popup. Verify completion through the existing durable callback path. - Add inline Access, configuration, and narrow Storybooks for all four providers. - Document the shared setup requirement and correct Executor's URL instructions. ## Verification - Focused Vitest selection: 24 passed. Covers all four inline forms, requester access, existing accounts, saved drafts, OAuth retry, callback validation, popup cleanup, and generic reconnect endpoint preservation. - UI typecheck, UI build, design-token gates, and Storybook build passed. - Live local browser: new Executor, Arcade, and Composio connections completed provider consent from task cards. Each appeared Connected with the requester selected. - Real Test calls returned Executor output `4`, an Arcade public GitHub star count, and Composio tool-discovery results. An ungranted agent was denied access. Real Paperclip process-agent runs discovered each provider catalog with only the requester’s connection installed. - Deployed implementation commit `5d722e89c` to the isolated staging tenant. The original failing Executor card now completes, discovers seven actions, limits access to its requesting agent, and resumes that agent. Its continuation completed real Executor calls and the provider resume flow, then returned an upstream Airtable authorization link. A real staging Test call returned `4` in 1.5 seconds. Later PR commits add regression coverage and popup-unmount cleanup. - Zapier fresh-token browser test remains pending a provider clipboard handoff. Its URL and token flow passes focused tests. - Latest-head CI (`d14f4c73c`): 53 checks passed; optional Storybook deployment and visual regression jobs skipped. Greptile 5/5, both review threads resolved. Canceled runners and unrelated chat timeouts passed the single retry on unchanged code. - The full local test suite was not run, as requested by the maintainer. CI runs the repository gates. ## Risks - OAuth popup behavior differs by browser. The explicit sign-in link and durable server completion checks provide recovery. - Saved task drafts store only a connection ID in browser storage. Credentials remain in the existing server vault. - No database or server protocol changes. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code execution, and browser tools. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4721f55803 |
fix(apps): recover MCP OAuth setup after consent errors (#13855)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external tools. > - MCP aggregator setup can return from provider consent to a saved draft. > - The branded setup did not explain failed or cancelled authorization. > - It also offered identity changes that the server does not apply when a saved connection resumes. > - This pull request explains OAuth return outcomes and keeps the displayed identity consistent with the saved policy. > - Users can understand the outcome and retry the same connection. ## Linked Issues or Issue Description Related: #13755 introduced the MCP aggregators. #13758 retired the legacy Composio broker. #13584 proposes changes to the provider handoff window; this fix retains the current handoff behavior. **What happened?** Cancelling Composio consent returned to the setup form with no explanation. The same controller ignored failed OAuth callback outcomes. Returning to Access on a saved draft also offered personal/shared choices, although the server retains the saved identity. This could make the OAuth request disagree with that identity. **Expected behavior** Explain cancellation or failure, preserve the draft, and offer Try again. Display the retained credential identity and start OAuth with the policy returned by the server. **Steps to reproduce** 1. Enable MCP aggregators and start a Composio connection. 2. Continue to provider consent and cancel it. 3. Observe the return screen. Before this change, it showed the form without cancellation feedback. 4. Go back to Access. Before this change, the form offered personal/shared choices even though resume retains the original identity. **Paperclip version or commit** Observed before the fix on `8c6cc7dccf91523e0720bd86f95487e66b4b0e63`. The original report of successful consent leaving setup unfinished did not reproduce. This PR addresses the recovery defects observed during that investigation. **Deployment mode** Isolated local development instance and authenticated staging deployment, with real Composio consent and provider calls. ## What Changed - Read the OAuth callback outcome in branded MCP setup and show cancellation or failure feedback. - Retry the same saved draft without displaying untrusted callback error text. - Keep saved personal/shared identity fixed in Access and select OAuth identity from the returned credential policy. - Add six focused regression cases and authorization-failure Storybook states for Arcade, Composio, and Executor. - Document return-screen recovery and retained identity in the connector playbook. ## Verification - `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx -t 'OAuth return' --maxWorkers=1`: 6 passed, 143 skipped, including after integration with current master. - `pnpm --filter @paperclipai/ui exec tsc --noEmit`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed for the implementation commit. - Browser recovery: cancelled real Composio consent, observed the new feedback, returned to Access, retried the same personal draft, completed consent, and ran a real tool call. - Staging on implementation commit `84bb40aa70662e0c8955bf692c5714661b4bea93`: fresh shared and personal connections each completed on the first consent attempt and loaded 11 tools. Real discovery, execution, and schema calls succeeded. Both connections stayed Connected after reload. A real agent used the shared connection through the Paperclip gateway and returned the public repository documentation hierarchy with one success and zero errors. - All current-head CI gates passed on `662f84a67e867a52a2e5526026adbed00f6b59bf`. The Cursor execution and agent-chat browser shards each had an initial timeout; both passed on one targeted rerun without code changes. Greptile reviewed this exact head at 5/5, with no open review threads. - No full local suite was run, as requested. CI provides the broader checks. The PR adds a master merge and documentation after the live-tested implementation commit. ## Risks - This shared setup controller also serves Arcade and Executor. Their callback rendering and retry behavior have focused test coverage; this investigation used Composio for live provider testing. - Saved identity remains fixed during resume. A different identity requires a new connection, consistent with server behavior. - No database, protocol, credential storage, or gateway policy changes. ## Model Used OpenAI GPT-6 via Codex, with code editing, shell tools, and browser testing. The exact runtime model ID and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7944ed3d97 |
fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Native runner agents need governed tools, durable runtime state, and useful execution evidence. > - A first activity trace waited 53.467 seconds even though tool activity took 6.274 seconds; provider input arrived before the server API call executed. > - Native agents also need a safe way to hire teammates without asking the model to rebuild runtime configuration. > - This pull request separates the observed ACP input-stream window from the actual server `tool.execute` span and adds a server-owned native hire contract. > - The benefit is clearer latency evidence and safer native teammates with existing approval, auth, and company boundaries preserved. ## Linked Issues or Issue Description Related Daytona provenance work is in [#13814](https://github.com/paperclipai/paperclip/pull/13814). No duplicate public PR was found for this combined timing and native-hire change. **What existing behavior does this improve?** Native runner agents can use governed tools and request hires. The server did not expose a safe native hire operation that reused the caller's validated runtime settings. First-activity traces also mixed provider input timing with server tool execution timing. **Current behavior** A native hire must construct a separate runner configuration. Full configuration copying could expose paths, instructions, secrets, or sessions. Timing evidence could make a provider or MCP identity join appear proven when the trace did not contain that join. **Proposed behavior** The native `hire_agent` operation accepts identity and persona inputs. The server sends `adapterType: "paperclip_runner"` with `inheritRuntimeFrom: "caller"`, then copies only validated provider, model, permission, lifecycle, and bounded execution settings. It inherits and validates the default environment, derives the managed AI binding through existing normalization, preserves approval and permissions, and creates fresh child instructions. Caller secrets, paths, prompts, and sessions are excluded. Provider events now include the optional boolean `inputUpdated`, with Rust forwarding support. Timing evidence separately records the ACP input-stream window and the actual server `tool.execute` activity. It does not claim a provider or MCP join without matching evidence. **Reason and benefit** Native agents can hire teammates that start with the caller's approved execution policy. Operators retain company boundaries, auth rules, approval gates, and requalification. Reviewers can distinguish provider streaming time from server API execution time when diagnosing first-activity delays. **Breaking changes** None for existing hires or tool calls. `inheritRuntimeFrom` is optional and only applies to same-company native agent callers. Conflicting explicit runtime settings are rejected. The provider event field is optional for existing producers. ## What Changed - Added the native `hire_agent` protocol action, catalog entry, API contract, and runner authority checks. - Added `inheritRuntimeFrom: "caller"` validation and a closed native runtime inheritance allowlist. - Preserved managed AI binding normalization, default-environment validation, approval snapshots, permissions, requalification, and fresh child instructions. - Added provider `inputUpdated` schema support and Rust forwarding. - Added first-activity and server tool timing evidence with conservative identity-join handling. - Added route, authority, provider-event, sidecar, API, catalog, and Rust-focused tests. - Kept private Honeycomb links, raw traces, and local result paths out of this description. ## Verification Focused checks passed: - 458 timing/session checks. - 61 native hire inheritance checks. - 20 hire authority checks. - 1,741 API checks. - 106 catalog checks. - 54 provider sidecar checks. - 12 Rust provider checks. Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2). R1 stopped at missing Docker image setup. The final trace is available at https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces. Latest-head CI passed all required build, typecheck, Rust, static, Vitest, serialized-server, workspace, chat, and E2E jobs. The focused local checks listed above passed; the broad local suite was not run before the live evaluation, while CI provides the full repository verification. ## Risks - Timing fields describe separate observed windows. They do not prove a provider or MCP owner without a valid trace join. - The inheritance allowlist must stay synchronized with native runner configuration fields. - Approval snapshots include resolved safe inherited settings and should be reviewed when native configuration fields change. - The focused local suite is narrower than the full repository suite; latest-head CI covers the broader repository checks. > Roadmap review: `ROADMAP.md` places this work within Paperclip's bring-your-own-agent direction. It extends existing native runner hiring and observability behavior. ## Model Used OpenAI GPT-6 (exact serving model ID is not exposed), with extended reasoning and repository tool use; GPT-5.6 Luna assisted with focused implementation and verification work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e4237c45f3 |
fix(ui): prevent organization title flicker during plugin loading (#13854)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The sidebar identifies the current organization. > - An optional plugin can replace this navigation surface. > - The built-in title appears before discovery and module loading finish. > - This pull request reserves the trigger until its owner is known. > - The organization name appears once, while failures retain built-in navigation. ## Linked Issues or Issue Description Refs #13832. Searched related pull requests and issues; no duplicate fix found. **What happened?** The organization switcher renders a provisional built-in title before an installed replacement loads. Unrelated plugin imports can also affect its loading state. **Expected behavior** Reserve the trigger with a neutral placeholder, then show the resolved navigation surface. Keep the built-in menu on failed or absent contributions. **Steps to reproduce** Install an organization-switcher contribution. Delay session, company, contribution, and module responses. Reload the page and watch the trigger through each stage. ## What Changed - Reserve the trigger through account, company selection, slot discovery, and module loading. - Distinguish failed session lookup from pending lookup so errors retain usable navigation. - Load and await only contributions matching the requested slots. Observe completion of imports started by another consumer. - Document loading behavior and add regression coverage for loading, failures, unrelated modules, and identity transitions. ## Verification - `pnpm -r typecheck` passed, including Rust checks. - `pnpm build` passed. - All 629 UI test files passed: 6,593 tests. The 42 focused UI/API/plugin tests also passed. - `pnpm check:token-gates` and `git diff --check` passed. - `pnpm test:run` was also attempted. The broad local server run was stopped after recording skill-cache/channel fixture failures outside this diff (for example, runtime skill source status `missing` instead of `available`). The original cause is not established. All latest-head Linux CI gates pass; the complete UI suite and affected local checks pass. - Desktop (1440px) and mobile (390px) Chromium checks passed with real host components, dynamic module loading, and the built Account bundle. Delayed fixture responses produced exactly two title states: empty placeholder, then the resolved name. A slow refresh preserved the title and trigger dimensions; absent/failed plugin fallback and Escape dismissal passed, with zero uncaught browser errors. This is browser component integration, not a live signed-in tenant test. ## Risks A cold load displays a neutral placeholder until discovery completes. Absent, ambiguous, failed, and invalid contributions still use the built-in menu. No migrations or authorization changes. Scoped module loading changes when an unrelated contribution is imported; each surface loads its own matching modules. ## Model Used OpenAI Codex, GPT-6, with reasoning, code execution, and browser verification. The exact deployment ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cdf04a33fa |
feat(adapters): refresh current coding models and reasoning controls (#13829)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its adapters supply model catalogs and reasoning controls to agent setup. > - Several provider releases are missing from the fallback catalogs. > - Some newer models also have effort levels that the UI does not offer. > - Operators need the exact supported IDs and controls when discovery is unavailable. > - This pull request updates the existing adapters from current provider documentation. > - Operators can select current coding models without entering custom IDs. ## Linked Issues or Issue Description **What existing behavior does this improve?** Model selection and reasoning controls across the existing coding-agent adapters. **Subsystem affected** Claude, Codex, Grok, Gemini, Cursor, Kimi, and OpenCode adapters; model discovery tests; agent creation and editing. **Current behavior** The catalogs omit Opus 5.5, GPT-6 Sol/Luna, Grok 4.7/4.6/4.5, current Gemini Flash models, and several Cursor/Kimi choices. Bedrock has obsolete IDs. The UI omits supported effort levels and saves Grok effort under a key the runtime does not read. **Proposed behavior** Offer verified current model IDs and model-specific efforts. Remove retired Gemini 2.0 choices. Keep configured defaults and saved model IDs. Keep runtime discovery for account-specific choices. **Reason and benefit** Catch up with provider releases through September 22, 2026. Correct the picker and runtime controls together. **Breaking changes** No database or API change. Gemini 2.0 options leave the picker after their June 1 shutdown. Existing saved IDs remain unchanged. Corrected Bedrock catalog IDs do not rewrite saved configuration. **Additional context** Supersedes the separate GPT-6 Sol PR #13830. Fable 5.1 was already merged in #12730, and GPT-6 Astra in #12851. The Grok 4.6/4.5 proposal #11324 was closed and parked by its author. This change retains the default-sentinel fix from #12062. Related discovery proposals #13127 and #13565 do not supply these catalog and effort updates. Searches found no open PR for the additional model IDs. See [the dated audit](https://github.com/paperclipai/paperclip/blob/feat/claude-opus-5-5/doc/adapter-model-audit-2026-09-22.md) for exact scope, primary sources, runtime observations, and account-specific limits. This updates existing adapters and does not duplicate planned core work. ## What Changed - Add Opus 5.5 for direct Claude and Bedrock, with a Claude Code 2.1.280 gate. Correct and extend Bedrock model IDs. - Add GPT-6 Sol/Luna and Fast mode. Offer Ultra for Astra/Sol and GPT-5.6 Sol/Terra, and Max for both Luna generations. - Add Grok 4.7/4.6/4.5, expose supported Extra High effort, and save Grok edits under `reasoningEffort`. - Add Gemini Flash 3.8/3.7/3.6/3.5, Flash Lite 3.5/3.1, and 3 Flash Preview. Remove retired 2.0 choices. - Add the current documented Cursor fallback models, including Fable 5.1, Composer 2.5, and Muse Spark 1.3. - Refresh OpenCode fallback IDs used in remote environments from its installed provider registry. - Add Kimi K3 256K. Update the existing coding alias to K2.8 Preview and enable its CLI effort settings. - Use model-specific Claude/Grok efforts in creation and editing. Clear unsupported effort when switching models. - Add catalog, CLI/ACP forwarding, compatibility, and UI persistence coverage. Record the audit and sources. ## Verification - Latest head `6e63c9ef53b54ba869cd4fb431a8570bebe289f4`: 53 CI checks passed, 2 skipped. This includes full workspace typecheck, build, and all test shards. Greptile is 5/5 with zero unresolved threads. GitHub reports no merge conflicts. - 340 focused tests passed across adapter metadata, CLI/ACP arguments, Claude version checks, Kimi effort, Grok execution, server model discovery, and UI effort selection/persistence. - `pnpm --filter @paperclipai/adapter-claude-local --filter @paperclipai/adapter-codex-local --filter @paperclipai/adapter-grok-local --filter @paperclipai/adapter-gemini-local --filter @paperclipai/adapter-kimi-local --filter @paperclipai/adapter-cursor-local --filter @paperclipai/adapter-opencode-local typecheck` — passed. The same filters with `build` passed. - `pnpm check:token-gates` and `git diff --check` — passed. - Full workspace and UI typechecks were attempted locally. They stop on existing missing `three` dependencies in `packages/shared/src/cliplab`. - Full `pnpm test:run` and `pnpm build` were not run locally. Worktree creation exhausted disk space, so a clean dependency install is not feasible on this host. Focused checks reuse existing dependencies. CI supplies full workspace verification. - No provider inference was run. Account-specific runtime model lists were inspected where available. - Manual check: select the new models in agent setup and editing. Confirm Luna has Max but no Ultra, Grok 4.7 has Extra High, and Fable 5.1 has Extra High/Max. Save Grok effort and confirm `adapterConfig.reasoningEffort` contains the selection. ## Risks - Catalog presence does not grant account access. Older CLIs and restricted accounts can reject a model. Opus 5.5 has an explicit upgrade check. - Higher effort can increase cost and latency. Existing agent defaults are unchanged. - Cursor fallback IDs come from public model documentation; the local account exposed no live catalog. Runtime discovery still adds account-specific variants. - Kimi effort remains supported only on its explicit CLI engine. This does not add effort support to its default ACP engine. - Saved obsolete Bedrock or retired Gemini IDs are not migrated automatically. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7badae6981 |
feat(plugins): add an optional organization switcher slot (#13832)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Plugins can add UI surfaces to the board. > - Organization navigation is still fixed in the host sidebar. > - A distribution needs a supported way to supply its own organization menu. > - This pull request adds one optional React slot with host-owned navigation controls. > - The built-in menu stays available when the optional contribution cannot render. ## Linked Issues or Issue Description **Subsystem affected** Plugin SDK, server capability validation, and sidebar UI. **Problem or motivation** An installed plugin cannot replace the organization switcher without editing the host menu. Existing sidebar and overlay slots do not provide this replacement surface. **Proposed solution** Add an `organizationSwitcher` slot that requires `ui.sidebar.register`. Pass display state, an icon renderer, and navigation/logout callbacks. Keep the built-in menu for absent, ambiguous, missing, failed, or unsupported contributions. **Roadmap alignment** This keeps distribution UI in plugins and adds a small host contract. It does not add an account system or change company authorization. The maintainer requested this extension. Related prior menu changes: #12788, #10917, and #10850. No duplicate replacement-slot PR was found. ## What Changed - Add the slot to shared validation, SDK types, and server capability checks. - Wrap both sidebar menu variants with the optional replacement. - Resolve selection against the current account query before mounting, and reset replacement state on account or company changes. - Add host-specific component props and a fallback to `PluginSlotMount`. - Document the React-only contract and its trust boundary. - Report runtime-supervisor fixture startup details when CI readiness fails. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - `pnpm check:token-gates` passed. - `pnpm exec vitest run --project @paperclipai/ui`: 6,580 tests passed. - Targeted manifest, replacement, and built-in menu tests: 26 passed, including the incoming-account selection regression. - Installed a local test contribution into an isolated server. CLI inspection reported `ready`. Browser checks covered the loaded production UI, keyboard dismissal, current-organization selection, and an expired remote session. Remote account responses were fixtures. - `pnpm test:run` was attempted. Its first server phase passed 8,323 tests but failed in 36 files due to embedded PostgreSQL startup and filesystem permission errors on this Mac. Later phases did not run. The current Linux CI run is green: 54 checks passed and two were skipped, including build, typecheck, and browser gates. See https://github.com/paperclipai/paperclip/actions/runs/35796722770. - The earlier runtime-supervisor readiness failure did not reproduce locally. The complete affected shard passed locally: 58 files and 843 tests. The six supervisor tests also passed on Node 24.21.0 with CI flags. Added fixture startup diagnostics for the selected Node executable and listener port. The affected Linux shard then passed all 843 tests. The original root cause remains unconfirmed; no production runtime behavior or timeout was changed. ## Risks - Plugin UI remains trusted same-origin code. Display props do not authorize account requests. - A replacement can change navigation behavior. The host retains the built-in menu when discovery or rendering fails and keeps logout/session cleanup host-owned. - No database migration. Existing menus and portfolio behavior remain available. - Local full-suite verification is limited by the environment failures listed above. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser testing. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a10702a878 |
feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people start and continue agent tasks from other services. > - A Slack conversation needs access to its surrounding discussion and Slack collaboration tools. > - The agent must use the linked requester's access and keep private material within its permitted audience. > - This pull request adds Slack tools through the existing connector contribution and approval framework. > - People can ask an invited bot to read a discussion, create follow-up tasks, and collaborate in Slack. ## Linked Issues or Issue Description **Subsystem affected** Chat connectors, connector runtime, tool gateway, and connection Settings/Access. **Problem or motivation** Slack-origin tasks can receive messages but cannot inspect the rest of a channel or act through the originating bot. People must paste context or configure a separate integration. **Proposed solution** Supply typed Slack tools and a bundled skill only to the originating task and assigned agent. Resolve the linked requester on the server. Check bot and requester access before reads and writes. Use existing durable actions and approvals. Retrieved messages remain source material. **Alternatives considered** Slack's user-OAuth MCP server does not replace the customer-created chat bot. An unrestricted Web API proxy would not provide suitable permission or publication boundaries. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps work with a provider contribution. It does not add a task dispatcher or a separate Slack task lifecycle. Related: #11144 covers generic per-user MCP grant execution; this change binds Slack bot operations to chat-origin tasks. ## What Changed - Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and shared native/HTTP execution. - Bind tools to company, endpoint, task, run, assigned agent, and admitted linked requester. Check membership and revocation on each call and before queued writes. - Add paginated reads, bounded history search, source links, messages, file uploads, reactions, pins, bookmarks, topics, canvases, lists, and approved channel operations. - Restrict private-source publication, including automatic replies and uploaded deliverables. Keep other people's bot DMs inaccessible. - Reuse action receipts, idempotency, approvals, and reconciliation. Suppress an identical explicit-send/final-reply duplicate. Return governed results through their verified originating conversation. - Add endpoint-bound personal search OAuth storage and lifecycle. Keep native real-time search disabled until a runtime meets Slack's transient-result requirements. Current runtimes use bounded history search. - Show capabilities, scope upgrades, and personal search authorization in Settings/Access and Storybook. Document provider and runtime limits. ## Verification - Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no unresolved findings. One unchanged rapid-callback timing test passed on a single CI retry. - Approval presentation regressions cover board-comment precedence and exact Slack publication; the expanded database assertion passed in CI. The local PostgreSQL startup probe later became unavailable, so that final assertion was verified in CI. Slack setup and failed-run retry browser tests also passed locally. - Full workspace typecheck and build passed. Server typecheck/build passed again after the approval routing fix. - Broad local suites passed in separate groups: server 12,958 tests, UI 6,555, shared 770, skills catalog 20, and other workspace packages 2,652. CLI and serialized server checks passed after environment/timeout retries. These are composite results, not one uninterrupted green full-suite invocation. - PostgreSQL authority regression covers admitted identity, cross-company/task/agent rejection, recovery, retained-session revocation, OAuth refresh/disconnect races, approval execution, exact publication lineage, retries, uncertain sends, and duplicate suppression. - Gateway/response regressions cover separate-origin approval batches and durable continuation. Focused provider, access, search, native runtime, route, and AgentMail regressions pass. - Storybook capability, missing-scope, OAuth configuration, authorization, and disconnect states were inspected in the browser. - Live staging: read a channel decision and full thread, create exactly two assigned backlog tasks, add a reaction, paginate discovery to exhaustion, and return bounded search matches with source links and coverage. - Live staging: create/edit/read a canvas and list, inspect the canvas in Slack, post/edit one message, and create a channel only after approval. New channels remain disabled for responses. - Live staging: read a response-disabled channel from the requester's DM; writes to that channel were denied. The test setting was restored. - Final live retest passed: explicit file upload and exact content read-back; approved deletion of only the disposable bot message; continuation confirmation returned to the original Slack thread without repeating the action. - Optional OAuth, private multi-user boundaries, native RTS, and CLI provider execution are not fully live-qualified. The staging agent initially supplied malformed tool arguments; valid arguments succeeded, and the tool/skill descriptions now emphasize UUID write keys. ## Risks - Existing Slack apps must add scopes and reinstall for new capabilities. Provider plans and document permissions can still restrict operations. - Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET` for governed tool actions. The staging instance was configured with explicit operator approval; fleet provisioning is a separate gap. - Native RTS is not exposed on current transcript-retaining runtimes. Bounded history scans are deliberately reported as incomplete. Inline file reads support text/canvas content up to 256 KiB; other types return metadata. - Private document edits fail closed when the full audience cannot be verified. Uncertain effects other than posts/uploads require inspection instead of blind retries. - Shared approval-delivery code now separates outcomes by source run to preserve origin boundaries. No database migration is required. - A separate completion-validator gap remains when the agent cites a prior run's registered artifact during finalization. It asked for registration again even though Slack delivery was confirmed. This change does not add a connector-specific task-completion policy. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and browser testing. The exact deployed model identifier and context-window size were not exposed in the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a959e47508 |
fix(apps): reduce Google Chat scopes and block unread filters (#13820)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents controlled access to external services. > - Google Chat uses OAuth profiles and reviewed MCP tools. > - Those profiles request membership and read-state access that the supported feature set does not need. > - Removing read-state access also requires us to block unread search filters, including on existing connections. > - This pull request reduces both OAuth methods and enforces the reduced search contract before dispatch. > - Users retain conversation lookup, message history, ordinary search, and approved message sending. ## Linked Issues or Issue Description **What happened?** Both Google Chat profiles request membership and read-state scopes. The supported tool set does not include membership listing or read-state updates. Message search still advertises an unread filter. Related work: Refs #12619. **Expected behavior** Managed and customer-owned OAuth request only the scopes needed for supported features. Unsupported unread filters fail clearly before any provider call. Existing cached catalogs and broader grants must not bypass that policy. **Steps to reproduce** Start Google Chat OAuth from either connection method and inspect the requested scopes. Inspect the message-search tool schema, then submit a search with `searchParameters.isUnread` set to true or false. **Paperclip version or commit** The scope change is based on master at `110d176fc`. **Deployment mode** Managed Cloud and self-hosted instances with Google Chat Apps enabled. ## What Changed - Remove `chat.memberships.readonly` and `chat.users.readstate.readonly` from shared profiles and all four Chat connection methods. - Hide unsupported read-state fields and instructions in agent and board Test tool schemas. - Reject explicit unread filters, including false, null, snake-case fields, and encoded filter objects, before provider dispatch. - Recheck previously approved calls and support existing profile-bound and URL-only Chat connections. - Add scope, signed broker request, OAuth URL, allowlist, schema, and dispatch regression tests. - Document coordinated app/broker rollout, existing-grant reconnects, and the remaining deployment checks. ## Verification - All seven focused OAuth and Chat gateway test files pass: 498 tests on the rebased branch. - `pnpm -r typecheck` and `pnpm build` pass, using pinned pnpm 9.15.4. - The complete sharded CI test matrix passes, including general server, Chat, workspace, serialized server, Runner, and all eight browser e2e shards. The duplicate unsharded local `pnpm test:run` was stopped after CI passed; it did not complete locally. - `git diff --check` passes. - `node scripts/ingest-app-definitions.mjs` succeeds and leaves the branch unchanged. Google Workspace JSON is the durable reviewed input used by the generator. - All current-head CI checks pass at `7ee755371714dc036fff1c7da844776fee3f2ec1`, including typecheck, build, and canary dry run. Greptile is 5/5 with zero unresolved threads. The generator concern was withdrawn after review of the source and regeneration evidence. - No production deployment or live Google consent test was performed. After coordinated deployment, verify reduced consent scopes, normal search/history, message sending, and rejection of unread filters. ## Risks - Coordinate deployment with the companion Cloud broker scope change. Mixed versions can reject exact-scope requests. - Existing tokens are not narrowed or revoked. Grants with old scopes need new consent. Do not revoke a shared Google client to migrate one profile. - Explicit unread filters now return an error instead of being sent to Google. Ordinary search and the approved send tool remain available. - No database, UI, lockfile, or workflow changes. ## Model Used OpenAI Codex, a GPT-5-based coding agent, with tool use and code execution. The exact runtime model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1ccae464c5 |
feat(ui): add task artifact media gallery and full-row links (#13825)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks collect the files and work products that agents create. > - The Artifacts tab shows these outputs in rows, which makes videos hard to compare. > - Small text links also make artifact rows harder to open. > - This pull request adds image and video tiles with previews and makes the full artifact row clickable. > - Users can compare outputs and open the existing media viewer with one click. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task Artifacts tab and media previews in task chat. **Current behavior** Video outputs appear as file rows or icons. Users must click a small link to open a work product. A task with eight video outputs gives little visual context. **Proposed behavior** Show images and videos in a responsive gallery. Show a paused video frame as the thumbnail. Keep documents, links, and other files in rows whose entire area opens the item. **Reason and benefit** Users can compare generated media without opening each item. Larger click targets also make the sidebar easier to use. **Breaking changes** None. This uses the existing artifact URLs, media viewer, run grouping, and attachment filters. No API or database changes. Related work: #11226 added the task sidebar output surface, and #7361 added rich attachment previews. #3524 concerns a separate reviewed-assets panel. This PR improves the existing task artifact components. The duplicate search found no active PR for this change. This is polish for the shipped Artifacts & Work Products roadmap item. ## What Changed - Add a shared media tile for work products and agent attachments. - Reuse video and image previews in task artifacts and chat. Seek up to one second into videos and reset preview state when the source changes. - Use the existing task gallery for playback and downloads. Preserve grouping and attachment deduplication. - Extend native links and buttons across work-product rows, including keyboard focus indicators. - Add eight offline Storybook examples for video outputs, mixed media, clickable rows, narrow and wide panels, missing previews, empty state, and light mode. - Register the component in the design guide and document its use. ## Verification - 108 focused component tests pass, including thumbnail seeking, source changes, gallery activation, and attachment deduplication. - `pnpm build`, `pnpm -r typecheck`, Storybook build, token gates, and `git diff --check` pass locally. - Reviewed the production components in the embedded browser. Checked all eight video thumbnails, mixed media, narrow layout, light mode, blank-area row clicks, keyboard gallery activation for generic-MIME images, and playback from chat video thumbnails. - Storybook: open **Tasks / Artifact Gallery** and select **Eight Video Outputs**, **Mixed Media And Files**, or **Whole Row Clickable**. The small local clips are synthetic fixtures. - All build, typecheck, unit, runner, and end-to-end CI jobs pass for `b8ace289b5e07df5b9f2c319b159f3923ce427b4`. Greptile gives 5/5 with both review findings resolved. All 54 PR checks pass, including the external security scan. The full test suite passed in CI. The duplicate serial local test run was stopped after CI finished; the 108 focused tests, full build, and recursive typecheck passed locally. ## Risks - Video thumbnails require the browser to load metadata and a frame. A slow server or unsupported codec can leave the fallback visible; opening and downloading still use the existing viewer. - Full-row click targets change pointer interaction with work-product cards. Native link and button semantics remain in place. ## Model Used OpenAI GPT-6 in Codex, with code execution and browser tools. The exact deployment ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
92d4868e79 |
fix(server): isolate run errors and redact runtime capability headers (#13826)
## Thinking Path
> - Paperclip manages AI agents and their work.
> - Operators use Sentry to investigate failed runs and server errors.
> - Run reports attach a task ID, run ID, error code, and adapter
fingerprint.
> - The server skips Sentry's OpenTelemetry setup to preserve its
separate tracing and privacy settings.
> - Without an async context manager, a scope mutation can attach old
run data to later errors.
> - The HTTP logger also retains a runtime credential capability header.
> - This change isolates run metadata and redacts that header so
diagnostics identify failures without leaking credentials.
## Linked Issues or Issue Description
Refs #13446 and #13719.
**What happened?**
After a terminal run failure, an unrelated server exception can inherit
that run's tags, context, and fingerprint. Sentry then groups a database
error with an earlier adapter failure. The real SDK reproduces this with
the application's `skipOpenTelemetrySetup: true` setting. HTTP request
logs also retain the `x-paperclip-github-capability` header, which must
be treated as a credential.
**Expected behavior**
Run metadata belongs to the terminal run event. Later exceptions must
not inherit it. Every genuine error must still be captured. Runtime
capability headers must be redacted on success and failure logs.
**Steps to reproduce**
1. Initialize the optional Sentry SDK with the application's options and
an in-memory transport.
2. Capture a terminal run failure.
3. Capture an unrelated exception.
4. Inspect the second event. Before this fix, it contains the first
run's identity and fingerprint.
5. Send a request with a fixture runtime GitHub capability header.
Before this fix, HTTP logs retain the fixture value.
## What Changed
- Pass tags, context, and fingerprint directly to `captureException`
instead of mutating the ambient scope.
- Preserve the existing run fields, grouping keys, ordinary exception
capture, and privacy settings.
- Test two run identities interleaved with unrelated exceptions against
the real optional SDK.
- Update the capture contract tests and document event-local run
metadata.
- Redact the runtime GitHub capability header through the existing HTTP
logger policy. Test successful, denied, and failed requests.
- Add a dedicated GitHub-hosted CI check that installs the exact
optional SDK version declared in `server/package.json`. It fails if the
real-SDK regression would be skipped. The SDK stays outside the
workspace and production dependency graph.
## Verification
- The real-SDK regression failed before the fix because the unrelated
event contained `contexts.run_failure`.
- Five focused suites passed: 123 tests, including all optional SDK
tests. Suites: `run-failure-sentry-real-sdk.test.ts`,
`run-failure-sentry.test.ts`, `sentry.test.ts`,
`run-failure-report.test.ts`, and `http-log-redaction.test.ts`. A custom
in-memory transport prevented outbound Sentry delivery.
- All three new header-redaction cases failed before the policy fix and
passed afterward.
- The dedicated CI command passed locally with
`PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1` and the audited SDK available
through `NODE_PATH`.
- Server TypeScript check passed with a scratch configuration that
resolves this checkout's workspace packages. The existing dependency
links point to another checkout.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- Gitleaks and a separate private-data scan passed before push.
- Full local workspace typecheck, test, and build were not run. The
machine has less than 2 GiB free and those commands include Rust builds.
Full PR CI must pass before merge.
- The dedicated real-SDK GitHub check passed with 1 test executed and no
skips: https://github.com/paperclipai/paperclip/actions/runs/35774449002
- Greptile reviewed
|
||
|
|
5f1100e3b3 |
refactor(server): remove retired operator UI snippet injection (#13789)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators can extend its interface through trusted plugin UI contributions. > - The server also accepts executable HTML through two legacy environment settings. > - This older path bypasses the plugin installation and lifecycle model. > - This pull request removes snippet injection from static and development pages. > - Operators must migrate existing integrations before upgrading. ## Linked Issues or Issue Description **What existing behavior does this improve?** Retire the operator HTML injection path from the server. Related public changes: #13168, #13245 and #13496 introduced the legacy settings; #13646 supplies the generic plugin host contract. **Current behavior** Managed instances append operator-supplied HTML or a decoded script body to every page. The same integration can use supported trusted plugin UI slots. **Proposed behavior** Ignore both retired settings. Serve normal branded static and development HTML, and keep plugin contributions unchanged. **Breaking changes** Installations that rely on `PAPERCLIP_CLOUD_UI_SNIPPET` or `PAPERCLIP_CLOUD_UI_SNIPPET_B64` must migrate before upgrading. The maintainer-owned staging and production deployments have completed the migration prerequisite. Other operators must migrate their integrations before adopting this change. ## What Changed - Remove the snippet injector and its static/dev rendering integration. - Remove injector-specific tests and retain a regression that old settings no longer change served HTML. - Replace setup instructions with a retirement and plugin-migration note. ## Verification - Rebased onto current master; `pnpm exec vitest run server/src/__tests__/static-index-html.test.ts server/src/__tests__/vite-html-renderer.test.ts`: 5 tests passed. - With repository-pinned Rust/Cargo installed, full `pnpm -r typecheck` and `pnpm build` pass after rebase. - Previous full local `pnpm test:run` encountered unrelated macOS runtime-cache rename `EACCES` errors and a missing AgentMail skill path; a focused reproduction confirmed 5 failures / 95 passes. That is not a passing full-suite result. The unchanged focused suites, build and typecheck were repeated after rebase; the full local suite was not repeated. All 54 refreshed GitHub checks/contexts passed on `bd62bff63f9d7980bfd10e54cb0102693d27ecfb`; fresh Greptile is 5/5 with no unresolved threads. - No UI component styling, database or API contract changed. - Maintainer approved the remaining rollout and cleanup. Staging snippet retirement and sleep/wake verification are complete. Production migration and removal of the legacy settings are complete for serving tenant instances. Remaining old warm inventory is excluded from new signups until configuration reconciliation completes. The maintainer authorized upgrading the remaining old deployments and clearing their pins. ## Risks - Removing the settings disables integrations that still depend on them; operators outside the completed maintainer rollout must migrate before upgrading. This is an intentional behavior change, documented at the existing setup-doc path. - Keep a previous image and its configuration for rollback. Existing browser tabs need a refresh to unload already-injected code. - Plugin UI remains trusted same-origin code. This does not add a security sandbox or change ordinary branding. ## Model Used - OpenAI GPT-6 (Codex; exact deployment variant and context-window size are not exposed in this session). Reasoning, repository inspection, local code execution and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused rendering tests; unrelated local full-suite failures are disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All refreshed checks pass on `bd62bff63f9d7980bfd10e54cb0102693d27ecfb` - [x] Fresh Greptile is 5/5 on the current head, with no unresolved findings - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3d78e3a4ec |
fix(runner): keep warm sessions alive with managed GitHub access (#13815)
## Thinking Path > - Paperclip manages AI agents and their work. > - The native Runner keeps a live provider process between task turns. > - Managed GitHub access used a token tied to one run. > - A new run forced Paperclip to replace that process to replace its token. > - This PR gives the session a stable credential transport and binds each operation to the active run. > - The agent can keep its process while Paperclip checks current identity and grants. ## Linked Issues or Issue Description Follow-up to #13738. Related credential-rotation work: #11770 and #8208 use process replacement for other adapter credentials; this change applies to managed GitHub access in the native Runner. **What happened?** A configured GitHub connection forced a warm native provider process to close at each new run. The saved conversation survived, but the live process did not. **Expected behavior** Keep the warm provider process. Resolve GitHub access for the current run when each command starts. Deny access while idle or after the run ends. **Steps to reproduce** 1. Configure managed GitHub access for a native Runner agent with a warm session. 2. Complete a turn, then send another message to the same task. 3. Observe the provider process close with the reason `warm native session configuration changed`. **Paperclip version or commit** Reproduced on master `8326e33ad`. Rebased onto `e3d8fb087` before submission. **Deployment mode** Local and remote native execution, including the sandbox callback bridge. ## What Changed - Move configured native GitHub transport and launcher ownership from the run to the provider session. - Bind the broker only after the executor acquires session ownership. Clear that binding when the run exits. - Keep the shared live-run, identity, grant, and trust-policy checks for each credential request. - Reject wrong scopes, idle requests, and credential responses that arrive after their run binding changes. - Retire transport and launcher files with the provider session. Keep anonymous commands available if bridge startup fails. - Add red/green executor tests, real subprocess and callback-bridge tests, and database checks. Update the runtime documentation. ## Verification - Before the fix, both new local and remote warm-session reuse tests failed. - After the fix, 435 targeted tests passed across the executor, broker, launcher, token, and database suites. - A real long-lived test process kept the same PID and original environment across two runs, including through the production callback bridge on local test processes. - Server typecheck and TypeScript compilation passed. - Full workspace typecheck and build passed. Server typecheck passed again after the review fix. - The fallback-logging regression failed before the fix; all 9 broker tests pass afterward. - The exact chat sidebar browser scenario passed locally. The initial CI timeout showed failed Vite module downloads; all eight browser shards pass on the latest commit. - All 53 latest-head checks passed, including the full CI test matrix and security checks (two unrelated conditional checks skipped). - The duplicate full local test run was stopped after CI passed; it is not claimed as a completed local pass. Targeted local tests, workspace typecheck/build, and the browser scenario passed. - Greptile reviewed the latest commit at 5/5 with no unresolved findings. - No fresh paid provider or Daytona campaign has run for this change. ## Risks - The broker now lives as long as the provider session. Tests cover idle denial, late cleanup, late responses, shutdown, and failed startup. - Its in-memory authority does not survive a controller restart. Existing checkpoint and process-recovery rules still apply. - Raw GitHub credentials remain confined to individual command processes. The session transport token cannot select a different task, agent, company, or run. - No database migration or public API change. ## Model Used OpenAI Codex, GPT-6, with reasoning, terminal tools, and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
110d176fc9 |
fix: preserve chat message bindings in review recovery (#13818)
## Thinking Path > - Paperclip manages AI agents and their work. > - External chat messages can start work on tasks that remain in review. > - Paperclip queues one recovery run when a task loses its review path. > - That recovery keeps the chat source but loses the admitted message IDs. > - The authorization check then rejects the recovery before execution starts. > - This change retains the message IDs so the existing check can verify current access. ## Linked Issues or Issue Description Refs #13809. **What happened?** A successful external-chat run can leave an active task in review without a maintained review path. Its automatic recovery then fails with `reviewed_chat_execution_binding_not_authorized`. The recovery context retains `chat:slack` but drops `wakeCommentIds`. **Expected behavior** An eligible recovery should retain its admitted message references and pass a new authorization check. Missing or revoked access must still prevent execution. **Steps to reproduce** 1. Finish a chat run whose task remains in review with no maintained review path. 2. Build the bounded review recovery from that run's context. 3. Dispatch the recovery through the reviewed-chat authorization check. The new regression tests fail before this patch. Related PR #13809 handles answered conversations that become idle. This patch handles recovery when the conversation remains active. ## What Changed - Retain the admitted message batch for external-chat review recovery. Derive the current comment reference from that batch. - Share the existing supported-provider selector between recovery and run-bound chat authorization. AgentMail remains on its separate email inbox path. - Keep the existing authorization check. Do not copy prior checkout, authorization, session, or prompt state. - Test Slack and Discord recovery, missing batches, revoked access, and changed task or company bindings. - Document the recovery authorization contract. ## Verification - The new tests reproduced the missing-message failure before the fix. - Six focused suites passed: 257 tests covering review recovery, reviewed-chat authorization, issue liveness, Slack lifecycle, comment-wake batching, and external-chat waits. PostgreSQL integration tests ran against a temporary local PostgreSQL database. - Focused test files: `review-path-recovery.test.ts`, `heartbeat-reviewed-chat-binding.integration.test.ts`, `heartbeat-issue-liveness-escalation.test.ts`, `slack-conversation-lifecycle.test.ts`, `heartbeat-comment-wake-batching.test.ts`, and `external-chat-wait.integration.test.ts`. - Server TypeScript check passed with scratch configuration that resolves this checkout's workspace packages. Existing dependency links point to another checkout; the default check reports stale shared-type errors. - `node scripts/check-module-boundaries.mjs` and `git diff --check` passed. - `git diff | gitleaks stdin --redact --no-banner` passed. The diff was also checked for private identifiers and user data. - [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35760757895) passed on the latest commit, including build, workspace typecheck, general and serialized tests, Rust checks, and all eight browser shards. The unchanged local-service readiness test and agent-chat page-load assertion passed when their failed shards were retried. The local-service suite also passed locally (6 tests). The first, superseded run lost a chat runner; its stuck browser job was cancelled to unblock the current run. - Full local workspace typecheck, tests, and build were not run. The machine has about 2 GiB free, and those commands include Rust builds. The full PR CI checks passed before merge. ## Risks Small context-construction change. Dispatch still checks current execution ownership and access for every admitted message. The recovery remains bounded to one attempt per consumed path. No schema changes. Existing failed runs are not retried by this patch. ## Model Used OpenAI GPT-6 (Codex), with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
8725d6ce09 |
fix: make answered Slack conversations idle (#13809)
Settle published, successful Slack turns as Idle; resume the same conversation on an admitted message. Preserve unfinished work, delivery errors, and explicit dispositions. Verified through focused lifecycle/API/UI tests, full CI, and a real staging Slack conversation in the embedded browser. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e3d8fb0876 |
feat: attest standard production images at full source commits (#13797)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators deploy its standard production container on several CPU architectures. > - Downstream image builders need to identify the exact source of their base image. > - A short commit tag does not provide signed source evidence. > - This pull request adds a full commit tag and signed image digest for canonical master pushes. > - Consumers can verify the source and compose from the immutable digest. ## Linked Issues or Issue Description **What existing behavior does this improve?** Publication of the standard multi-platform production image. **Current behavior** The Docker workflow publishes short commit tags and channel tags. It does not provide a signed standard-image contract tied to the complete master commit. **Proposed behavior** Canonical master pushes also publish `sha-<full-commit>` and attest the exact index digest after platform validation and an immutable-image orphan-reaping check. The signer certificate binds the source repository, commit, workflow and ref. Existing tags and the separate cloud producer remain available. **Reason and benefit** Downstream builders can prove the source of a standard base without adding their dependencies or repository details to the public workflow. No matching open issue or duplicate PR was found. ## What Changed - Add the canonical full-SHA tag without changing existing tag mappings. - Validate amd64 and arm64 descriptors and hash the exact registry response bytes and require its digest header to match. - Verify the immutable image and sign it with GitHub artifact attestations. - Run the contract tests in trusted PR verification and document the consumer contract. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/cloud-source-verification.test.mjs scripts/standard-image-contract.test.mjs`: 37 passed. - A read-only check against an existing published index returned its exact expected digest. - `actionlint -shellcheck='' .github/workflows/docker.yml .github/workflows/pr-trusted.yml`: passed. Normal ShellCheck reports only existing `ls` and word-splitting warnings. - `pnpm build`: passed locally with Cargo available. - `pnpm -r typecheck`: passed locally. - Full local Vitest was attempted: 8,260 passed, 14 failed, with 34 failing suites. The failures were missing embedded-PostgreSQL library aliases in this fresh install and existing macOS runtime-skill-cache rename errors. Native aliases are now restored. Rerunning the 33 affected database suites produced 577 passes and two unrelated AgentMail skill-root lookup failures (32 suites passed). The four directly failing database tests also pass independently. This is not a claim that the full local suite passed. - All final-head CI checks pass. One unrelated routine-route mock assertion passed on the single-shard retry; its 15 tests also pass locally. Greptile is 5/5 on this exact head, with no unresolved threads. - Actual signing requires a canonical master push. This draft PR does not publish trusted provenance. ## Risks The new attestation step requires OIDC and attestation write permissions in the merge job. Signing failure leaves the image available but without the new admission proof. Consumers must fail closed when proof is missing. Existing release tags, the legacy producer, and image retention remain unchanged. No database or application behavior changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, shell execution and test tools. The session does not expose a more specific model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8326e33ada |
Fix oversized sandbox process launch payloads (#13793)
## Thinking Path > - Paperclip manages agent work and preserves context across retries. > - Sandbox ACP runs encode their command and environment into one launch value. > - A long continuation can make that encoded value exceed Linux's exec limit. > - The launch shell then exits before the agent can initialize. > - This PR transfers large command envelopes through a private temporary file. > - The agent receives the complete environment and can start normally. ## Linked Issues or Issue Description Refs #13777. **Bug description** Sandbox tasks with long retry context fail during ACP initialization with exit 127. The streamed bridge also drops the shell error that explains the failure. **Steps to reproduce** Launch the streamed sandbox process bridge on Linux with one valid 110,000-byte environment value. Its base64 command envelope exceeds the limit on one exec argument or environment string. Daytona reports `argument list too long: env`, then exit 127. **Expected behavior** The command envelope must not make a valid child environment too large to launch. A shell startup failure must retain its diagnostic in the run log. ## What Changed - Keep envelopes up to 64 KiB on the existing launch path. Upload larger envelopes in bounded chunks inside a mode-0700 session directory. Set the final payload file to mode 0600. - Read the file without following symlinks and delete it before spawning the child. Remove incomplete uploads on failure. Both streamed and polled bridges use the same envelope. - Preserve stderr when the launch shell fails before the wrapper emits a terminal event. Emit a fixed terminal error and shutdown acknowledgement if a payload cannot be read or parsed, without exposing its contents. - Cover large environments, file permissions, payload deletion, interrupted uploads, missing or malformed payloads, and startup diagnostics. Document the transfer and cleanup behavior. ## Verification - Both large-envelope regressions fail before the fix and pass after it. - Targeted bridge, ACP engine, real-spawn, and stdin-race checks pass: 370 tests. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - A gated live Daytona probe on the current sandbox image reproduced exit 127 with the old bridge. The fixed bridge launched the same command successfully. A second probe completed real Claude ACP initialization with a 110,000-byte context value. It did not run an agent task. All temporary sandboxes were deleted. - `pnpm -r typecheck` and `pnpm build` reach the unchanged Rust runner step and stop because this machine has no `cargo` executable. - Full CI passes on `ea816cf58d`: [run 35683853754](https://github.com/paperclipai/paperclip/actions/runs/35683853754). All 53 checks pass; two optional checks are skipped. The local full-suite run was stopped after equivalent CI suites passed; it has no final local result. - Greptile is 5/5 on `ea816cf58d`, with no unresolved review threads. The branch is mergeable. ## Risks Large envelopes require extra upload calls during startup. The temporary data stays inside the private session directory and is removed before child startup or during failure cleanup. Individual child environment values still obey the operating system's native limits. No migration or configuration change is required. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted checks; full-workspace limits described above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c496acb570 |
Return client errors for known OAuth reconnect states (#13794)
Map missing OAuth refresh credentials and terminal reauthorization to HTTP 422. Preserve reconnect instructions and reporting of unexpected provider failures. All 363 focused tests and server typecheck pass. Required CI and review checks passed; Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a7d3b17a97 |
Explain disabled Slack MCP app access during discovery
Recognize Slack's exact disabled-app response and return actionable setup instructions from catalog and health routes. Bound response parsing and keep unknown upstream errors reportable without exposing provider settings links. Verified 361 focused tests, server typecheck, and authenticated discovery. The three route regressions fail before this change and pass afterward. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ded156a904 |
Keep interaction continuations scoped to the target task
Separate an interaction's producer from explicit resume history. Do not import another task's comments or results. Filter newly captured foreign origins and recover older inherited origins only when the saved producer context and comment row prove their source. Keep missing context and company boundaries fail-closed. Verified 46 continuation tests, 201 related recovery tests, and server typecheck. Added 20 database regressions for provenance and scope guards. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
878734a061 |
fix(tools): treat OAuth sign-in challenges as client errors (#13786)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connected apps can require OAuth sign-in before they list their tools. > - Remote discovery recognizes this condition as `oauth_challenge`. > - Discovery and catalog refresh currently return it as HTTP 502. > - The server error handler reports that response as a crash. > - This pull request returns HTTP 422 for the explicit sign-in challenge. > - The operator keeps the sign-in instructions, while unexpected upstream failures remain reportable. ## Linked Issues or Issue Description **What happened?** Connecting a remote MCP app that answers with a recognized OAuth challenge returns HTTP 502. Reading an empty catalog or explicitly refreshing it does the same. The server error handler then sends the expected sign-in condition to error monitoring. **Expected behavior** A known sign-in requirement returns HTTP 422 with the existing `oauth_challenge` code, message, and setup/reconnect links. An unexplained upstream HTTP 400 or an unavailable upstream service still returns 502 and reaches error monitoring. **Steps to reproduce** 1. Configure a remote MCP app that returns HTTP 401 with a Bearer challenge. 2. Connect the app, read its empty catalog, or request a catalog refresh. 3. Observe HTTP 502 and a server error report before this change. **Paperclip version or commit** Reproduced on `6de50ba594b15efaa3eae6ed869cd39b3a436456`. **Deployment mode** Server with remote MCP connections. The regression coverage uses local PostgreSQL and mocked upstream HTTP responses. Related: #9750 addresses MCP initialization and session recovery. It does not change the classification of this recognized sign-in condition. Targeted searches found no duplicate classification PR. ## What Changed - Return 422 for `oauth_challenge` from discovery and from catalog health-error normalization. - Preserve the existing structured error and remediation links. - Test automatic empty-catalog reads and explicit refreshes. Verify that OAuth challenges produce no Sentry capture and that upstream 400/503 failures still do. - Update the direct-connect and blocked-redirect expectations and document the monitoring behavior. ## Verification - Before the fix, three sign-in route regressions fail with 502 instead of 422; all four upstream-error controls pass. - After the fix, all 339 tool-access and error-handler tests pass, including authorization and redirect protections. - These suites ran against disposable Homebrew PostgreSQL 16.14 through the existing test-constructor seam. The temporary setup and config remain outside the repository. CI uses the ordinary embedded PostgreSQL setup. - Direct server `tsc --noEmit` passes. - Full build and recursive typecheck were attempted; the Runner Rust step cannot run because `cargo` is absent on this machine. - Full `pnpm test:run`: 8,211 passed, 14 failed, 4,759 skipped. The 36 failed files match the existing embedded PostgreSQL startup/cleanup and macOS runtime-cache `EACCES` limitations. The changed database-backed service suite passed separately with local PostgreSQL. - Greptile: 5/5 with no unresolved comments. Seven CI workers received a simultaneous shutdown signal; the failed jobs are being retried through the normal workflow. Other completed checks passed. ## Risks Low risk. Clients now receive 422 instead of 502 for the explicit `oauth_challenge` condition. The code, message, and remediation links remain available. No permissions, credential handling, OAuth discovery rules, retry policy, or schema change. Other upstream failures retain their existing behavior. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, code editing, and test execution. The session does not expose an exact model snapshot or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (339 service and error-handler tests; full workspace limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6de50ba594 |
fix(sentry): carry the deployment environment to the browser (#13784)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators can enable Sentry for the server and the signed-in browser. > - The server SDK reads `SENTRY_ENVIRONMENT` from the process environment. > - The browser receives its DSN through the session response, but receives no environment. > - A browser in staging therefore reports errors under the SDK's production default. > - This pull request passes the configured environment through the existing session and monitoring gate. > - Browser errors then identify the deployment environment while preserving the existing privacy settings. ## Linked Issues or Issue Description **What happened?** With `SENTRY_ENVIRONMENT=staging`, browser exceptions are tagged `production`. This can send errors to the wrong environment's alerts and makes deployment follow-up unreliable. **Expected behavior** The browser uses the server's configured Sentry environment. A reused image works in either staging or production. A signed-out browser still sends no events. **Steps to reproduce** 1. Configure a frontend Sentry DSN and `SENTRY_ENVIRONMENT=staging`. 2. Sign in and capture a browser exception. 3. Inspect the event environment. Before this change, it is `production`. **Paperclip version or commit** Reproduced on `a3749aac4680a901fa0fe1cc898907887abc9908` with the real browser SDK and a local test transport. **Deployment mode** Authenticated server and browser with optional Sentry monitoring enabled. No duplicate environment-attribution issue or pull request was found in the targeted GitHub search. ## What Changed - Add `sentryEnvironment` to the authenticated session response and shared schema. The optional field supports a newer browser reading an older server response. - Pass the environment to the browser SDK. An environment change restarts the client through its existing serialized lifecycle. - Cover environment attribution with a real SDK event, session authorization, unchanged-session refetches, environment changes, and legacy responses. - Document configuration and compatibility. Keep the loaded bundle's release identity and existing privacy filters. ## Verification - The regression test emits `production` for a requested staging environment before the fix. - Focused route, schema, browser lifecycle and real-SDK tests: 69 pass. - UI and shared-package typechecks, direct server `tsc --noEmit`, and token gates pass. - Full `pnpm build` and `pnpm -r typecheck` were attempted. Both stop at the Runner Rust step because `cargo` is absent on this machine. - Complete UI suite: 6,540 tests pass in 626 files. - Full `pnpm test:run`: 8,210 passed, 14 failed, 4,753 skipped; 36 files fail due to embedded PostgreSQL startup/cleanup and macOS runtime-cache `EACCES`. These match the existing local baseline; none touch the changed behavior. - Greptile: 5/5, no unresolved review threads. Linux CI has passed Build, Typecheck + Release Registry, and the completed test jobs so far. Remaining jobs are running or queued: the AWS runner provisioner is retrying EC2 CreateFleet `InternalError` responses. Full results will be recorded before merge. ## Risks Low risk. This adds one optional session field and changes Sentry attribution only. No migration or new monitoring opt-in is introduced. Missing settings keep the browser SDK default. Agent and unauthenticated requests still receive 401 without monitoring settings. Existing loaded browser bundles keep their old behavior until refreshed. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, code editing, and test execution. The session does not expose an exact model snapshot or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass (focused and full UI suites pass; full-root environment failures documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5842185e4f |
fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat needs reliable native execution before native runners become the onboarding default. > - Existing stories covered idle reassignment and controller restart, but not an executing worker handoff or worker process loss. > - Status answer tests also need to reject stale claims and invented facts. > - This pull request adds six opt-in full-stack cells with independent state assertions and retained evidence. > - The probes exposed a misleading Retry across server projection and recovery-banner paths; the fix reports the blocked recovery honestly. > - The tests preserve failures without changing recovery policy, production prompts, or onboarding defaults. ## Linked Issues or Issue Description Refs: #13762. Related: #13765 (Retry targets the latest failed attempt), #13753 (task context ownership), #13746 (native recovery work). ## What Changed - Add active reassignment with saved draft and plan preservation, old-worker cancellation, and successor completion checks. - Preserve recovery-needed projection when native cleanup fails before its coordinator exists, refuse a generic retry that would immediately fail again, and replace the recovery banner's misleading Retry with Inspect run. - Add verified local worker process loss with a required successful continuation; retain a failing qualification result when recovery is unavailable, while independently verifying the UI/API refuse doomed retries. - Add two-turn factual answer checks for current blockers, stale claims, inactive backlog work, and unknown facts. Retain prose for separate semantic review. - Add positive and negative oracle calibration and document fault isolation, cleanup, billing, and qualification limits. ## Verification - Eval TypeScript check passes. - All 442 eval support tests pass locally. The 89 focused server tests and server typecheck pass. Six recovery-banner UI tests and token gates pass. - Initial new-cell campaign: https://github.com/paperclipai/paperclip/actions/runs/35657128077. All six results are retained; four failed on fixture-contract issues and two exposed real worker cleanup quarantine. - All 26 existing native onboarding cells: https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26 passed on master |
||
|
|
846336e5a0 |
test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path > - Paperclip lets people manage agents through ongoing conversations. > - Chat users can change instructions while a provider is already working. > - Existing chat evals wait for each turn to settle before the next message. > - They cannot prove delivery during active work or the saved effect of a correction. > - Existing fixtures also enable Agent Chat through the API rather than the settings UI. > - This PR adds bounded browser workflows and checks their persisted outcomes. ## Linked Issues or Issue Description Refs #13741, #13752, #13750. **What happened?** The chat suites cover planning, delegation, status, and recovery. They lack active-turn follow-ups and the experimental settings lifecycle. A sequential conversation can pass even if messages sent during work are lost. **Expected behavior** A follow-up submitted during a provider turn survives and affects the final reply. A changed launch day appears in the saved plan. Disabling Agent Chat rejects new messages while preserving history; re-enabling resumes the same conversation. **Steps to reproduce** Run the explicit `agent-chat-stories` suite. It selects three local cases for each native Claude and Codex profile. An ordinary provider command waits for a fixture brief file so the browser can send the follow-up at an observed active-run boundary. ## What Changed - Add six opt-in Product E2E cells for settings, active follow-ups, and plan corrections. - Drive experimental settings through the UI and verify disabled sends are rejected by the public API. - Use a bounded file wait in the actual isolated agent workspace, with provider-written readiness and an undisclosed brief reference. - Grade persisted user messages, final replies, native run outcomes, and exact saved plan fields. - Accept active-turn steering or one queued successor; reject lost input, duplicate input, and stale outputs. - Allow one steered run or two sequential runs throughout the shared harness, while preserving exact counts for other cases. - Require a single marker-bearing response attributed to the final provider run. - Unload the development browser client before restarting the server, avoiding reconnect/navigation races without weakening the post-restart memory check. - Add browser regressions for restart isolation and asynchronously saved settings switches. - Document prepared-agent setup, native onboarding limits, and the separate API-tool rollout gate. ## Verification - Eval TypeScript check passed. - Eval support suite: 436 tests passed in 39 files. - New oracle calibration: six tests passed, including plausible invalid outcomes. - Browser support regressions: seven tests passed; the restart regression was observed failing before the fix. - Catalog discovery selects exactly six local native cases and leaves default paid selection unchanged. - [Consolidated existing native chat report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/): master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a browser navigation timeout across restart; the page request returned 200 and the chat rendered. - [Nine targeted restart/replay cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/) passed on `1fe2fe275`, including the original failure, across native Claude/Codex and local/Daytona; all cleanup passed. - [Initial six-story campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817) retained all six failures: asynchronous switch assertions, unavailable fixture paths, and rich-text escaping in raw command comparisons. The corrected fixtures preserve the same behavioral assertions. - [Six-story campaign v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/) on `8232773a0`: 4/6 passed (both settings cases and both Claude interruptions). Codex could not see the host-temp fixture outside its workspace; this failed before follow-up delivery was exercised. - [Four affected interruption cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/) all passed, including cleanup, on definition v3 / `ad6ac0545`. Files live inside the actual agent workspace and the observed run workspace is verified. Both providers saved Friday in the real plan with the undisclosed brief reference; follow-ups persisted while the original run was active. Together with both unchanged settings cases from v2, all six new scenario variants have passing live evidence. - Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful checks, two intentional skips, zero pending/failing checks; mergeable and clean. Fresh Greptile 5/5, zero unresolved findings. - Full typecheck, tests, build, and browser CI passed remotely. One earlier head encountered a signoff-policy browser timing failure; the final head passed that shard. - Local pnpm wrapper could not fetch its version/signature metadata in the restricted environment; local eval checks used the installed Node executables. Repo-wide validation was completed by GitHub Actions. ## Risks These are eval-only changes. The file wait is a timing fixture in the isolated agent workspace, not a production runner hook. Native Codex host-filesystem isolation stays unchanged. It has a two-minute limit and is released in `finally`. The prepared-agent settings case is not full native onboarding: the wizard currently offers legacy adapters. The disabled-entry assertion uses full document navigation, which clears the prior React Query cache; preserved history is checked through the public API and re-enabled chat. No production prompt, rollout default, adapter behavior, or credential policy changes. Active-task reassignment and worker-crash recovery remain outside these new cases. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8813a50105 |
feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path > - Paperclip manages agent work as tasks and runs. > - GitHub chat brings repository conversations into those tasks. > - A review bot needs the assigned agent, its authority, and governed provider tools. > - The existing channel connection did not supply that review workflow or a complete setup journey. > - This pull request adds GitHub App setup, account access, event prompts, task-bound review tools, and exact-commit checks. > - Operators can inspect each review through the same task, run, and activity systems. ## Linked Issues or Issue Description **Subsystem affected** GitHub chat, governed connection tools, task execution, shared/database contracts, and connector setup UI. **Problem or motivation** Operators need a GitHub review bot that runs their assigned Paperclip agent. Mentions and PR events must preserve task ownership and requester authority. Provider publication must use the bot App identity and enforce the configured permissions. **Proposed solution** Extend the existing GitHub chat connector with resumable App onboarding, linked-member and sponsored-guest access, editable event prompts, and governed review operations. Validate structured assessments on the server and compute a stable Paperclip Review check for the exact head commit. **Alternatives considered** A separate review scheduler would duplicate Paperclip execution and permissions. Reusing personal GitHub credentials would change the bot identity and credential boundary. **Roadmap alignment** This extends the existing Connected Apps and governed-tool infrastructure. The project owner requested and approved this design. Related PR #8645 imports external Codex review feedback; this change runs an assigned Paperclip agent and publishes its results through the existing chat connector. ## What Changed - Include the current Paperclip instance origin in the copied setup prompt. Storybook uses its configured Paperclip origin; callback parameters and URL credentials are excluded. - Add a Claude/Codex copy button in the real setup and Storybook opening step. Its detailed prompt asks four setup questions and guides embedded-browser setup, verification, and optional required checks. Clipboard failure exposes selectable instructions. - Add a tutorial that explains why App installation, review scheduling, and required checks are separate choices. - Add manifest registration, an existing-App path, separate installation and repository selection, repository refresh, and explicit account confirmation. - Add low-trust agent guidance, effective capability verification, member selection, and explicit restricted guests with a sponsor. - Add configurable PR events, prompts, repository overrides, rating thresholds, and separate formal-review permissions. - Give the assigned agent governed App tools to read PRs, comment, begin an assessment, submit findings, and optionally submit a formal review. - Bind review history, root PR events, and inline replies to ordinary tasks. Deduplicate deliveries/findings and reject stale publication. - Link check Details to the underlying task on the current trusted hostname, or to Reviews before task creation. - Add schema migration 0283, API contracts, production UI, and 49 interactive Storybook states. - Repair local lease recovery. Keep the Cloud Dockerfile identical to master; no provider-pack layer or runtime-default environment variable is added. - Retry only rolled-back wake-admission transactions after transient endpoint-lock contention. A deterministic held-lock regression proves one accepted wake. ## Verification - Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero diff against master. Final workspace typecheck and build passed. The new PostgreSQL migration regression passed and preserves existing relation and constraint identities after replay. - Greptile reviewed this exact head at 5/5. There are zero unresolved review threads and no merge conflicts. - All current-head checks are green: 54 passed and two conditional Storybook jobs skipped. This includes complete server/workspace test suites, build, typechecks, policy checks, Runner suites, browser suites, and security status. One timing-sensitive callback-ordering test passed in isolation and its CI shard passed one retry. The duplicate local full-suite run was stopped after CI completed; it is not counted as a local full-suite pass. - Before the final Slack rebase and migration renumbering, 186 focused GitHub tests, 14 native bootstrap cases, token gates, and Storybook build passed. The final rebase retained the new Slack communication guidance. - The embedded-browser setup test copied the full detailed prompt, including the configured Paperclip instance URL. Desktop and narrow layouts were checked. Component tests cover successful copying and clipboard failure with selectable text and retry. - Live local and hosted GitHub acceptance evidence refers to application revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks exercised issue mentions, automatic PR reviews, inline findings, repeated mentions, task continuation, and failing-to-passing checks after a push. The Storybook agent generated, built, and browser-rendered pages; missing acceptance text failed, matching text passed, and broken JSX produced an incomplete result. - Live cases also covered independently disabled push events, prompt injection, duplicate signed deliveries, rapid pushes, stale-result rejection, finding deduplication, and restart recovery. Formal reviews were denied while disabled and published only after explicit enablement. Check Details links pointed to the underlying task on the trusted hostname. - Those hosted native Claude runs used the provider-pack layer now removed from this PR. They do not prove native Claude works on the standard Cloud image. A replacement hosted native Codex run is not yet verified: the disposable QA tenant has only an Anthropic AI connection. No new staging or production deployment was made for the packaging removal. - Required-check merge enforcement could not be tested because the private disposable repository's GitHub plan rejected the rules configuration. Published success/failure/incomplete check states were verified directly. ## Risks - Latest master allocated migration 0282 to Slack. The GitHub migration is regenerated as 0283 with replay-safe table/index/constraint creation; a PostgreSQL regression verifies existing relations and constraints are preserved. Existing preview tenants remain subject to the fleet migration-history compatibility preflight; no bypass is introduced. - Migration 0283 adds company-scoped configuration, registration, review, and publication records. Existing connections retain their behavior until reviews/tools are enabled. - Signed webhooks and expiring registration state remain required. Hosted installations also need the companion narrow Cloud gateway exemptions. - Agent assessments can be incomplete or wrong. The server enforces coverage/result structure, current-head publication, rating policy, and separate formal-review permission; it does not replace code-review judgment. - No Cloud image packaging changes are included. Remote native ACPX/Claude and OpenCode retain their existing operator-supplied provider-pack prerequisite. Native Codex and Codex with managed MCP tools do not require that pack. Earlier staging deployment evidence refers to its stated revision, not this packaging-removal head. Production rollout and merging remain outside this change. ## Model Used OpenAI GPT-6 through Codex, with repository, code execution, API, and embedded-browser tools. The exact serving model ID and context-window size were not exposed by the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d9b3a5653e |
feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people use the same tasks and agent tools from external conversations. > - Agents need communication guidance that fits the conversation medium. > - That guidance belongs in the original task context, without repeated instructions on each turn. > - Connection owners also need clear settings and a consistent way to remove a connection. > - This pull request adds initial Slack guidance, optional connection instructions, and chat connection menus. > - The benefit is clearer Slack replies with the existing Paperclip workflow and permissions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent replies in Slack and chat connection management in the Apps catalog. **Current behavior** Slack tasks do not carry a saved communication profile. The catalog shows a separate Manage button and does not offer removal on every chat connection row. **Proposed behavior** Save Slack guidance when a new conversation creates a task. Restore that original guidance when a model session is rebuilt. Do not append it to ordinary follow-ups. Expose optional additional instructions in Slack Settings. Put Manage and Remove connection in a three-dot menu for all chat providers. Keep Finish setup visible for drafts. **Reason and benefit** Small answers fit in Slack. Substantial deliverables use ordinary document or artifact tools with a useful Slack summary. Connection settings apply to new tasks and cannot change permissions. Users can remove both active and unfinished chat connections from the catalog. **Breaking changes** Two additive database columns store endpoint preferences and the initial conversation snapshot. Existing endpoints default to empty preferences. Existing conversations keep their original behavior. Non-Slack guidance is unchanged. Related public context: https://github.com/paperclipai/paperclip/pull/13741 improves native chat recovery. This change adds communication context to those existing execution paths. A search found no duplicate communication-guidance PR. ## What Changed - Add a provider-guidance registry, enabled for Slack first. - Persist optional endpoint communication instructions and capture an immutable snapshot when a conversation creates a task. - Resolve guidance from the verified company-scoped connection. Restore it for fresh native and legacy sessions without per-turn reminders, extra model calls, or extra context queries. - Add the Slack Settings field, validation, audit coverage, and Storybook save/error states. - Add Manage and Remove connection menus for all seven chat providers. Keep the draft setup button. Require removal confirmation and allow retry after failure. - Add regression coverage, an active/draft menu story, and connector documentation. ## Verification All CI checks are green for |
||
|
|
b82661b561 |
refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path > - Paperclip manages agents and their access to external tools. > - Connectors expose these tools through a governed MCP gateway. > - PR #13755 added a direct Composio MCP connection behind the experimental MCP aggregators flag. > - The old project API-key broker still created toolkit child connections and showed a separate Services tab. > - Keeping both paths leaves obsolete setup and session code in the product. > - This change removes the broker and preserves direct MCP setup, credentials, permissions, and execution. > - Saved legacy records fail closed and remain available for explicit removal. ## Linked Issues or Issue Description Related: #13755. This retirement supersedes the legacy-path fixes proposed in #12630, #12632, #12634, and #12906. It does not close those PRs. **What existing behavior does this improve?** Composio connector setup, management, and runtime dispatch. **Current behavior** Composio offers both direct MCP and a project API-key broker. The broker mints sessions and creates one child connection per toolkit. **Proposed behavior** Offer only direct MCP. Remove the toolkit Services UI, REST routes, API client, and session broker. Block saved legacy parent and child records from discovery, execution, health checks, reconnect, and OAuth. Preserve their records and credentials until the operator removes each connection. **Reason and benefit** The direct MCP connector becomes the single supported Composio workflow. Provider accounts remain managed in Composio. ## What Changed - Remove the API-key catalog method and its generated-source definition. - Delete Composio broker clients, session creation, account synchronization, child lifecycle, and toolkit routes. - Remove the Services tab, service rows, child provenance, and cascade-removal controls. Keep Vercel provenance intact. - Retain a shared retirement guard for stored legacy records. Show Retired status and replacement/removal guidance in the connection list and details; hide obsolete runtime controls. - Preserve the experimental MCP aggregators flag and direct MCP infrastructure. - Replace broker fixtures with retirement tests and extend direct Composio catalog/reconnect coverage. ## Verification - Focused shared, server, and UI tests passed with one worker. Server retirement tests use a name filter; no full local test suite was run, as requested. - Server and UI TypeScript checks passed. - Token gates and UI build passed. - Real browser: opened the saved Composio connection, refreshed all 11 tools, and ran the provider's read-only GitHub account-list operation through the standard Test dialog as an agent. The provider returned success using the existing OAuth credentials. - See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live evidence. - Storybook build passed. A fresh real agent used `COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway audit records confirm both calls succeeded. - Browser retirement check: a credential-free legacy fixture showed the guidance, opened the direct MCP replacement flow, and was removed through the standard confirmation. - Focused regressions for the experimental settings copy and exact OpenAPI route coverage passed. All latest-head CI checks passed (54 successful, two intentionally skipped); Greptile scored 5/5 with no unresolved review threads. The PR has no merge conflicts. ## Risks This intentionally breaks the old Composio project API-key and child-connection workflow. Existing legacy records cannot run, even if their stored status is active. Operators must create a new direct MCP connection and choose access rules; credentials and grants are not migrated. Remove each old record separately to delete its credentials. No schema migration or data deletion runs automatically. Direct MCP connections keep their existing grants and secrets. ## Model Used OpenAI GPT-6 via Codex, with reasoning, code execution, and browser tools. The exact runtime variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e8c8ba3c19 |
feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its tool gateway applies company access rules and approval controls to connected apps. > - MCP aggregators expose many apps through one provider endpoint. > - Each aggregator needs its own credential, catalog, grants, and lifecycle in Paperclip. > - This pull request adds independent Zapier, Arcade, Composio Connect, and Executor setup with a common Access → Connect layout. > - A default-off MCP aggregators flag lets operators opt in while we complete provider acceptance tests. > - Agents use the normal Paperclip permissions, Test screen, and gateway after setup. ## Linked Issues or Issue Description **Subsystem affected** Apps, connection setup, shared contracts, and the remote MCP gateway. **Problem or motivation** Aggregator endpoints need clear provider setup and correct MCP sessions. Generic setup does not explain each provider's authentication or broad execution tools. Provider approval must preserve the original execution instead of replaying a write. **Proposed solution** Add four separate connectors behind Settings → Experimental → MCP aggregators. Start with human and agent access, then connect the endpoint and read its tools. Enable tools by default. Use the existing Permissions and Test screens after setup. Keep legacy Composio API-key and child connections intact. **Alternatives considered** A shared connection for all providers would mix credentials and access rules. Separate provider-specific permission and test screens would duplicate existing controls. Vercel Connect is outside this change. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps capability and the Connected Apps roadmap area. This work was requested and reviewed by the maintainer. Related work: #11894, #12630, #12632, #12634, and #12906 concern the legacy Composio broker. #13102 also covers remote MCP pagination. This change preserves the broker path and adds initialized sessions, response matching, and provider resume handling alongside pagination. ## What Changed - Add branded setup and interactive Storybooks for Zapier, Arcade, Composio Connect, and Executor. Use the existing access controls and normal action tests. Do not request a connection name or action choices during setup. - Add the default-off `enableMcpAggregators` flag to settings, managed feature metadata, the catalog, and setup guards. Hidden connections keep running. Legacy Composio connections remain unchanged. - Reuse the vault, grants, policy, and catalog models. Support OAuth discovery, bearer tokens, custom headers, and credential-bearing URLs. Add no database tables or migrations. - Initialize and retain Streamable HTTP sessions by connection and effective credentials. Read paginated catalogs and match streaming responses to request IDs. - Classify unfamiliar aggregator tools as writes despite upstream read-only hints; only exact reviewed read capabilities enter the read-only allowlist. Legacy Composio child behavior is preserved. - Preserve provider authorization links and execution IDs. Support Executor approve/resume, decline, and cancel without automatic replay of uncertain writes. - Preserve Off and Ask first choices during refresh and reconnect. Allow new tools and retire removed tools. Keep agent access updates atomic and preserve an empty agent selection. - Document connector UX rules, provider branding sources, and live acceptance results. - Stabilize the existing Sentry release fixture after its repeated CI failure by reusing one module mock; production Sentry behavior is unchanged. ## Verification - Final head `d11781970`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/35633534900) passed, including broad typecheck, test shards, build, and E2E. All 54 checks pass; 2 optional checks are skipped. Greptile is 5/5, Security Scan passes, and all review threads are resolved. - Passed 27 focused connector Vitest checks and 18 connector-only Storybook browser checks before the flag change. All 85 stories rendered at desktop and narrow widths. - Passed 5 connector lifecycle/server checks and 7 selected flag checks after adding the flag. The latter cover settings, managed defaults, cached catalog visibility, and all four setup routes. - Review fixes passed 13 risk/handoff/lifecycle checks, dedicated session-expiration and transport regressions, 13 selected connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A real Composio connection-list call also succeeded through the refreshed UI on `9ab115f71`. - UI and server TypeScript checks passed. UI build, Storybook build, token gates, and diff whitespace checks passed during implementation. - Real browser and real Paperclip agent tests passed for Arcade, Composio, and Executor. Tested action permissions, denied agent access, reconnect, disconnect, and isolation. Tested Arcade catalog additions/removal and Executor provider approve/resume, decline, and cancel. - Zapier live acceptance is incomplete. Its dedicated provider server is configured, but its credential-copy dialog returned an empty clipboard through browser automation. No live Zapier action is claimed. - The three isolated Sentry release cases pass after the CI fixture fix. - Local verification is deliberately narrow at the maintainer's request. The full local suite, recursive typecheck, and repository-wide build were not run. CI provides the broader checks. ## Risks - Shared MCP transport changes affect other remote MCP servers. Protocol fixtures cover initialized sessions, streaming response matching, pagination, and isolation. - Broad execution tools remain broad permissions. The provider governs actions inside those tools. - Provider handoff links are retained briefly in memory. After a server restart, a one-time link may require reopening the provider dashboard. Paperclip does not replay the original call. - Zapier remains unproven live. Custom-header imports and self-hosted endpoints have fixture coverage rather than a separate live account for every variant. - Turning the experimental flag off hides setup; it does not revoke existing credentials or stop existing connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, shell execution, and browser automation. The exact runtime model ID and context-window size are not exposed in this session. A separate Anthropic-backed Paperclip agent performed live gateway acceptance tasks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9d19f98b50 |
fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent chat uses native runner sessions to plan, delegate, and track that work. > - A user can press Stop while the native session is still starting. > - The server can acknowledge that Stop without dispatching it, then let the session submit a turn. > - This leaves chat recovery waiting for an execution that the user expected to stop. > - This PR waits for the startup handle, dispatches cancellation, and prevents a late startup from submitting a turn. > - New full-stack evals check the resulting records and outputs across Claude and Codex. > - Those evals also exposed missing ACPX readiness fields, unbounded polling, and an old-run identity check that rejected valid warm handoffs. ## Linked Issues or Issue Description **What happened?** Stop during native startup could record an acknowledged cancellation with `dispatched: false`. The provider could then begin work. A subsequent `/new` stayed queued. A remote Claude follow-up also exhausted the command journal while probing warm-session readiness: ACPX never returned the readiness fields required by the shared transport. Once readiness worked, attachment incorrectly compared the next run descriptor against the old run ID. The 25 ms polling loop could issue 4,800 commands during its two-minute wait, beyond the 500-command bound. The existing chat eval treated lifecycle logs as proof of an active provider turn, so it did not distinguish startup cancellation from active-turn cancellation. **Expected behavior** A Stop during startup must reach the pending session. A late session must not submit a prompt after Stop. Recovery must retain control when startup exceeds the bounded wait. Chat evals must check saved task state, document contents, worker identity, account binding, and duplicate effects. **Steps to reproduce** 1. Start a native Claude or Codex chat turn. 2. Press Stop after process startup is requested but before the provider turn starts. 3. Send `/new`, then send a fresh message. 4. On the affected base, cancellation can be acknowledged without dispatch and the reset stays queued. **Paperclip version or commit** The live Claude baseline reproduced this on `29d6b3509`. The branch also includes master commit `0f5fafe16`. Related work: #13678, #13686, #13693, #13291, #13738. A separate runner reliability branch also contains a startup-wait fix. Its overlap must be reconciled before merging; this branch additionally prevents prompt submission after a late startup. ## What Changed - Wait for a pending native startup before acknowledging a run-scoped Stop. Preserve the existing recovery error when that wait expires. - Keep a Stop guard on startup. Cancel a late handle before it can submit a provider turn. - Add regression tests for normal handle publication and publication after the Stop deadline. - Back off blocked warm-attachment probes. Keep the fast two-snapshot barrier, fail closed, and record changed blockers. - Add red/green tests for delayed readiness, persistent blockers, alternating readiness, and readiness near the deadline. - Publish ACPX readiness and blockers. Preserve the old authority’s event acknowledgement barrier; only settled sessions can proceed to attachment. - Bind warm ACPX descriptors to the validated next authority while retaining old-run event correlation until activation. Preserve session identity and provider profile checks. - Exercise two consecutive run rotations through a qualified fake sidecar, verifying checkpointing, provider identity, pre-activation rejection, and new-run work admission. - Separate startup and active-turn cancellation checkpoints in the browser eval. - Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells across Claude and Codex. - Cover hiring and reuse through managed AI accounts, source-based review, current blocked-task status, request replay after a lost HTTP acknowledgement, server restart continuity, and Stop/reset continuity. - Use ordinary production agent instructions. Enable API tools only for the two coordination cases that need them. - Calibrate the matchers with invalid records and outputs. Require remembered context after restart and a structured status snapshot that distinguishes the current blocker from history and task status from active execution. Compare the public issue mutation contract and relationships during read-only reporting. Preserve before/after source records in failed eval evidence. - Fix the lost-ack browser harness and verify it against a real HTTP server. Check the chat composer after restart instead of waiting for an unrelated document lifecycle event. - Document the scope and limits of each case. ## Verification - The startup regression failed on the unfixed executor and passed after the fix. - `pnpm test:e2e:runner:typecheck` passed. - `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files. - `pnpm exec vitest run server/src/services/native-runtime/native-session-executor.test.ts` passed: 385 tests. - [Baseline live campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868): Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed. Codex delegation was blocked by provider capacity. - [Eval-only startup campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786): both providers failed as expected. Both persisted `dispatched: false` and left `/new` queued. - [First fixed campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706) on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both providers. Failed cases exposed eval harness defects and remote continuity failures. All attempts remain available. - [Original workflows and stronger memory checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649) on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and startup Stop passed for both providers; Codex remote restart passed. Claude remote restart exposed the missing readiness contract. Two Codex planning cells hit provider capacity. - [Unchanged-model retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548): Codex planning and backlog creation both passed. - [18-cell campaign with ACPX readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963) on `6a98ef743`: 16/18 passed, including all local/remote Stop and committed-send cases. Claude remote continuity exposed the next-authority check, now fixed. Codex hiring produced its checklist, but the runner redacted the requested marker after it appeared as “Tracking token: …”. That content-redaction policy is unchanged and remains an explicit limitation. - [Structured status grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725) on `50448c228`: both providers passed on their first attempt, including cleanup. - [Complete read-only state grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011) on `551e13892`: both providers passed. - [Final ACPX handoff and hiring retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456) on `cbd637587`: all three Claude Daytona cases passed (restart continuity, active Stop/reset, and lost-ack replay). Codex hiring reproduced the content-redaction failure: the saved checklist contained `Tracking token: [REDACTED]` instead of the required business marker. All four cases completed cleanup successfully. [Published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/). The only subsequent commit adds the qualified-sidecar integration test; production code is identical to this live proof. - `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without paid models. - Runner TypeScript typecheck passed. All 5 warm-readiness tests pass; two failed with the prior fixed-rate loop, and the late-readiness test failed before the pacing correction. - ACPX readiness and warm-identity regressions each failed before their fixes. All 292 runner-core Rust library tests passed. The qualified-sidecar integration test passes. Rust formatting is checked. - Status-grader regressions for misleading historical mentions and previously unchecked mutations each failed before tightening the oracle and pass now. - [Latest-head CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307) passed on `a4093c8f1`: full build, type checks, test partitions, browser E2E, and native runner checks. Two unrelated tests initially failed (Sentry fixture release attribution and local-service fixture readiness); both passed locally together (35 passed, 5 optional SDK tests skipped) and on the failed-job retry. No changes were made to those tests. - Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed and all review threads are resolved. - The paid live suite is not fully green: the reproducible content-redaction case remains red. This is separate from the passing PR merge checks. No production content-redaction, prompt, model, or completion-policy change is included. - Managed-account hiring and review cases explicitly enable API tools; these do not qualify default new-user onboarding. ## Risks - Stop can wait up to 30 seconds for startup, then use the existing pending-recovery path. This does not prove that remote cleanup has finished. - Blocked warm readiness adds up to 750 ms between later probes with the two-minute remote budget, or about 32 ms with the default five-second budget. Ready sessions retain the short second barrier. - Paid evals can fail because of provider capacity or agent decisions. Each failure needs evidence-based classification. - The HTTP request replay case checks comment idempotency and duplicate effects. It does not prove replay safety for an ambiguous provider tool call. - The new suite is opt-in. It does not increase the default paid campaign. - No production prompts or model selection change. Review-handoff behavior and content-redaction policy remain separate product decisions. The latter can remove harmless business content that looks like credential syntax; the failing attempt is retained. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
57fd8b70d2 |
feat: add agent avatar download to Slack setup and settings (#13740)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Slack connections let a team talk to those agents in Slack. > - Agents now have a saved avatar, but Slack setup did not offer that image. > - A matching avatar helps a team recognize its agent. > - This pull request adds an optional avatar step and a download in connector Settings. > - Users download a PNG and upload it directly in Slack with clear instructions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Slack connector onboarding and its Settings page. **Current behavior** Setup does not offer the assigned agent's avatar or explain how to upload it in Slack. **Proposed behavior** After Slack connection verification, users can download a 512 × 512 PNG of their agent's saved avatar. They can upload it in Slack, confirm, or skip. Settings keeps the download and upload instructions available after onboarding. **Reason and benefit** The same avatar helps people recognize the agent across Paperclip and Slack. Users who skip the optional step can return to it in Settings. **Breaking changes** None. No schema, authentication, Slack scope, or provider API change. Completed connections keep their existing completion state. Searched existing Slack avatar and Cliptoon PRs; no matching implementation was found. ## What Changed - Add an optional avatar step before personal Slack account linking. Keep the numbered sidebar and shared footer. - Resolve the selected agent's saved appearance for the preview and PNG download. - Add the same download and expandable upload instructions to connector Settings. - Remember uploaded or skipped per company and endpoint in browser storage. Treat uploaded as user confirmation, not provider verification. - Reject failed or non-PNG download responses and allow retry. - Reuse the production avatar components in onboarding and Settings stories. - Test wizard progression, resume, Settings, download recovery, storage isolation, and terminated assigned agents. - Exercise real PNG downloads in the Slack browser flow and keep default app names consistent with app creation. - Fetch the assigned agent directly so its saved avatar remains available after termination. ## Verification - Focused chat suites: 48 passed; the two affected suites passed again after the final naming fix (32 tests). - Slack browser E2E passed through setup, avatar download, account linking, and Settings download. PNG signature and 512 × 512 dimensions verified. - UI token gates passed. - Browser: downloaded the real 512 × 512 PNG; checked confirmation, return, mobile layout, and Settings instructions. - Full workspace typecheck, application build, and production Storybook build passed. - All latest-head CI checks passed (54 passed, 2 skipped), including all browser, chat, general, and serialized test groups. The unrelated Sentry test failed once and passed on the single CI rerun; its suite also passed locally. - Local full-suite attempt encountered a rapid Slack callback ordering failure under concurrent build load; that test passed in isolation, and all three chat shards passed in CI. The remaining local run was not used as the merge gate. - Review the Connections / Slack / Add avatar and Avatar in Settings stories. ## Risks - Slack upload is manual. Confirmation does not claim to verify the Slack icon. - Optional step progress is browser-local. Clearing storage or changing browsers can show it again. Setup still works when storage is unavailable. - The existing avatar API remains the image source. Download failures show a retry message. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c65fc9e3c8 |
fix: recover authentication and browser connection failures (#13724)
fix: recover authentication and browser connection failures Include connect timeouts in the bounded retry policy for idempotent actor synchronization. Handle WebSocket constructor failures through existing reconnect paths and preserve HTTP polling while realtime is unavailable. Refresh visible company queries until the socket recovers and clear all fallback timers on hiding or unmount. Verify 172 focused tests, server/UI typechecks, UI build, and design token gates. Full workspace build/typecheck require the unavailable Rust toolchain; the full test run is tracked separately. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9f30eb10dd |
fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path > - Paperclip manages AI agents and keeps their work attached to tasks. > - Chat connectors carry user messages and agent replies between a provider and those tasks. > - Each extra startup and context reset delays a reply. > - Managed account metadata was lost during adapter decoding, so compatible follow-ups started fresh. > - This branch fixes the reset and measures the remaining preparation, execution, and delivery costs. > - The changes must preserve account isolation, authorization, durable output, and recovery ownership. ## Linked Issues or Issue Description Refs #13699. The related service lifecycle work in #13410 and #13408 is separate; this branch focuses on task-bound chat response latency. **What happened?** Managed AI follow-ups started new provider sessions even after their configuration fingerprint stayed stable. The Codex codec removes unknown fields. The resume check then read the removed credential identity and treated it as a credential change. **Expected behavior** Compatible follow-ups resume the correct provider session. Changes to credentials, responsible users, permissions, or task configuration retain their reset behavior. **Steps to reproduce** 1. Use a Slack connector with a managed AI connection. 2. Send a message, then send a same-thread follow-up. 3. Inspect the configuration reset reason and the provider session identity. **Paperclip version or commit** Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`. **Deployment mode** Cloud staging with a native Codex runner. ## What Changed - Read saved credential identity before adapter decoding discards it. - Remove the internal credential identity from adapter-facing session params. - Test the real Codex codec and missing, changed, or unmanaged identity cases. - Preserve configured warm Codex runners and flush refreshed credentials after every turn. - Fence detached or closing session handles from successor credential ownership. - Stage current Codex launch credentials after restoring durable session history, uploading launch assets only once. - Reuse a runner binary already in the retained sandbox only when its SHA-256 matches the controller-owned artifact; still verify required capabilities before launch. - Lock the task before the run when saving results, preventing deadlocks with task updates. - Scope reusable projectless sandboxes to the company, environment, task, agent, and runtime configuration; verify Daytona sentinels for that scope. - Admit a new authorized chat message after a fully committed failed run and verified process cleanup. - Send compact deltas for verified plain-text Slack continuations. Match the actual prior run and current comment identity/body; exclude edited historical comments and prior agent output, preserve genuine brief edits and the full bootstrap fallback. - Keep attachments, omitted input, questions, approvals, recovery, and other providers on their existing framing. - Document managed session compatibility, credential lifecycle, and compact continuation boundaries. ## Verification - Workspace/session coverage: 156 tests passed. - Native session and credential ownership coverage: 390 tests passed, including exact artifact reuse, mismatches, failed probes, timeouts, and explicit artifact overrides. - Explicit continuation and durable chat authorization coverage: 172 tests passed. - Session resume and launch preparation coverage: 416 tests passed. - Result persistence coverage: 15 tests passed. The new concurrency test reproduced a PostgreSQL deadlock before the lock-order fix. - Environment lifecycle coverage: 92 tests passed, including projectless reuse and task/agent isolation at both selection and atomic handoff. - Daytona plugin coverage: 237 tests passed; 6 gated tests skipped. Standalone plugin build passed. - Compact Slack continuation and native resume coverage: 69 tests passed, including full-bootstrap retention, matching message authors/bodies, current-delivery selection, rejection of duplicate identities and historical comments, brief edits, and attachment/recovery fallbacks. - Final frozen-head `pnpm test:run` on repository-supported Node 26: 668 suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped. The sole failure was a local `socket hang up` in `issue-recovery-actions.test.ts`, not an authorization assertion mismatch. All 57 tests in that suite passed three fresh reruns, and the suite passed latest-head CI. The full local invocation is therefore not claimed green. - An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup failure; that 11-test catalog suite passes on Node 26 and in CI. No test behavior or timeout was relaxed. - Full local typecheck and build passed. Latest-head CI is green; Greptile is 5/5 with no unresolved review threads. - Two real Slack baseline replies took 25.1 and 24.6 seconds (24.9-second mean). Three same-thread signed probes on this head took 23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample and a modest wall-clock improvement, not a large or statistically established speedup. - In that same thread, uncached provider input fell from 8,514 tokens before compact input to 694–765 tokens afterward. The current delivery uses a 362-character delta; the full 19–21k-character bootstrap remains available for failed resume. Verified runner artifact preparation fell from about 1.2 seconds to 0.6 seconds. - A fresh thread created a separate task, sandbox, and provider session with full bootstrap (24.4 seconds). Its follow-up reused its own sandbox/session and compact input (28.3 seconds, including 16 seconds of model execution). Model variability and process startup remain substantial. - A signed duplicate webhook produced exactly one user comment, one successful run, and one final Slack reply. Slack's API independently confirmed the actual replies and a public task URL without an internal or pool hostname. - Earlier signed probes verified recovery after a failed run and reuse across a server deployment. The final idle test observed Daytona report the sandbox as stopped, then delivered a new reply in 19.9 seconds using the same sandbox/provider-session identity and compact input. Slack’s API confirmed that reply. - Live probes use signed synthetic inbound webhooks and real outbound Slack delivery, read back through Slack’s API. The final browser recheck found the Mac locked and the Slack tab blocked by another extension, so this is not claimed as full UI E2E proof. - This is a review branch. Do not merge until the maintainer reviews it. ## Risks - Incorrect session reuse could mix account or task context. Missing or changed identities continue to reset, and existing authorization checks remain in place. - Warm mode remains opt-in. Remote warm mode requires a reusable sandbox lease. Retained processes keep credentials until they close, so idle expiry and ownership fences are required. - A fresh user message may continue after a committed provider failure. Approval, current authorization, process termination, and prior-result checks remain required. - Projectless sandbox reuse is task- and agent-scoped. Missing or mismatched ownership cannot replace an existing lease; existing workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox per task/agent, so provider auto-stop and deletion policies still determine idle compute and storage costs. Fleet defaults are unchanged. - Compact prompts apply only after proven resume and a matching prior-run delta. Missing or specialized context falls back to full input; fresh sessions always receive the full bootstrap. - No schema or migration changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, tool use, and test execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
600e552d7b |
fix: attribute Sentry errors to the loaded source release (#13719)
Attribute optional server and browser Sentry events to their source build. Use validated build commits for Docker and source/npm artifacts, preserve explicit server release overrides, and keep cached browser bundles tied to the commit they loaded. Verify 127 focused tests, server/UI typechecks, Docker and source build stamps, all 53 CI checks, and Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2a99de80ec |
fix: supply public task links and preserve managed AI sessions (#13699)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors deliver agent replies to external conversations.
> - Agents need a public task link when a user asks to open the task.
> - The prompt and task tools lacked that link, so an agent could invent
an internal address.
> - Managed AI credential directories also changed the session
fingerprint on each run.
> - This change supplies public task URLs and excludes only those
temporary directory values from the fingerprint.
> - Follow-up messages can reuse compatible sessions while real
configuration changes still reset them.
## Linked Issues or Issue Description
Refs #13680 and #13694 for the related Cloud-origin fixes. No duplicate
open PR was found.
**What happened?**
An external chat reply could contain an invented internal task URL. The
publication filter then removed the link. Follow-up runs also lost their
saved provider session because each managed credential home used a
different temporary path.
**Expected behavior**
Agents receive the current public board URL for a task. Temporary
credential directories do not reset an otherwise compatible session.
Account, credential, model, permission, and custom environment changes
still invalidate it.
**Steps to reproduce**
1. Use a chat connector with a managed AI connection.
2. Ask for the current task link.
3. Send a follow-up message with the same agent configuration.
4. Inspect the task URL and the session reset reason.
**Paperclip version or commit**
Reproduced on the source at
|
||
|
|
b193077582 |
fix: report Slack callback health correctly behind Cloud proxies (#13694)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack connections turn messages into governed agent runs and return replies to Slack. > - Remote runs must replace an incompatible sandbox runner with the controller's packaged binary. > - The fallback used a package-relative path that does not match the vendored server layout. > - Slack callback health also compared the internal proxy address with the public callback address. > - Master now contains the runner fallback fix; this pull request fixes callback health and extends missing-runner regression coverage. ## Linked Issues or Issue Description Refs #13677, #13680, #13691, and #13686. **What happened?** A packaged controller stopped a Slack-triggered remote run with `runner_remote_artifact_unavailable` when the sandbox runner needed replacement. Working Slack callbacks also showed a stale URL warning behind the Cloud gateway. **Expected behavior** The controller stages its packaged runner when needed. Callback health uses the observed public address and still detects real address changes. **Steps to reproduce** 1. Run a packaged server with a sandbox that has an older runner or no runner. 2. Send a Slack mention to an agent that uses that sandbox. 3. Route signed Slack callbacks through a claimed Cloud gateway that rewrites the upstream host. 4. Check the run and the Slack callback health panel. **Paperclip version or commit** Reproduced on master at `aeef493f4a7603b7b1254421b80fb00212982390`. The fix branch also includes #13691. **Deployment mode** Packaged server with a Cloud gateway and a Daytona sandbox. ## What Changed - Extend the controller-owned runner fallback tests with a missing sandbox binary case; retain the fix now merged in #13686. - Prefer dedicated gateway diagnostic headers that survive provider rewrites of standard forwarded headers. Use validated host hints only for callback-health evidence on claimed Cloud instances after provider acceptance. Preserve request bodies, routing, authentication, and configured callback URLs. - Cover current, stale, and missing sandbox runners, all Slack callback surfaces, rejected callbacks, malformed proxy hints, real host and port changes, and the existing self-hosted behavior. - Document the artifact lookup and callback-health boundaries. ## Verification - Native session executor and binary resolver suites: 379 tests passed after merging current master (`45c99a0d0`). - Targeted callback integration suite with disposable PostgreSQL: 4 tests passed before rebase. - Full `pnpm -r typecheck` and `pnpm build` passed after merging current master. The full local suite passed 12,668 tests; one suite failed to start its disposable PostgreSQL. Rerunning that suite alone passed all 31 tests. - The built server resolver selected the executable under `server/dist/vendor/paperclip-runner/bin/`. - Final callback regression: all 4 targeted integration tests pass, covering provider header rewrites, default ports, uppercase/trailing-dot hosts, and ignored self-hosted hints. - Live staging proof of the runner fix: a previously failed Slack thread recovered, a new mention received its requested response, and the account-connect command succeeded. Unsigned callbacks returned 401. - Live browser and Slack acceptance passed: generated callback URLs, account linking, a new mention, an interactive question and answer, and all three callback-health indicators. The connector was activated through the onboarding UI. - All 54 latest-head checks pass, two optional checks are skipped, and Greptile is 5/5 with no unresolved threads. ## Risks - Runner fallback must select a binary for the remote platform. This preserves explicit remote artifact overrides and the existing capability checks. - Proxy headers are not identity proof. They are used only for diagnostics on claimed Cloud instances after the provider accepts the request. Self-hosted instances ignore them. Wrong public hosts and ports still warn. - No database migration or new public API contract. ## Model Used OpenAI GPT-6 via Codex. Used reasoning, repository tools, code execution, and browser/native-app testing. Exact model variant and context window were not exposed by the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |