mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
144083fd481464f4a328f0d184a82263de63228d
635
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
144083fd48 |
fix(interactions): wait for workspace readiness before enabling approval (#14893)
## Thinking Path > - Paperclip lets people manage AI agents and review their work. > - Task confirmations must use the work produced by their source run. > - The server blocks approval while that run still needs to sync its workspace. > - The card currently enables approval before that check can pass, so an ordinary click produces an error. > - This PR exposes the existing readiness check and shows “Preparing approval…” with acceptance disabled. > - The card refreshes itself and enables approval when the source workspace settles. ## Linked Issues or Issue Description **What happened?** A confirmation appears while its source run is still preparing or syncing its workspace. Its enabled approval button returns a conflict asking the user to retry after syncing. **Steps to reproduce** 1. Run an agent in an isolated workspace. 2. Have it create a confirmation before workspace finalization completes. 3. Click the approval button while the source workspace is still active. **Expected behavior** The card explains that approval is preparing. Acceptance becomes available automatically when the same server check permits it. Reject and revise remain available. Related work: #10770 handles this conflict after a click with retries. #9520 proposes changing the workspace acceptance barrier. This PR preserves that barrier and exposes readiness before the click, including in compact task chat. It preserves the terminal-finalize behavior from #10099. ## What Changed - Add an optional, read-only `acceptanceBlocker` to interaction responses. Readiness uses the existing source-run workspace predicate, with one check per pending source run. - Disable acceptance and show a shared preparation notice in classic and compact confirmation cards, including checkbox and secret-binding confirmations. - Refresh preparing cards every two seconds in task detail, attention, pipelines, and Skill Studio. Restore each surface's previous polling cadence when preparation clears. - Preserve live tool reviews, questions, rejection, revision, and the server acceptance barrier. No automatic acceptance occurs. - Document the preparation state and cover readiness, terminal sync outcomes, unrelated runs, historical cards, and automatic refresh. ## Verification - Focused service, card, query-refresh, and helper tests: 217 passed across five files. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm build-storybook`: passed. - `pnpm check:token-gates`: passed. - `pnpm test:run`: incomplete locally. Stopped after about 17 minutes once it reproduced seven existing environment failures: two Slack tests and two email tests lack ancestor-directory skill fixtures; three company-skills tests fail on macOS runtime-cache staging permissions. These are outside this change. The full CI test matrix passed. - CI: all 53 checks passed; two optional Storybook jobs were skipped. The branch has no conflicts with `master`. - Greptile: 5/5 on commit `8a6f216d8d`, with no review threads. - Reviewed added lines and new files for credentials, private URLs, internal task references, user paths, and run artifacts. None found. ## Risks - No database migration or change to acceptance authorization. Readiness is advisory; the server still enforces its existing gate at acceptance. - An open preparing card adds a read every two seconds. This cadence stops after readiness clears; historical cards add no workspace checks. - Failed or stale finalization retains the existing server behavior. This PR does not change recovery policy. ## Model Used OpenAI GPT-6 through Codex. The exact serving model ID and context-window size are not exposed in this session. Used reasoning, repository inspection, tool execution, and automated tests. No sub-agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (217 focused tests; full local-suite limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d034ba7491 |
fix(interactions): derive question storage from canonical forms (#14946)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents request human input through durable issue interactions. > - A question form has a canonical presentation and a compatibility storage format. > - The creation API required agents to write both formats. > - Tool guidance told agents to split text and choice questions across those formats. > - This pull request accepts one complete canonical form and derives storage fields on the server. > - The benefit is a complete question card with stable answer and retry behavior. ## Linked Issues or Issue Description Related work: Refs #13630 and #14430. PR #13630 addresses the display of historical partial forms. This change fixes creation and keeps the check that rejects conflicting new forms. **What happened?** A question save supplied three compatibility questions and one canonical text question. The API correctly rejected the incomplete canonical form. The Runner's tool description encouraged this split. Sending only a complete canonical form also failed because the API required compatibility questions. **Expected behavior** An agent sends one complete `payload.questionSet` with every text and choice question. Paperclip derives `payload.questions` for storage and answer compatibility. Existing legacy requests remain valid. Explicitly conflicting dual forms remain invalid. **Steps to reproduce** 1. Call `paperclip_request_human_input` with `interactionKind: "questions"`. 2. Send `payload: { version: 1, questionSet: ... }` with a required text question and a required choice question. 3. The old API rejects the missing compatibility questions. With this change, it stores both questions and preserves the canonical form. 4. Retry with the same idempotency key. Confirm that only one interaction exists. 5. Submit both answers. Confirm that the normal resolver and continuation rules apply. **Paperclip version or commit** The branch is based on `cf8ad63c8`. The problem affects the native Runner and the interaction creation API. **Deployment mode** Server deployment with the native Paperclip Runner. Integration tests use the real interaction service and an embedded test database. ## What Changed - Add one shared canonical-to-storage projection. Reuse it for native harness question requests. - Accept canonical-only question creation at the shared validator and server boundary. - Export the input type and update the plugin SDK and its RPC contract. - Advertise a typed, complete question form in the live and scenario tool schemas. - Enforce canonical text and custom-answer constraints before ordinary or native resolution. Preserve harmless display whitespace. - Run regex matching in isolated workers with a deadline and resource limits. Both answer paths await the result before persistence. Saved native delivery uses the validated answer without taking another worker slot. - Update agent guidance and generated Runner contracts. - Test mixed forms, option-ID collisions, retries, answers, legacy requests, and conflicting forms. ## Verification - Interaction service, HTTP route, native bridge, and Runner authority suites: 221 tests passed after correcting an obsolete tool-description assertion. - Shared validator, plugin SDK, CLI, and UI compatibility suites: 67 tests passed. - Runner core tool-contract suite: 20 tests passed. AJV validates live and scenario schemas. - Final review regressions: 172 shared, service, native bridge, and authority tests passed. These cover text length, pattern, numeric limits, whitespace, custom option IDs, and historical pending cards. - Runner session suites: 67 tests passed. Published example tests: 4 tests passed. - Server typecheck and the shared/server builds passed after the compatibility fixes. - Final delivery verification: 35 response-delivery tests passed. The native delivery regression proves saved answers do not enter pattern workers; server typecheck and build passed. - Pattern security and answer-flow verification: 205 tests passed after repairing the child fixture loader. These cover pathological matching, event-loop responsiveness, worker concurrency, slot cleanup, HTTP routes, native delivery, and the full helper in a child process. - `pnpm -r typecheck` passed on the bounded-worker revision. - `pnpm build` passed on the bounded-worker revision. - All 55 GitHub checks passed on `fe457af`; four optional jobs were skipped. An unchanged Cursor adapter test timed out once in CI, passed locally, and passed on one failed-job rerun. - Reviewers can send the canonical-only mixed form above and verify that the saved interaction contains both canonical and compatibility questions. ## Risks - The creation API accepts a new input shape. Stored rows and answer contracts keep the existing shape. - The shared projection must preserve synthetic free-text option IDs. Collision and native round-trip tests cover this behavior. - Historical partial rows remain readable. New conflicting dual forms, including written-answer mismatches, remain rejected. - Existing pending cards retain the written-answer paths offered by their stored options. Canonical text constraints still apply. - Ordinary answers now enforce declared canonical constraints before persistence. Invalid answers leave the card pending. - Regex validation has a one-second deadline and a four-worker capacity limit. A complex pattern or capacity error leaves the card pending with a validation error. - No database migration or change to company authorization is required. ## Model Used - OpenAI GPT-6 through Codex. The session exposes the GPT-6 model family; its exact runtime model identifier and context window size are not exposed. Used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7a52dcdc74 |
fix: repair MCP validation and cancelled execution recovery (#14951)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The tool gateway gives agents access to connected services. Recovery controls what happens when a run stops. > - Generated tool names can exceed the provider limit after the MCP client adds its prefix. > - The same invalid definition can fail each automatic retry. A cancelled run can also hold saved messages without showing its cause. > - This pull request bounds tool names, stops configuration retries, and retains cancellation evidence. > - It shows the stopped run and admits saved input only after the existing safety checks pass. > - The benefit is a clear recovery path that preserves operator Stop and prevents duplicate message delivery. ## Linked Issues or Issue Description **What happened?** A long connected MCP tool name makes the provider reject the entire request. Automatic recovery repeats the invalid request. Separately, unexpected legacy cancellations can leave saved input behind a recovery hold. The notice does not identify the stopped run or its cause. **Expected behavior** Complete MCP names fit the provider limit. Tool-definition errors require configuration repair. Cancelled runs retain their source and reason. The recovery notice shows the cause and saved-message count. Verified unexpected cancellations can start a fresh turn through the existing admission checks. **Steps to reproduce** 1. Assign an App gallery connection with a long application key and tool name to a Claude agent. 2. Start a run. The provider rejects a name over 128 characters, including its MCP prefix. 3. For cancellation recovery, stop a legacy provider turn without an operator Stop request and send a user message while the recovery hold is active. 4. Inspect the recovery notice and the deferred message queue. **Paperclip version or commit** Rebased onto master at `cf8ad63c806685bfd7c48e3ed4a919d61a7c55f1`. **Deployment mode** Hosted or self-hosted server with legacy Claude or Codex execution. Related public work: - Refs #14017. That PR caps name segments. This PR preserves existing short names and uses stable hash aliases for long complete names. It also covers classification and recovery. - Refs #4510. That PR adds a cancellation-source column. This PR records bounded evidence in the existing run result, without a migration. - Refs #12552 and #4506. Those PRs suppress recovery after operator cancellation. This PR preserves operator intent and uses the existing continuation gates. ## What Changed - Bound gateway names with the full provider prefix in the 128-character budget. Retain the original upstream tool name for dispatch and permissions. - Classify invalid tool definitions as configuration failures before diagnostic redaction. Stop automatic retries and continuation attempts for that error code. - Persist cancellation source, expectedness, initiator, reason, and time. Preserve recorded Stop intent when adapter results arrive. Report unexpected started cancellations with closed diagnostic labels. - Show the run cause, saved-message count, and Inspect run link. Offer Continue for eligible unexpected cancellations. Require verified provider stop, empty tool inventory, ownership, and the existing pause, budget, approval, and dependency gates. Use the existing queue for single delivery. - Add regression coverage and update the execution, MCP gateway, and run-log documentation. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - `pnpm check:token-gates` passed. - Ran `pnpm test:run` and completed its workspace and serialized groups. Initial resource and timing failures passed on isolated reruns. All 149 serialized route suites passed. - Reran the changed server, adapter, and UI suites after the rebase. Coverage includes long-name upstream dispatch, configuration retry suppression, cancellation evidence retention, privacy labels, oversized run projection, and concurrent saved-message delivery. - `pnpm test:e2e tests/e2e/legacy-failure-continuation.spec.ts` passed all six browser scenarios. The recovery notice shows the run cause and inspection link, and each recovery entry point reaches one new response. - Added database-backed checks for active, removed, paused, unavailable, and disabled chat connections. The final continuation and recovery-notice suites passed 167 tests. Externally bound chats hide board Continue and show a usable next action. - All 55 GitHub checks passed on `42afbf1371dcaeb72646e3d8f65c19ff7cddf8de`. Two unrelated Storybook jobs were skipped by their normal conditions. Greptile reviewed that commit at 5/5 with no findings and no open review threads. ## Risks - Long tool names change to aliases. Existing short names stay compatible. The original connection and upstream name remain the dispatch authority. - Invalid tool definitions no longer get automatic retries. An operator must repair the configuration before a new attempt. - Continuation changes apply only to positively identified unexpected legacy cancellations with complete empty tool inventory. Operator Stop, unknown historical cancellations, outstanding tools, and unverified provider termination keep their holds. - No database migration. The added projection fields are optional. Cancellation reason and initiator IDs remain local run evidence; Sentry receives only closed source and initiator-type labels and expectedness. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository editing, shell execution, and GitHub tool use. The runtime does not expose the exact model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7d59de6113 |
feat(connections): probe provider usage limits on demand (#14936)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections store the AI accounts used by legacy and native runners. > - Subscription accounts can reach session, weekly, model, or paid usage limits. > - Operators need to read these limits for a specific stored account before making a routing decision. > - This pull request adds an on-demand usage probe to the connection service and account detail. > - The result preserves provider limits, reset times, paid usage, and unknown values for later consumers. ## Linked Issues or Issue Description **Subsystem affected** Shared contracts, the connection service and API, and the account detail UI. **Problem or motivation** Managed AI accounts lack a common operation to read their current usage limits. A local harness probe can read a different login from the account selected for an agent. **Proposed solution** Add `aiConnectionService.probeUsage()` and a board-only connection usage endpoint. Probe the selected credential grant on request. Support Codex, Claude, and Grok subscriptions, plus OpenRouter API key limits. **Alternatives considered** Harness-specific automatic polling would couple the read to execution and can read ambient credentials. This change uses the managed connection credential and leaves scheduling and admission decisions to later work. **Roadmap alignment** This extends the existing Personal & Shared AI Accounts capability. It adds no routing or quota enforcement. Related: Refs #14459 for managed OpenAI quota reads; Refs #14781 and Refs #13379 for downstream pacing and budget work. This operation reads one requested account across all three subscription providers. ## What Changed - Add typed usage snapshots and a probe capability flag to managed AI connections. - Normalize Codex, Claude, Grok, and OpenRouter responses. Keep model scopes, provider admission, reset periods, and paid allowances separate. Preserve unknown values. - Enforce company membership, credential audience, grant identity, and connection lifecycle before reading the stored secret. - Add a board-only `GET /api/companies/:companyId/ai-connections/:connectionId/usage` endpoint with `no-store` responses. - Add manual **Check usage** and **Refresh** actions to account details. Show compact usage bars, resets, admission and overage status; remove repeated descriptions and account-default copy. Clear previous results during a new request or error. - Add Storybook previews using the production account components for all four providers, initial checks, loading, and permission errors. - Add provider, authorization, runner selection, API, and UI coverage. Document provider sources and live qualification. ## Verification - Initial provider, authorization, selection, API, and UI validation passed (96 focused tests): `pnpm exec vitest run server/src/services/ai-connection-usage.test.ts server/src/__tests__/ai-connections.test.ts ui/src/components/ai-connections/AiConnectionUsagePanel.test.tsx server/src/__tests__/openapi-routes.test.ts`. - `pnpm -r typecheck` passes for the initial implementation. After simplifying the UI, 9 usage-panel and date-helper tests, UI typecheck, token gates, and Storybook build pass. The initial feature module boundary check also passed. - Real Codex, Claude, and Grok credentials were saved to encrypted disposable connections. The actual usage HTTP route returned 200 with `status: ok`. Legacy and native runner selection checks passed. The tests started no model turn and exchanged no refresh token. The disposable databases and vaults were removed. - Live Claude responses added structured scoped limits. Live Grok responses omitted included-plan usage. Tests now cover both shapes and preserve the Grok omission as unknown. - The full workspace build passes. A full local test run hit a heartbeat feedback timeout. That case passes in isolation. The duplicate local run was stopped after all remote checks passed. The Slack ordering and OpenCode transport CI flakes also pass in isolation and on the CI rerun. - Current head: `ff3d479029a1c4248190323e221b2803cfb0d79d`. All 54 active checks pass. Two Storybook checks are intentionally skipped by the workflow. Greptile is 5/5 with no unresolved review findings; the branch is mergeable. ## Risks - Subscription usage endpoints can change. Credentials can lack usage-read permission. The probe returns explicit errors without fresh limits in these cases. - A successful probe can contain partial data. Missing utilization or admission remains unknown. An enabled paid-usage switch does not prove a funded balance. - This change adds no migration. It does not change runner admission or automatic provider selection. Provider requests use fixed endpoints, disabled redirects, bounded response sizes, and a 15-second deadline. ## Model Used OpenAI Codex, GPT-6, with reasoning, file editing, shell execution, and HTTP tools. The session does not expose the exact runtime model variant or context window size. Real provider credentials were used only for the authorized live checks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b2c565038b |
test(shared): make the worktree port registry lock suite deterministic (#12798)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Shared worktree services use lock leases and worker-thread heartbeats > - The lock test suite measured wall-clock timing across two threads > - Processor contention allowed a heartbeat tick to change the value during an assertion > - This pull request removes that timing race and restores a regression guard > - The benefit is a stable test suite that still detects slow heartbeats ## Linked Issues or Issue Description **What happened?** The worktree port registry lock suite failed at random under continuous-integration processor contention. The failure reported a fresh timestamp where the test expected an old timestamp. **Expected behavior** The suite must pass when the heartbeat runs at its supported interval. It must also fail when the heartbeat interval regresses. **Steps to reproduce** 1. Run `npx vitest run src/worktree-port-registry.test.ts` in `packages/shared`. 2. Repeat the run under bounded processor contention. 3. Set the heartbeat interval to 3000 ms and run the asynchronous critical-section test. **Paperclip version or commit** `a661caf74e704f7700a8b8a1e79b76ebd04e3483` **Deployment mode** Built from source. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific (core test). **Database mode** Not database-related. **Additional context** Related open pull requests are #11994, #11985, and #11922. This pull request keeps all five tests active and does not use `skip`, `skipIf`, or `todo`. ## What Changed - Build the fallback-probe lock state by hand so no live heartbeat changes the timestamp during the assertion. - Count distinct heartbeat refreshes in the asynchronous critical-section test. - Close the fake probe and settle the pending lock attempt in a `finally` block. - Keep production code unchanged. ## Verification - `npx vitest run src/worktree-port-registry.test.ts` — 5 of 5 tests pass. - `npx vitest run` — 72 files and 704 tests pass at submit time. - `npx tsc --noEmit` — exit code 0. - Ten target-file runs pass under bounded processor contention. - A 3000 ms heartbeat interval fails with `expected 2 to be greater than or equal to 3`. - An inverted cleanup assertion exits normally in 379 ms without a leaked worker. ## Risks Low risk. This pull request changes one test file. It changes test setup and assertions only. ## Model Used OpenAI GPT-5 through Codex. Exact model ID: GPT-5. The model used tool calls and code execution. The context window is not disclosed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6c1a75da49 |
feat(connections): make AgentMail a default connection with inline setup (#14772)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents access to external services. > - AgentMail needs both a saved key and an inbox assigned to the agent. > - Chat requests offered a setup link instead of an inline card and could treat a saved key as complete. > - Inbox setup also hid address conflicts behind a generic server error and a separate review step. > - This pull request makes AgentMail a default connection, adds the inline card, reduces setup to two steps, and shows conflicts beside the address. > - Shared native dropdown styles also give every caret a consistent inset. ## Linked Issues or Issue Description **What happened?** AgentMail requests in chat did not show a usable inline connection card. Manual setup required extra screens, ignored saved account keys, and could trap new-address setup in a locked inbox dropdown. Agent selectors omitted the avatar from the selected value. A taken address could produce an HTTP 403 from AgentMail and appear as an internal server error. Native dropdown arrows also touched the right edge of their fields. **Expected behavior** Make AgentMail available as a default connection. Ask for the API key inline, with a direct link to its provider page. Default human access to the company and agent access to the requesting agent. Resume the agent only after an assigned inbox is active. Manual setup should ask for an agent and email address, then finish. Address checks should run as the user types. Taken addresses should show clickable alternatives. A domain dropdown beside the name should prefer a verified custom domain. Setup should suggest authorized saved AgentMail keys and show agent avatars in the picker and selected value. **Steps to reproduce** 1. Ask an agent to connect AgentMail when it has no assigned inbox. 2. Check that an inline API-key card appears and links to the provider's API-key page. 3. Open AgentMail setup, choose an agent, and request an address that is already taken. 4. Correct the inline error, refresh, and finish setup with the same request ID. 5. Inspect native dropdown carets in light, dark, disabled, and right-to-left states. Uses the bounded provider-error parser merged in #14768. Related work: #13256 introduced AgentMail; #14725 expanded connection search. ## What Changed - Stop recurring email queries for tasks that have no email thread. Share the query between the thread provider and activity view. Keep email-task updates and invalidation-based discovery. - Make AgentMail available without the experimental chat setting. Keep the catalog, setup and management routes, agent Channels tab, task email feed, receiving worker, and agent tools available by default. Other experimental chat providers stay gated. - Make the email address and copy icon a single clickable action with the shared Copied! confirmation. Add View inbox linking directly to the matching AgentMail console inbox, with the address encoded as one URL path segment. - Reorganize inbox Settings around the copyable email address, usage instructions, and receiving status. Move reconnect credentials into a disclosure and separate the Disconnect action. Add production Settings stories for active, paused, unassigned-address, revoked, webhook, long-address, mobile, and reconnect states. Show repair controls when the inbox has an error. Keep usage instructions tied to an active inbox with an address. - Add AgentMail channel intents and an inline key field with the direct API-key URL. - Keep setup and retry state tied to the interaction. Require an active inbox for completion. Preserve company and agent access checks. - Reduce manual setup to agent selection and email selection. Put the domain dropdown beside the address and default to a verified custom domain. Preserve explicit choices across reloads. Keep receiving settings under Advanced options. - Check the initial address and edits after a 350 ms pause. Abort superseded requests and ignore stale responses. Show clickable suggestions and retain known creation conflicts across reloads. - Add a company-scoped, manager-only address check using the saved credential. Search the visible inbox list instead of fetching an uncreated inbox: live AgentMail retains negative lookups that can break subsequent access-key creation. Unlisted addresses remain unknown; creation is authoritative. - Suggest labeled saved AgentMail keys in both manual setup and the inline card. Filter by company, provider, active credential, and current-user grants on the server. Prefer an account key and preserve the selected key or an explicit new-key choice across refresh. Use verified scope metadata and bounded concurrent checks for legacy keys. Never return secret values. - Catch an inbox-only key before the email step. Allow its existing inbox only after an explicit choice. Recover old locked drafts at the key picker. Save the replacement key before retiring an empty draft, then use a new setup URL so refresh preserves the switched account; stop if cleanup fails. Preserve already allocated addresses and their original accounts. - Use the shared AgentSelect in email setup. Show the canonical agent avatar in each option and the selected value, including other consumers of the shared component. Add regression coverage for legacy and current Lucide agent-mention icon formats. - Start each catalog Add connection with a fresh setup identity. Honor Finish setup's exact draft/account/address instead of resuming an unrelated browser draft. Return Cancel and Done to Connectors and Email settings to the inbox. Group the task/thread explanation in a How it Works card. - Route AgentMail catalog removal through the email inbox control API, including unfinished drafts. Refresh both the catalog and inbox views. - Render each inbox management tab separately. Access uses the saved account grants and agent controls; Conversations and Activity use the shared persisted email feed. Activity lifecycle actions use the email API. Reconnect returns to inbox Settings. Conversation failures show a retry instead of a false empty state. Email delivery recovery stays in the task. - Map documented provider address conflicts to a field error. Preserve actionable messages for other failures. - Preserve non-secret draft fields across refresh, scoped to the requested agent. Never save API keys in browser storage. Resume partial inbox creation with the original agent, address, and request ID. - Show an already-created address with explicit retry and new-address recovery instead of locked inputs. Preserve the original inbox and resumable draft when choosing another address. Distinguish runtime-key 404 errors and log safe provider status/operation/code. - Apply final agent access once within email setup authorization for a new account whose original installs are unchanged. Preserve later permission edits and reused account installs. Support in-place retry of progress loading. - Let a failed inline setup change keys after retiring an empty draft. Persist its replacement setup identity without storing secrets. Recover a server-saved account when refresh interrupts the save response, while preserving intentional account changes. - Render the production setup in Storybook and add error, recovery, and mobile states. - Inset native select carets in shared CSS. Preserve custom icons, listboxes, keyboard behavior, and forced-color controls. - Add browser regression coverage and an AgentMail Product E2E case with persisted-state and rendered-card evidence. ## Verification - Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `git diff --check` passed after the default-availability change. - All 485 focused tests passed. These cover setup, management, catalog and route gates, connection intents, email authorization, Cursor execution, and the OpenAPI contract. All 39 email integration tests run with the experimental chat setting off. - The shared polling change passed four behavioral tests, UI typecheck and build, and token gates. - `tests/e2e/agentmail.spec.ts` passed with the actual server setting off. This full-stack browser test uses simulated provider responses. It covers catalog entry, saved keys, editable address and domain controls, creation, conflicts, retry, all management tabs, clipboard feedback, the provider link, and task email rendering. - In the live local browser, Add connection reached the editable email step with the saved account key. The verified custom domain was selected by default. Both domain choices worked. The existing inbox Settings page remained available. Both active inboxes completed new mail checks with the setting off. No new provider inbox or email message was created for this pass. - Earlier live provider acceptance covered creation on a verified custom domain, Finish connecting on the reported draft, successful mail checks after refresh, and catalog removal of disposable draft and active connections. Clicking the email address copied the exact address and showed Copied!. View inbox opened the same inbox in AgentMail’s console. No email messages were sent. - Production setup and Settings Storybook builds and interactions passed. Settings states include active, paused, unassigned, revoked, webhook, long-address, mobile, and reconnect. Receiving and revoked-access stories had zero accessibility violations. - Full local `pnpm test:run` on an earlier revision completed with 14,709 passing, 87 skipped, and four transient failures. All four failed cases passed in focused reruns without product changes. That serial full local command was not repeated after each follow-up. The latest-head full CI suite is the final test gate. - CI found an obsolete browser assertion that hid every channel when the flag was off. Updated it to keep AgentMail and the Channels surface visible while preserving the GitHub chat route gates. All 11 provider browser tests passed locally after scoping the Channels selector to the agent sidebar. Two initial local attempts stopped at temporary Postgres initialization. The passing run used a separate disposable database on the existing local Postgres server; it was removed after the test. - Updated the remaining sidebar and aggregator discovery assertions for default AgentMail availability. Ordinary task fixtures now return no email thread. All 128 sidebar/task-page tests and all 42 aggregator tests passed locally. - Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI passed, with 54 successful checks including Snyk and two intentional Storybook skips. The CI run is https://github.com/paperclipai/paperclip/actions/runs/37020833647. A fresh Greptile review scored 5/5 with no unresolved threads. Live model evaluations and inbound/outbound email delivery were not run. ## Risks - AgentMail no longer needs experimental opt-in. Setup still requires a human to connect an account and assign an inbox. Inline setup creates an inbox after a human submits a new or saved key. Company access, agent access, inbox assignment, and completion checks remain enforced. - AgentMail read APIs cannot prove global address availability. The visible-list check is bounded to 100 entries and cannot see inboxes outside the key’s scope. The UI reports this limitation, suggests alternatives without claiming they are free, and keeps final creation conflicts inline. Lookup outages show an error without preventing the authoritative creation attempt. - Native select CSS affects the whole app. Custom-icon selects and multi-row lists are excluded. Forced-color mode keeps the browser caret. - Saved-key discovery uses stored verified scope metadata and checks authorized legacy credentials concurrently within a shared three-second deadline. Provider outages mark legacy choices unavailable; users can still enter another key. Final use rechecks authorization and provider access. - No database migration or transport default change. Live connection remains the default. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The exact served model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e00d10d5d5 |
fix(connections): repair stale AI defaults from agent settings (#14916)
## Thinking Path > - Paperclip manages AI agents and controls the credentials used for their work. > - Managed AI connections resolve each responsible user's provider default. > - Agent settings created another account but kept the old default selected. > - A rejected provider test left the old account marked as connected. > - Claude ACP reported a typed login failure as a generic terminal-access error. > - This pull request repairs the selected account or selects the new login explicitly. > - Agents can save and run with the repaired credential, and failed logins request sign-in. ## Linked Issues or Issue Description - Fixes #14831. - Refs #13867. Environment failures remain separate from credential-health failures. ## What Changed - Add an agent-settings action to reconnect an unavailable personal default in place. Keep its connection, grant, default, and agent access. - State that a new account becomes the user's provider default. Select its returned grant before changing the agent binding. Keep the actual sign-in method. - Show default-update errors and allow retry without another provider login. - Show the agent-access choice. Connection managers start with company-wide access for their own tasks. Other members start with access for the current agent. - Use the server's connection-manager permission in the shared list response. This includes members with a custom management grant. - Mark credentials as needing attention after an explicit login rejection in Test or Save. This includes API-key 401 and 403 responses. Network, quota, and server failures keep the credential health unchanged. - Reuse the credential-generation check so an old failure cannot invalidate a newer reconnect. - Route Claude's typed provider `access` failure to the existing login-recovery flow. Replace its generic terminal-access fallback with a sign-in message. - Add regression tests and update the AI Connections documentation. ## Verification - Red: the UI tests failed on the missing reconnect action, unused returned grant, missing access choice, and lost default-update error. The server tests failed because rejected credentials stayed connected. The real ACP fixture returned `acpx_turn_failed` for typed login failures. - Green: 156 tests passed across the AI connection, hiring, agent field, and New Agent suites. All 37 environment-route tests passed. The Claude ACP authentication fixtures also passed. - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - The full local `pnpm test:run` passed 707 files and 14,503 tests, then exited with an agent-conversation timeout and embedded PostgreSQL startup failures in unchanged suites. The isolated conversation and migration tests passed on rerun. Later local test groups did not run after this failure. - [All CI gates passed](https://github.com/paperclipai/paperclip/actions/runs/37012669356) on commit `38513dfe2`. This includes the full test matrix, browser tests, typecheck, build, Runner checks, and canary dry run. - Greptile reviewed commit `38513dfe2` and returned 5/5 with no open findings. - The regression tests use a real embedded database and a real ACP fixture process. Live provider sign-in requires a valid account and was not run. ## Risks - Connecting a new account from agent settings changes the user's provider default. The dialog states this before sign-in. - The displayed access choice can allow all company agents to use the account for its owner's tasks. Reconnect keeps the existing access. Server permissions still control installs. - Claude's typed `access` category maps to the provider's `auth_required` signal. Tool and workspace request failures retain their existing classification. - No database migration or provider credential format changes are required. ## Model Used - OpenAI GPT-6 through Codex. The exact served model identifier and context window are not exposed in this session. Capabilities used: reasoning, repository tools, code editing, and command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4ac374103f |
fix(connections): repair Asana MCP and add shared-app sign-in (#14756)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let agents use provider tools through the permission
gateway.
> - Asana provides an official remote MCP server, but its v2 server
requires a registered MCP OAuth app.
> - Setup can discover retired v1 endpoints and send a callback that
differs from the displayed URL.
> - This pull request repairs custom app setup and adds sign-in through
Paperclip's shared app.
> - Users can choose their own app without enrolling with Paperclip
Cloud.
> - Agents can use Asana tools after the user connects their account and
sets action permissions.
## Linked Issues or Issue Description
Related: #14739 supplies the personal credential repair used by resumed
Asana setup. No duplicate Asana authentication PR was found.
**What happened?**
Asana setup failed even with a user-created app. Root discovery metadata
still points at v1. MCP v2 uses the Asana OAuth issuer and requires an
MCP app with a client secret. Local setup also displayed a localhost
callback while an Origin header could make authorization use a numeric
loopback callback.
**Expected behavior**
Sign in with Paperclip's app when its broker profile is available. Keep
custom MCP app setup available without Cloud enrollment. Use the correct
issuer, callback, client credentials, and resource throughout setup.
**Steps to reproduce**
1. Open Asana in the connection catalog.
2. Supply an Asana MCP app's client ID and secret.
3. Start OAuth on a local instance opened with a numeric loopback
address, or resume a draft that cached v1 metadata.
4. Observe the wrong discovery endpoint or callback mismatch.
**Paperclip version or commit**
Reproduced from
|
||
|
|
467125fafb |
feat(connections): one-screen connector setup with stated defaults (#14811)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents use Connections (the Apps catalog) to act in services like
Notion, GitHub, Google Workspace and Railway
> - Each connector asked the user to answer setup questions before it
went to the provider. Most of the questions already had the correct
answer selected
> - ROADMAP.md lists "simpler setup" for Apps and Connections as ongoing
work. This change continues that work
> - This pull request removes the questions that Paperclip can answer
itself. It states the defaults in one line and moves the choices behind
"Change" and onto the Permissions tab
> - The benefit is that most connectors take one click in Paperclip and
then the provider's own consent screen
## Linked Issues or Issue Description
No public issue exists. This is the description, from the enhancement
template.
**What existing behavior does this improve?**
The setup flow for tool connectors in the Apps catalog.
**Subsystem affected**
Apps and Connections: `ui/src/features/connections`,
`ui/src/pages/apps`, the `packages/shared` app definitions, and the
OAuth routes in `server/src/routes/tool-access.ts`.
**Current behavior**
Every connector opened with an Access step. The step asked who can use
the connection and which agents get it, and both answers were already
selected. 18 connectors also asked "How do you want to connect?" when
Paperclip could rank the methods. The Google apps and Postman also asked
"What should Paperclip be able to do?" before sign-in. The four gateway
connectors (Zapier, Arcade, Composio, Executor) used a separate two-step
wizard. Asana was pinned to a customer-owned OAuth app, so the user had
to register an app in Asana's developer console. The "Set all" control
on the Permissions tab changed only one action. After the user approved
access, Railway's consent page showed "you can close this window" and
did not return to Paperclip.
**Proposed behavior**
One screen per connector, with one primary button. The screen states the
defaults in one sentence, for example "Connects for everyone in your
organization, available to all agents". A "Change" link opens one
Advanced panel. When the provider's metadata allows dynamic client
registration, Paperclip registers a client itself. Connecting lands on
the Permissions tab. On that tab, "Set all" changes every action in the
group.
**Reason and benefit**
The user makes fewer decisions before the connection exists. Most
choices are easier to make after the connection, on the Permissions tab,
where a change has an immediate effect.
**Breaking changes**
None. No schema or API change. Existing connections keep their settings.
## What Changed
- **No Access step.** `ConnectionSetupFlow` no longer has the Access
step. The flow shows the resolved default above the primary button and
on the completion screen. The access controls moved into one Advanced
panel. The panel opens automatically only when a setting in it is
required.
- **A default method for every app.** The flow always picks the ranked
default method. Alternate methods are in the Advanced panel. The Google
and Postman capability choice is not asked before sign-in. The
write-capable method is the default.
- **Gateway connectors.** `RemoteMcpProductionSetup` (Zapier, Arcade,
Composio, Executor) no longer has its own Access step. Its commit path
and the main commit path use one helper, `askFirstCatalogEntryIdsFor`,
for server-suggested defaults.
- **Dynamic registration from live metadata.**
`canRegisterOAuthClientDynamically` now allows registration when the
provider advertises a registration endpoint, even if the catalog entry
lists only customer-owned clients. The Asana and Linear definitions and
catalog text match live probes. Asana issues clients for loopback
callbacks only, so a hosted deployment still needs an Asana app.
- **Connection setup states.** New
`packages/shared/src/connection-setup-state.ts` sorts each method into
`instant`, `authorize`, `paste` or `register`. The gallery card verb
("Connect" or "Add key") comes from this resolver and the instance's
ownership availability.
- **Generic MCP.** The generic path no longer asks "Does it need a key?"
first. A credential challenge from the server shows the key field.
- **Permissions tab.** Each action row shows its risk level. Each group
has a "Set all" control. The control sends one change for the whole
group. Before, each row's save started from the same render, so the
saves overwrote each other. The Zapier/Arcade/Composio/Executor setup
screen had the same defect.
- **OAuth callback interstitial.** A cross-site browser navigation to
`/api/tools/oauth/callback` gets a small same-origin "Finishing your
connection…" page. That page repeats the request, and the repeat does
the code exchange. Railway's consent page replaces itself after about
two seconds, and the code exchange plus tool discovery takes longer than
that. The interstitial uses only a meta refresh, because the OAuth code
is single-use. Requests without `Sec-Fetch-Site: cross-site` take the
old path.
- **Linear registers through its MCP server.** Linear pins the console
endpoints at `linear.app`. Pinned endpoints now replace discovery only
when the method cannot register, or when the connection has an
operator-entered client. So a Linear connection now finds the
registration endpoint at `mcp.linear.app`.
- **Own-OAuth-app recovery stays on the one-click screen.** When the
method also accepts a customer-owned client, the client fields are in
the Advanced panel. The panel opens after a failed sign-in. "Try again"
resumes the draft with the operator's client.
- **E2E specs** follow the one-screen flow. The Access-step clicks are
removed, the specs open **Change** before they pick agents, and they
expect GitHub's **Add key** verb.
- **Default permissions do not change.** New connections still allow
every action. The user can set actions to Ask first or Off on the
Permissions tab.
## Verification
- `cd ui && npx vitest run src/pages/apps src/features/connections
--no-file-parallelism`
- `cd packages/shared && npx vitest run src/app-definitions.test.ts
src/connection-setup-state.test.ts`
- `cd server && npx vitest run src/__tests__/tool-access-service.test.ts
src/__tests__/remote-mcp-connectors.test.ts`
- `pnpm check:token-gates`
- New tests:
- `PermissionsPanel.group.test.tsx` checks that "Set all" sends one
change for the whole group. It fails on the old code.
- `action-permissions.test.ts` checks the group update.
- `connection-setup-state.test.ts` checks the four setup states.
- A server test checks that a cross-site callback gets the interstitial
and does not use the OAuth state, and that the same-origin repeat
completes the connection.
- Manual check on a hosted staging deployment. GitHub, Google Drive,
Composio, Notion, PostHog and Railway each connected from one screen and
returned to the Permissions tab. On Railway, "Set all" changed all 65
write actions, and the change remained after a reload.
- Visual changes: snapshot baselines are intentionally not updated. See
the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot
verification demoted to dormant (Jul 13 2026)".
## Risks
- **Fewer confirmation clicks.** Organization-wide access is the
default, and the user does not confirm it on a separate step. This was
already the preselected answer. The flow shows the default before the
user clicks and again after the connection.
- **Google write scope.** Google apps now request the write-capable
scope by default. A narrower scope needs a new sign-in.
- **Dynamic registration from live metadata.** A provider can advertise
registration and then reject a redirect URI. Asana rejects hosted
callbacks, for example. In that case registration fails, and the
customer-owned client path remains available for recovery.
- **Callback interstitial.** The OAuth callback adds one same-origin
step for cross-site browser navigations. Browsers without `Sec-Fetch-*`
headers use the old direct path.
- Chat and bot connectors (Discord, Telegram, Microsoft Teams, iMessage)
do not change.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, used through
Claude Code with tool use (shell, file editing, browser automation) and
extended thinking. It wrote the code, the tests and this description. A
human product owner directed the work and tested it by hand.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: scotttong <squadbot000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
33f2b3a159 |
fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Connectors catalog lets people give agents tools or connect agents to conversations. > - GitHub put these two uses behind one card and an extra choice. > - People should choose the connection they need from the catalog. > - This pull request keeps GitHub for tools and adds GitHub Code Review Bot as a separate card. > - Each card opens its setup directly. Both use the existing connection code. ## Linked Issues or Issue Description **What existing behavior does this improve?** GitHub connector discovery and setup. **Current behavior** With chat connectors enabled, GitHub opens a menu that asks whether to use tools or create a bot. Saved tools and bots share the same catalog entry. **Proposed behavior** GitHub opens tool account access. GitHub Code Review Bot opens agent selection. Saved bots and drafts appear under the bot card. Chat-disabled instances show only GitHub tools. **Reason and benefit** The catalog names the two uses and removes an extra setup choice. The bot keeps the existing GitHub provider, credentials, endpoint IDs, setup steps, and runtime. **Additional context** Related work: https://github.com/paperclipai/paperclip/pull/12843 and https://github.com/paperclipai/paperclip/pull/14594 established GitHub account identity. This change preserves that tool flow. No duplicate catalog split was found. ## What Changed - Split the generated app definitions into GitHub tools and GitHub Code Review Bot. Reuse the existing GitHub logo and channel method. - Open bot setup directly, including old resume and reconnect links. - Put existing bot endpoints and drafts under the bot card. Hide duplicate internal chat applications. - Keep pasted GitHub URLs mapped to the tool connection. - Add seven Storybook states for the catalog, saved connections, disabled chat, both setup paths, mobile, and light mode. - Fix narrow-screen bot rows so the label cannot overlap status and setup actions. - Update catalog, route, browser, and API tests, plus the GitHub connector guide. ## Verification - [Hosted Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog): seven states built from this branch. The deployment passed its public-file verification. - All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`. Two optional Storybook jobs skip under their normal trigger rules; the manual Storybook deployment passes. The branch has no merge conflicts. - Greptile: 5/5 on the current head, with no review comments or unresolved threads. - `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm build-storybook` passed. The final Storybook fixture also passed UI typecheck and the hosted build. - Targeted catalog, URL matching, routing, grouping, brand, and chat UI contract tests passed. - GitHub provider browser tests: 2 passed. These cover direct tool setup and the bot setup and management lifecycle with provider responses mocked. - Embedded-browser test on an isolated local instance: opened both cards, selected an agent, saved a bot draft, and resumed the same endpoint under the bot card after a reload. - Storybook Tool Setup and Bot Setup assertions pass in the published preview. Chat Disabled assertions pass locally. Inspected mobile and light mode, including the draft-row layout and official GitHub marks. - Local full-suite limitation: `pnpm test:run` was not clean. A cross-company route assertion failed in the aggregate run and passed in isolation; a workspace-runtime test reached its 30-second hook timeout. Some isolated database reruns skipped when the embedded-PostgreSQL availability probe failed. The local aggregate was stopped after CI completed. The corresponding full CI suites pass all 360 tool-access tests and all 162 workspace-runtime tests. - No live GitHub authorization or installation was performed. The isolated instance correctly stopped at the cloud enrollment or public HTTPS prerequisites. ## Risks - Low scope: catalog presentation and routing change. There is no database migration or provider credential change. - Existing GitHub bot URLs now open bot setup directly. The tool route remains `/apps/connect?source=github`. - The bot remains behind the existing chat-connectors feature flag. Existing endpoints retain `provider: github`. - Channel applications are represented by endpoint rows. Regression tests cover legacy bot applications, tools, active bots, and drafts together. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and embedded-browser tools. The exact deployed model ID, context window size, and reasoning setting are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018993140f |
feat: let agents name prompt-only tasks (#14761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad55d0a281 |
fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use Apps through a gateway that checks identity, company access, and action policies. > - Personal pasted credentials can point to company secrets. Setup can show success while the gateway rejects every call. > - Several OAuth methods also omit the scopes needed for their supported write actions. > - This pull request gives setup, health checks, and invocation the same credential rules. Owners repair existing connections by reconnecting. > - New connections request reviewed permissions for their supported actions. Read-only choices remain available under Advanced. > - Agents can use the connections people give them, while existing consent, identity boundaries, and action restrictions remain enforced. ## Linked Issues or Issue Description Refs #14009 and #14008. This addresses the personal-credential defect. The separate GitHub organization-identity selection defect is outside this change. Related work: #13942 fixed part of new personal-key setup. #14200 independently fixes legacy personal reconnect and protects managed-agent profile credentials during removal. This PR covers that ownership invariant across key and secret-URL setup, reconnect, health, discovery, and invocation, and keeps owner reconnect as the repair path. #14059 tracks requested versus provider-asserted OAuth scopes; it remains separate work. I searched open PRs and issues for Zapier, Airtable scopes, connector writes, and personal credential failures. **What happened?** A Zapier secret URL saved through personal setup can become a company secret referenced by a user grant. Health checks bypass the gateway's ownership check, so the connection appears healthy but calls fail with `grant_credential_invalid`. Custom-header paths can also receive a duplicate `credentials.` prefix. Omitted OAuth scopes make write access depend on provider defaults. **Expected behavior** Personal invocation credentials belong to the selected user. Setup, health, and actual calls enforce the same rule. New connections request documented permissions for supported read and write actions. Existing tokens gain no permissions without provider consent. **Steps to reproduce** 1. Connect Zapier or a generic secret URL with the personal identity. 2. Allow an agent to use the connection and complete setup. 3. Invoke a tool through a run-scoped gateway. The legacy layout fails ownership validation despite successful setup. **Paperclip version or commit** The implementation started from `44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master at `94e8dec56`. **Deployment mode** Built from source. Regression tests use isolated PostgreSQL fixtures and controlled MCP transports. ## What Changed - Share credential writing, ownership validation, and canonical paths across initial setup, resume, reconnect, rotation, health, discovery, and gateway calls. Keep OAuth client-registration secrets separate from invocation credentials. - Existing personal connections with company-scoped credentials require owner reconnect with a fresh key or secret URL. Reconnect creates a correctly owned value and updates the existing grant and declarations. There is no automatic ownership backfill or new startup hook. - Preserve PostgreSQL timestamp precision when reconnect checks whether a grant changed. Previously, converting the timestamp to a JavaScript Date could reject reconnect with a false concurrent-change error. - Protect credentials used by other grants, connections, bindings, managed-agent profiles, routine triggers, or secret proposals from connection removal. - Review all 117 tool methods, including 84 OAuth methods. Record explicit scopes or documented provider-default exceptions with official evidence. Add Airtable's seven scopes, Hugging Face repository/job scopes, and other documented MCP permissions. - Prefer available write-capable methods. Put explicit read-only choices under Advanced. Explain pasted-key permissions and offer reconnect for missing OAuth consent. Preserve existing grants, policies, Google availability gates, and curated scope allowlists. - Reconnect generic secret URLs and custom headers using their stored credential fields. Refresh the catalog after setup, correct reconnect feedback and error guidance, and let Cancel exit invalid setup while Save & exit retains draft-saving behavior. - Apply ownership checks to the new GitHub repository/skill connection picker. Align the permission audit with the Google scope reductions merged on master. - Add run-scoped gateway, ownership, owner-reconnect, OAuth URL, insufficient-scope, UI, and catalog-wide regression coverage. Update the connector playbook and permission audit. ## Verification Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all CI/status gates (55 completed check runs, no failures or pending checks) and has a completed Greptile review at **5/5 with no outstanding findings**. GitHub reports the PR as mergeable/CLEAN. - **Embedded browser:** used the actual server and built UI from this worktree, a fresh isolated database, and local HTTP MCP fixtures. Completed personal bearer-key, secret-URL, and custom-header setup; reproduced the legacy ownership failure; reconnected through the owner’s form; and completed writes afterward. Read-back was verified for bearer-key and secret-URL connections. Public organization-wide setup appeared immediately in Browse without reload. Zapier URL validation/Cancel and Google’s enrollment gate were also exercised. - **Persistence and invocation:** verified user ownership, canonical `credentials.authorization` / `remote.url` / `headers.X-Api-Key` declarations, and unchanged connection/grant identity. The old company secrets retain their ownership. Separate HTTP calls through an actual run-scoped gateway session completed a write and read-back. - **Backend coverage:** the final gateway suite passes all 82 cases, including catalog Zapier and generic inline reconnect. It checks company/user isolation, canonical declarations, same-endpoint URL validation, fresh credentials, retained restrictions, and real gateway read/write execution using fixture transport. A timestamp with PostgreSQL microseconds covers the former false reconnect conflict. - **Local checks:** 368 catalog, gateway, repository, and UI tests passed before the final extra Zapier case; 49 GitHub skill access tests also passed. All three Apps browser regressions pass, including reconnect through the actual form and catalog visibility without reload. Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final patch, and token gates passed. Full tool-access service runs hit varying 15-second Google fixture timeouts; both affected cases and the updated reconnect assertion pass in isolation (3 tests). The complete test matrix passes in CI on this head. - **Verification limits:** no live provider account was available for Zapier/Airtable/OAuth consent or account-bound write proof. Public metadata and local fixtures do not establish provider consent. The original development database clone failed on a pre-existing missing `tool_connections_transport_check` constraint; browser acceptance used a fresh isolated database created by the normal CLI onboarding flow. ## Risks - Existing broken personal connections stay unusable until their owner reconnects. Health, discovery, and invocation return an actionable ownership error; startup does not rewrite credential ownership. - Scope changes affect new authorization requests. Providers may still require resource selection, account roles, paid plans, or app verification. Existing consent and action restrictions remain unchanged. - Shared credentials are retained rather than reassigned or revoked. Provider-default exceptions and unavailable live checks are documented in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`. - No new endpoint, database table, lockfile change, or CI workflow change is included. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell execution, web research, and browser tools. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d432dc7fa3 |
Add GitHub-synced skill sources (#14713)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Company skills supply instructions and files to those agents. > - GitHub imports already exist, but users cannot manage repositories as skill sources. > - Repository refresh also needs caller-authorized access and complete local packages. > - This pull request adds Sources inside Skills and reuses GitHub connections from Apps. > - Installed snapshots let agents use skills without fetching GitHub during a run. > - Manual refresh preserves skill identity and leaves failed imports on their last good version. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: skills UI, server, database, shared contracts, and runtime materialization. **Problem or motivation** Users keep skills in GitHub repositories. They need a clear way to select, import, and refresh those skills. Existing imports do not expose repository management or consistently preserve supporting files. **Proposed solution** Add company-scoped skill sources. Browse repositories from all accessible GitHub connections, or paste a public repository or branch URL. Select whole skill packages, inspect included files and reference warnings, and install complete, immutable snapshots. Refresh each source manually. **Alternatives considered** Project repository settings hide the workflow from Skills. A second GitHub connector would duplicate credentials and grants. Upstream editing and PR creation are separate work. **Roadmap alignment** This implements the Skills Manager direction in ROADMAP.md. The maintainer requested this scope and reviewed the component and full-app journey stories before implementation. Related reports: Refs #10285, Refs #10949, Refs #13464. Related work: #14356, #13656, #9268. ## What Changed - Add source and entry records, an idempotent migration, company-scoped APIs, and legacy GitHub import adoption. - Reuse current caller grants and credential refresh. Combine and deduplicate repository inventories across accessible connections. Pasted public URLs also prefer the active user’s authorized connections. Tokens stay in the Git child environment, never argv or disk. - Fetch a shallow Git snapshot at one immutable commit. Scan the full local tree, including hidden and nested folders. Read Git objects without checkout or archive transformations and enforce nested package boundaries. - Bound Git downloads to 128 MiB and three minutes. Cancel active process groups and remove incomplete downloads. Preserve cancellation and deadlines while progress drains; close stalled HTTP progress streams after 30 seconds. Reuse caller-scoped temporary snapshots for preview/import after reauthorization. - Index repository package boundaries once and cap expanded work at 1,000 packages, 10,000 files, and 100 MiB, including repeated copies of shared blobs. Bound path depth and the shared path index. Discovery keeps audited manifests without retaining all package bodies. - Resolve moving branches before fetching so unchanged discovery reuses caller-scoped snapshots. Limit active scans, scan frequency, and new downloads per caller and company; quotas apply before metadata reads and across connections, and cached scans do not consume the download quota. - Store complete versions with script content, binary bytes, and executable modes. Preserve these through copies, runtime caches, and runner packaging. - Stage downloads before publication. Use source leases, revision checks, and transactional activity records. Keep installed versions after failures, upstream deletion, deselection, and disconnect. - Add the approved import flow, Sources page, selection tree, provenance, read-only Studio behavior, and saved return from GitHub setup. - Add package manifests, commit-pinned file previews, and separate runtime requirements and reference warnings. Supporting files are included together; nested skills remain independently selectable. Preview requests reauthorize the caller and re-audit package content. - Show installed skills as compact links beneath each source. Repository titles open GitHub. Keep Refresh, Select skills, and Disconnect source in a three-dot menu. Source rows omit the branch, imported count, and refresh timestamp; action alignment and repository titles work at narrow widths. - Stream discovery metadata over an opt-in NDJSON response. Show measured Git download progress and real package/file counts, animate newly checked skills, support cancellation, and require a complete scan before selection. Keep the existing JSON API. - Retain component stories and add a separate full-app journey story group. Include fixed progress states and interactive scan, large-repository, interruption, and saving stories. - Update Skills documentation and product contracts. Suppress private GitHub skill references in telemetry. Privacy review requested for the telemetry changes. ## Verification - Local repository typecheck, full build, token gates, and Storybook build passed during this work. Focused transport, authorization, scanner, persistence, route, and UI tests pass. The final UI refinement passes all eight focused UI tests, UI typecheck/build, and token gates. The scanner resource and repeated-discovery fixes pass 132 focused scanner, transport, authorization, source-service, route, and rate-limit tests, plus server typecheck/build. Full-suite verification comes from CI; the older full local Vitest run was stopped after unrelated chat failures and a font-test failure, all of which passed in fresh focused runs. At commit `1098d5996`, all 54 active checks pass; two optional Storybook jobs are skipped. CI covers repository typecheck, build, the full test suites, browser shards, and the canary dry run. Greptile is 5/5 with no open findings; the security scan also passes. - Adversarial scanner tests verify repeated-blob byte accounting with and without declared sizes, package/file/path caps, one-time repository indexing, metadata-only discovery audits, and nested package boundaries. Additional tests cover branch movement, snapshot reuse, caller/company quotas, isolation across connections, active-lease cleanup, quota recovery, and rejection before any metadata API call. - Real Git tests verify hidden paths, exact binary bytes, executable modes, export-ignore preservation, symlink/submodule reporting, pinned commits, caller-scoped cache reuse, cancellation, cleanup, and credential isolation. Regression tests hold both download slots with permanently blocked progress callbacks, verify timeout/cancellation cleanup and retry, and exercise HTTP backpressure cancellation. Access tests cover automatic public-URL connection selection and revoked grants. Database tests verify company and grant audiences. - Live isolated browser test: the public `anthropics/skills` scan now completes and discovers all 20 skills without connecting an account. Imported canvas-design with all 83 files, opened it from Sources, and verified the installed binary-font preview/download control. Package previews also expose the complete file inventory before import. Cancelled an active Git download and retried successfully to all 20 discovered skills; the browser displayed measured download progress. The current audits reject four other packages; eligible selections remain importable. - Browser checks verify the simplified source rows at desktop and narrow widths, keyboard navigation into the actions menu, Refresh from the menu, selection, and fixture disconnect with installed skills retained. Storybook includes a menu-open checkpoint and a 320px layout. - Storybook includes receiving/preparing download checkpoints and a timed full-app import journey, plus cancellation, retry, large-repository, and saving states. Streaming tests cover split UTF-8 frames, incomplete streams, late responses, cross-company requests, HTTP errors, and JSON compatibility. - Earlier live acceptance on this PR imported `stitch-skill` with `DESIGN.md`, assigned it to an agent, disconnected its source, and ran a successful Studio test that read both installed files. An editable copy changed independently. Both Skills variants, mobile selection, and return from GitHub setup were exercised. - Private access, revoked credentials, OAuth success return, binary/script preservation, concurrent refresh, transaction rollback, version pins, and legacy adoption have automated coverage. A real private-repository OAuth grant was not created during this test. ## Risks - The migration groups recognizable legacy imports without provider calls. Their first successful refresh completes the local package snapshot. - Reference checks are advisory. They cover Markdown links and explicit relative resource paths, not arbitrary runtime dependency graphs. Preview text is capped at 64 KiB; imported bytes remain complete. - Git must be installed on the server. Shallow fetches still download the branch snapshot, including files outside selected packages. Downloads have size/time/concurrency limits. Temporary caches are bounded and caller-scoped. GitHub API quota still applies to repository metadata and the connection picker; content no longer uses per-file API requests. Failed scans retain installed content. - Sources depend on the current caller's GitHub access. A saved connection does not grant access to another person's token. - GitHub script support and immediate manual refresh are explicit maintainer-approved requirements. The operator trusts the selected repository and accepts upstream script and executable-mode changes on refresh. Static audits are not a sandbox or a guarantee of safe code; agents may later invoke installed helpers under their runtime permissions. Import and refresh do not execute scripts, hooks, package installation, or builds. Raw URL and skills.sh imports keep their prior script restrictions. - Source originals remain read-only. Refresh affects subsequent unpinned runs; explicit pins and active runs retain their versions. - The telemetry change removes source-managed GitHub identifiers from skill-reference events. It introduces no event or field. Please review the privacy boundary. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
25c422ba7e |
fix(apps): request minimal Google service scopes (#14740)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Google connections give agents service-specific tools through
governed credentials.
> - Each connection has a reviewed OAuth scope set.
> - Docs, Sheets, and Slides request Drive permissions in addition to
their own service scopes.
> - Google documents these permissions as alternatives, not combined
requirements.
> - This pull request removes those extra permissions and redundant
Calendar write free/busy access.
> - Users grant fewer permissions without adding tools or changing
connection access policy.
## Linked Issues or Issue Description
Refs #13820. Related #14739 changes other connector permissions; it does
not reduce these Google profiles.
Companion broker PR:
https://github.com/paperclipai/paperclip-cloud/pull/615. Ship the
matching changes together after fresh-grant validation.
**What happened?**
Seven Google profiles request redundant scopes. Docs, Sheets, and Slides
request Drive scopes. Calendar write requests free/busy even though
calendar.events authorizes its availability tool.
**Expected behavior**
Each profile requests only the scopes required for its reviewed tools.
Managed and customer-owned OAuth methods use the same set.
**Steps to reproduce**
Inspect the Google profile registry and the four app definitions on the
base commit. Compare their scope sets with Google's MCP authorization
alternatives linked in the updated documentation.
**Paperclip version or commit**
Base:
|
||
|
|
3c561642b4 |
fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask for decisions and optional details through cards in chat. > - A clear approval in a message can leave the matching card pending. > - An unanswered question can also block an unrelated later reply. > - Decisions need a saved source message, while optional questions need to remain answerable in history. > - This pull request records conversational decisions and lets users move on from questions and answer them later. ## Linked Issues or Issue Description **What happened?** Native Claude and Codex could act on approval in chat while the original approval card stayed pending. Pending question forms stayed above the composer, were absent from history, and could suppress later chat replies. A late native question answer could wait for a finished run to reconnect. **Expected behavior** The active agent records a clear approval or refusal against the exact card and user message. Ambiguous replies do not grant consent. Users can send another message without answering a question. The question remains pending in history and can be reopened and answered later. The saved answer reaches the agent. **Steps to reproduce** 1. Ask an agent to propose work with a confirmation card, then approve it in chat. 2. Check that the original card records that approval before work starts. 3. Ask an interactive question, send an unrelated message, and reload. 4. Open the unanswered question from history and submit an answer. Related work: #14408 added completion delivery. #14607 tests completion reporting turns. Neither records conversational answers on approval cards. ## What Changed - Add a confirmation endpoint backed by a user comment, with schema validation, OpenAPI discovery, and native Plan-mode access. Ask mode remains read-only. - Check company, active run, actor, current session, message provenance, revision, and resolver policy. Save the decision and audit in one transaction. Retries do not repeat effects. Emit resolution telemetry after commit. - Give fresh and resumed chat turns the actual pending confirmation identities. Teach agents to save clear conversational decisions before acting and to clarify ambiguity. - Keep unanswered Agent Chat questions as compact history entries. A newer user message closes the old form. Question cards never contribute to composer pending counts or navigation, including after dismissing a fresh form. The history card is the sole reminder; clicking it restores that exact form and draft. - Preserve Agent Chat questions when later messages or questions arrive. Historical ordinary inputs no longer gate later chat replies. Current-run requests, task execution, and governed approvals keep their gates. Remove the special acknowledgement-publication proof helpers that this rule replaces. - Route answers to finished native runs through durable fresh-wake delivery, with existing idempotency and source-question context. Settle late replies against contiguous completed conversation turns and freeze their history replay; failed, unhandled, and newly arriving messages remain actionable. - Add real-component Storybook scenarios, database and UI regressions, and a three-turn native Claude/Codex E2E case. Capture distinct, UI-ready screenshots and report the individual assertions. ## Verification - Focused decision/publication/UI regressions after merging master: 288 passed; subsequent UI draft, failed-send, and conversation checks: 199 passed. - Native question and durable delivery regressions: 106 passed, including all four terminal run states and exactly-once late delivery. Seven targeted regressions fail against the original implementation and pass with the fix. - Latest conversation/decision/native-delivery regressions after the master merge: 121 passed. Covers completed progress, missing or failed intervening turns, new messages during a late reply, stale sessions, and frozen retry/replay boundaries. Four new assertions fail before the ordering fix. - E2E support suite after the master merge: 792 passed. Negative controls reject expired cards, wrong questions/answers, stale or missing replies, unrelated clarification forms, and unexpected tasks. - The embedded-browser walkthrough caught one additional defect: dismissing a fresh question still showed a composer badge. Both Cancel and close-button regressions failed before the fix. The fix at `65f2ade12` passes 170 chat-thread tests and 792 E2E support tests. After merging master, 232 chat-thread/confirmation tests, server/UI typechecks, and token gates pass. The preview and two-provider live E2E pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at that commit. All 55 checks are now successful at `e5512a206` (four conditional checks skipped), including the aggregate verification gate and clean-install canary test. The first attempt was interrupted by simultaneous CI worker shutdowns; one failed-job rerun passed without code changes. - [Published Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on): nine real-component scenarios. Manually exercised move on, reopen, preserve draft, answer later, answer one of multiple questions, and a custom mobile answer in the embedded browser. Retested fresh Cancel and close-button dismissal in the updated build, then reopened and submitted the preserved Green selection and inspected its answered receipt. Static preview has no live model/backend; its callbacks are fixture responses. - [First live campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/) reproduced the late-answer completion-state defect on both providers despite correct saved answers and acknowledgements. It also exposed a valid imperative clarification rejected by the old oracle. Both issues are fixed with regression controls; this failing run is retained as evidence. - [Four-cell qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/) passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous confirmation, each on native Claude and Codex. Inspected saved state, source-message decisions, visible cards, and agent replies. Both late-answer chats settled to waiting; no unrequested tasks were created. [Final branch rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/) passed 2/2 at `142630720`: the same unanswered-question journey after merging master, plus an additional screenshot and browser assertion for the actual late-answer acknowledgement. - [Composer-reminder E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/) passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each, with explicit no-badge assertions before and after reload. Inspected saved pending/answered state, both screenshots with a clear composer, and actual Blue acknowledgements; all five behavioral matchers passed per provider and neither created tasks. Cost coverage is partial; this is bounded workflow qualification. - [Fresh-dismissal E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/) passed 2/2 at `e5512a206`: native Claude and Codex, including fresh Cancel, clear composer, reopen, unrelated message, reload, late Blue answer, and actual agent acknowledgement. All five behavioral matchers pass per provider. Inspected the fresh-dismissal screenshots and saved pending/answered identity; neither created tasks. Cost coverage is partial (4/6 runs). - Prior evidence remains available in [the earlier campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/). Its early loading screenshot and overwritten final capture prompted the UI-ready, distinct screenshot fixes. ## Risks - The model interprets intent. The server verifies permission and provenance; it does not infer consent from text. Ambiguous and unrelated replies are not approvals. - Historical questions can accumulate. They remain visible, pending, and answerable; no automatic answer or expiry is invented. - The change to completion gates is scoped to Agent Chat and ordinary historical inputs. Current-turn and governed approvals retain their existing controls. - Live qualification is limited to the selected stories. Broader native onboarding finalization remains separate work. - No database migration. Telemetry adds no fields or values; the contract and README document the commit boundary. Privacy review was requested on the PR. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser-test orchestration. The exact model ID and context-window size are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a36cbffa9e |
fix(connections): broaden natural-language and aggregator search (#14725)
## Thinking Path > - Paperclip manages agents and the services they need for work. > - Agents use connection search to discover a setup path before they request access. > - Tool-only filtering hid channel and AI methods from this search. > - Requiring every query word to match rejected normal task descriptions. > - A small aggregator index also omitted supported apps such as Circleback. > - This pull request broadens retrieval and returns purpose-specific setup guidance. > - Agents can choose a relevant result while existing access and provider-choice checks still apply. ## Linked Issues or Issue Description **What happened?** A search such as “AgentMail create an email address and manage an agent mailbox” returned no usable result. “Help me find tools for circle back” also missed Composio's supported Circleback toolkit. Queries longer than 200 characters failed validation. **Expected behavior** Return useful native and verified aggregator matches from natural-language queries. Include channel/email methods when Chat connectors is enabled. Identify each method's purpose and the correct setup path. **Steps to reproduce** Use the queries above with `connections_search` from an active task. Enable Chat connectors for the AgentMail case. The regression suite reproduces these misses before the change. Related routing work: #13941. This change does not change the runner failure path or add channel setup to tool-only connection cards. ## What Changed - Rank name and capability matches. Accept extra words, split names, small spelling errors, and queries up to 4,000 characters. - Include tool, channel/email, and AI methods. Return company-prefix setup links for channel and AI flows. - Add a dated snapshot of 1,583 official Composio toolkit names and a refresh script. Merge duplicate MCP variants for search and link each support claim to official evidence. - Find authorized indexed aggregator namespaces within longer queries. Return multiple app matches when the agent needs to choose. - Prefer exact app names over fuzzy matches for other apps; retain existing AI readiness. - Preserve native preference, company and identity boundaries, administrative denials, and saved provider consent. - Add relevance and database regressions, extend native tool-authority coverage, and document search behavior. ## Verification - Red: 15 new assertions failed against the previous implementation; the existing baseline passed. Added red-green regressions for Motion versus fuzzy Notion and existing AI access during review. A further regression covers mixed ready/unconfigured AI results and their per-result setup guidance. - Green: all 103 tests in the eight focused shared, database, runtime-tool, fixture, and route suites pass on the latest commit. - `pnpm -r typecheck` and `pnpm build` passed. - Latest-commit CI passed: 54 successful checks and two skipped checks, including the full test matrix, browser E2E, typecheck, build, and canary dry run. - The long local `pnpm test:run` attempt began before the review fixes and retained transformed pre-fix search code; it also hit an unrelated timing failure. Fresh serial reruns of the affected search suites and three timeout cases passed all 124 tests. Parallel local route shards hit two additional database setup timeouts; both suites passed all 17 tests on a fresh serial rerun. The complete corresponding CI suites also passed. Duplicate broad local runs were stopped after CI completed. The local UI suite independently passed all 7,007 tests. - Greptile: 5/5 on `cc6a0180d`; all review findings resolved. - Browser verification passed in a disposable local instance through real process-agent search requests: AgentMail opened its setup flow with the requester selected; the saved Circleback choice produced the Composio setup card; a paragraph-length Notion query produced its setup card. No provider credentials or external accounts were created. - The browser test caught an invalid UUID-based setup URL. The fix uses the company prefix and has a regression assertion. - A 3,971-character catalog query found Circleback first in a local 10 ms spot check after sharing query preparation across the catalog scan. This is a single measurement, not a performance guarantee. ## Risks - Broader retrieval can return extra candidates. Named services rank first; agents must select the relevant method. - The public support snapshot can age. It proves catalog support, not account authorization or the availability of every requested action. - Channel and AI methods use existing setup links. The tool connection card still accepts tool methods only. - No schema, migration, credential, or runner lifecycle changes. ## Model Used OpenAI Codex (GPT-6). The session does not expose a more specific model identifier or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d72389bee2 |
feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps gateway gives agents governed access to external tools. > - Browser Use Cloud can run browser work, but a tool result alone does not let a person watch or take over. > - A task needs a durable browser session, a visible viewer, and recorded costs. > - This pull request adds a Browser Use Cloud v4 connection and interactive browser tabs on tasks. > - People can follow the work, interact with the page, and retain the browser after the agent finishes. ## Linked Issues or Issue Description **Problem or motivation** Agents need governed access to Browser Use Cloud. People need to see and interact with the same browser from the task. A browser must remain available after a run finishes and appear at the correct point in the task feed. **Proposed solution** Add a native REST connection for the v4 API. Bind each session to its company, task, agent, and credential grant. Open its interactive viewer in the task side panel. Record provider costs as financial events. Use `browser-use-cloud` as the app and connector key. Keep its skill with the connector and deliver it only with authorized connection tools. **Alternatives considered** A v3 MCP connection would expose tools without the v4 lifecycle integration. An external viewer link would leave the task. A fixed viewer size would prevent pages from responding to changes in the task pane. **Roadmap alignment** This extends the governed Apps gateway and Connected Apps roadmap. It uses the existing task, grant, secret, approval, and financial records. The work was requested by the maintainer. A search found no duplicate Browser Use connector PR or issue. ## What Changed - Add the Browser Use Cloud app, brand asset, API-key connection, and profile settings under the `browser-use-cloud` key. - Bundle the `browser-use-cloud` skill with the connector. Keep it out of global `skills/` discovery. Deliver it only with authorized task/run connection tools. Remove retired connector skill keys from runtime overlays and preserve unrelated browser skills. - Expose seven v4 tools through the governed gateway and deliver them to native and CLI agents. - Persist sessions, browsers, runs, event and recovery cursors, shutdown leases, and cumulative cost accounting. Recover uncertain paid starts without replaying them. - Enforce task ownership, credential grants, approvals, revoked access, and budget limits. - Add interactive task browser tabs and compact chronological feed entries. Retain the viewer across tab switches and keep visible idle browsers open. - Add debounced automatic viewport fitting, standard size presets, and a viewer ownership lease. - Add lifecycle, authorization, accounting, viewport, UI, and Storybook coverage. - Add an idempotent database migration after the current master migration. Preserve deployed migration hashes. Migrate pre-release Cloud connection and financial keys without replacing grants, credentials, or browser history. - Document provider behavior, live acceptance results, and the lack of documented passkey forwarding. ## Verification - Full workspace typecheck and production build pass on the updated branch. - Token gates, brand asset validation, module boundaries, and migration ordering pass. - Cloud tests verify global skill exclusion, authorized task/run delivery, unassigned agents, disabled connections, revocation, adapter isolation, and secret exclusion. The existing AgentMail connector assignment test also passes. - Migration replay runs twice against existing browser work and financial records. It preserves the records and avoids duplicate costs. - The focused provider, app catalog, OpenAPI, connection gateway, and migration regression suites pass. Recovery coverage includes lost replies, process crashes, provider rejection, and browser arrival acknowledgement. - All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`, including the full test matrix, browser E2E shards, build, typecheck, security, and release canary. Two optional Storybook jobs are skipped. - Greptile is 5/5 on the same commit, with zero unresolved review threads. The corrected review uses the actual master-to-head diff. - The local `pnpm test:run` started and was stopped after the full CI matrix passed. It did not complete locally; the full-suite result above comes from CI. - Earlier live acceptance used an isolated company with a capped provider credential. The agent opened paperclip.ing, the embedded viewer accepted navigation, and the same browser stayed available after completion and tab switches. - The local Storybook build passes. Stories cover the panel, footer, feed entries, settings, lifecycle failures, and viewport modes with an offline viewer fixture. ## Risks - Browser Use charges for hosted work. Provider caps and local budget checks reduce exposure; reported costs can arrive after work completes. - Viewer and CDP URLs grant access to the browser. The server validates and restricts them. They are excluded from agent results and durable event data. - Runtime resizing of v4 agent browsers uses a provider option confirmed by live testing but absent from its published agent schema. Resizing during a click may invalidate coordinates. Fixed presets remain available. - Viewport ownership is process-local and resets on restart. The lifecycle and accounting records remain in the database. - The original intermittent embedded-viewer stall has not been fully diagnosed. A bounded reconnect and active-session recovery cover the observed failure paths. - Live tests did not cover every revocation, approval, rate-limit, or restart case. Deterministic integration tests cover those paths. Passkey forwarding is not claimed. - Unknown create outcomes keep the credential available for cleanup. Run-list absence cannot prove a paid POST was rejected, so recovery stays pending until it can identify provider work. ## Model Used OpenAI Codex, GPT-6. Used reasoning, repository search, code execution, browser interaction, and test tools. The exact serving model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2f6fa3b6dc |
fix: recover provider authentication inside tasks (#14629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a working model provider connection to run a task. > - A provider can reject a stored credential after the task starts. > - The failed run must ask the responsible user to repair that connection. > - This pull request adds that request directly to the task and reuses Connections sign-in. > - The user can choose an API key or subscription, then continue the task with a fresh session. ## Linked Issues or Issue Description **What happened?** A run that ended with `acpx_auth_required` or another known provider authentication error did not immediately offer an inline way to connect the provider. A repair form could also lock the user to the failed account's sign-in method. **Expected behavior** Show a provider connection card in the task as soon as the authentication failure is saved. Allow the responsible user to connect or repair the provider with any supported sign-in method. Keep the connection name automatic and resume the task after successful setup. **Steps to reproduce** 1. Run a task with a supported provider and an expired or invalid credential. 2. Let the run fail with a provider authentication error. 3. Open the task and attempt to repair the connection. Related work: Refs #13724 and #13726. This change adds the inline task repair flow and method choice. ## What Changed - Classify provider authentication failures and create one connection request for the current task. A persisted blocked classification suppresses automatic retries only after the repair card is created; unsupported providers retain their existing recovery path. - Mark only the attributed, unchanged credential as needing sign-in. Preserve credentials that were refreshed after the failed run started. - Reuse the provider sign-in controls inside the task. Allow API key and subscription choices for Claude, Codex, and Grok. Keep names hidden and generate a default from the user, provider, and method. - Keep the existing account when reconnecting with the same method. Create and select another account when the method changes. Validate updates to explicit agent bindings through the normal agent save path. - Require explicit adoption for legacy agent authentication. Validate in the agent environment, then commit the binding, connection install, audit, and card completion in one transaction. Keep failed setup and account selection visible and retryable. - Add regression tests and update the specification and Connections documentation. ## Verification - Fresh local verification: 199 tests passed across the inline form, provider method selector, default naming, authentication and recovery classifiers, run liveness, OpenAPI routes, database adoption/rollback, and Cursor execution suites. The adoption database suite also passed against disposable Docker PostgreSQL. - Full repository `pnpm build` and `pnpm -r typecheck` passed on the latest commit. Token gates are clean. - Embedded browser: opened real task cards from seeded authentication failures; switched Claude from API key to subscription and back; switched Codex from subscription to API key; confirmed the name field stays hidden. Provider sign-in was not completed with real credentials. - The broad local `pnpm test:run` started before review fixes and was interrupted after the working tree changed; it is not counted as a passing full run. Fresh focused tests passed. CI supplies the full test and browser suite results for the current commit. - CI is green on commit `4b97a4e447045ff3d7516525a187a5d1d21e0d4c`: 54 checks passed and two Storybook checks were skipped by their path rules. The workspace preview job passed on one rerun after a local-server startup timeout; its rerun passed 835 tests. - Greptile is 5/5 on the same commit with no actionable findings and no unresolved review threads. ## Risks - Incorrect authentication classification could prompt for a connection unnecessarily. Tests exclude tool authorization, quota, and unrelated runtime failures. - A method change selects the new personal provider default, which also applies to other agents that use that user's default. Explicit account bindings use the existing permission and runtime validation path. - Credential invalidation must not race with refresh or reconnect. The code compares the saved credential generation and grant update time under locks. - No database migration or new credential storage format is required. ## Model Used OpenAI GPT-6 through Codex. The exact model ID and context window size were not exposed in this session. Capabilities used: reasoning, repository editing, shell commands, database tests, and embedded-browser interaction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused tests listed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f38b5693f6 |
fix: always enable keyboard shortcuts (#14643)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The web UI has keyboard shortcuts for the inbox, task lists, cases, and task detail, plus global shortcuts such as `c`, `/`, `?`, `[`, and `]` > - Shortcut enablement was an instance-wide General setting until #14141 moved it to a per-user preference that defaults to off > - The move did not carry the old instance value over, so every existing user lost shortcuts on upgrade and had to find a new toggle under Profile settings > - A toggle that only turns off a standard, input-safe feature costs a setting, a database column, two API routes, and a React context for little benefit > - This pull request removes both the instance setting and the personal preference and enables keyboard shortcuts for every signed-in user > - The benefit is one less thing to configure, no silent loss of shortcuts on upgrade, and less code to maintain ## Linked Issues or Issue Description Refs #14141 (the change that introduced the personal preference). **What existing behavior does this improve?** Keyboard shortcuts in the web UI stay off unless each user turns them on in Profile settings. **Subsystem affected** Web UI shortcuts, Profile settings, instance general settings, the `/api/auth/preferences` routes, and the `user` table. **Current behavior** Shortcuts default to off per user. #14141 moved the toggle from Instance settings → General to Profile settings and did not carry the old instance value over. Users who had shortcuts on lost them after the upgrade and had to find the new toggle. **Proposed behavior** Keyboard shortcuts are always enabled for every signed-in user. There is no instance setting and no personal preference. Shortcuts already ignore key presses inside text inputs and modal dialogs, so an opt-out is not needed. **Reason and benefit** Fewer settings, no silent loss of shortcuts on upgrade, and removal of a database column, two API routes, a query hook, and a React context that existed only to gate this feature. **Breaking changes** `GET` and `PATCH /api/auth/preferences` are removed. `PATCH /api/instance/settings/general` no longer accepts `keyboardShortcuts`; that schema is strict, so the key now returns 400. `instance.general.keyboardShortcuts` is no longer a valid `PAPERCLIP_HIDDEN_SETTINGS` key; the parser ignores unknown keys with a warning. ## What Changed - Removed the Keyboard shortcuts section from Profile settings, the `useUserPreferences` hook, `queryKeys.auth.preferences`, and `authApi.getPreferences` / `authApi.updatePreferences`. - Removed `GeneralSettingsContext`. The inbox, legacy inbox, task list, legacy task list, cases, and task detail pages no longer gate their key handlers. - Removed the `enabled` option from `useKeyboardShortcuts`. The app shell always registers the global shortcuts. - Removed `GET` and `PATCH /api/auth/preferences`, their OpenAPI entries, and the `currentUserPreferencesSchema` / `updateCurrentUserPreferencesSchema` validators. - Removed `keyboardShortcuts` from `InstanceGeneralSettings`, the general settings zod schema, the settings service defaults, and `HIDEABLE_GENERAL_SECTIONS`. - Added migration `0289_drop_user_keyboard_shortcuts`, which drops `user.keyboard_shortcuts`. - Updated `AGENTS.md`, `doc/SPEC.md`, `doc/SPEC-implementation.md`, and `docs/deploy/environment-variables.md`. - Parsed the stored general settings row with `instanceGeneralSettingsSchema.strip()` in the feedback vote path, so a retired key left in the row cannot reset the sharing preference to `prompt` and overwrite the stored choice. - Kept every bare global shortcut (`c`, `?`, `[`, `]`, `/`) out of open modal dialogs in `useKeyboardShortcuts`; only `/` had that guard before. - Updated the affected tests and added a Profile settings test that asserts the toggle is gone, a hook test for the modal dialog guard, and a feedback service regression test for the retired-key case. ## Verification - Typecheck passes for `@paperclipai/shared`, `@paperclipai/db` (including the migration numbering and safety checks), `@paperclipai/server`, and `ui`. - `pnpm exec vitest run server/src/__tests__/instance-settings-routes.test.ts server/src/__tests__/openapi-routes.test.ts server/src/__tests__/auth-routes.test.ts server/src/__tests__/sentry.test.ts` → 119 passed. - `pnpm exec vitest run ui/src/components/Layout.test.tsx ui/src/pages/ProfileSettings.test.tsx ui/src/pages/IssueDetail.test.tsx ui/src/pages/Inbox.test.tsx ui/src/pages/Cases.test.tsx ui/src/hooks/useKeyboardShortcuts.test.tsx ui/src/pages/Agents.test.tsx ui/src/pages/InstanceGeneralSettings.test.tsx` → 286 passed. - `pnpm exec vitest run packages/shared/src/settings-visibility.test.ts` → 16 passed. - `pnpm exec vitest run ui/src/hooks/useKeyboardShortcuts.test.tsx` → 7 passed. - `pnpm exec vitest run server/src/__tests__/feedback-service.test.ts` (embedded Postgres) → the new retired-key test passes with the fix and fails without it. - Manual: sign in with no settings changed, open the inbox, press `j` and `k` to move the selection, press `?` to open the cheatsheet. Open Settings → Profile and confirm there is no Keyboard shortcuts section. ## Risks - The migration drops a column. It uses `DROP COLUMN IF EXISTS`, and the column has no readers after this change. If you roll back to a build from before this PR after the migration has run, re-add the column first: `ALTER TABLE "user" ADD COLUMN "keyboard_shortcuts" boolean DEFAULT false NOT NULL;`. The older build's ORM selects that column when it loads users. - Any external client that still sends `keyboardShortcuts` to `PATCH /api/instance/settings/general` receives a 400. No in-repo client does. - Stored `instance_settings.general.keyboardShortcuts` values are stripped on read and ignored. - Users who never turned the toggle on now get shortcuts. The handlers skip text inputs, contenteditable regions, and modal dialogs, so typing is unaffected. ## Model Used Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
7636966452 |
fix(inbox): keep other users’ failed runs out of Mine (#14572)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Mine inbox shows work that needs the current user. > - Failed-run rows used the latest run for every agent in the company. > - A failure from another user therefore appeared in Mine and its badge. > - Run list responses also omitted the responsible user needed to filter these rows. > - This pull request uses run ownership for personal failure routing. > - Users see their own failures and can still inspect company failures in All. ## Linked Issues or Issue Description **What happened?** An agent run started for one user failed. Its row and failure badge appeared in another user's Mine inbox. **Expected behavior** Mine and its badge include failed runs for the current responsible user. Other users' failures remain available in All and run details. **Steps to reproduce** 1. Use a company with two human users. 2. Create a failed or timed-out run attributed to the first user. 3. Open Mine as the second user. Before this fix, the failed run appears there and increases the badge. **Paperclip version or commit** Reproduced in regression tests on master at `24beb0057`. **Deployment mode** Authenticated deployment with multiple users. Tests also cover the local single-user board. Related prior work: #933 addressed inbox dismissal and badge consistency. No duplicate ownership fix was found. ## What Changed - Return `responsibleUserId` in normal and summary run lists. - Share one ownership rule across both inbox versions and client/server badges. - Select the latest run per agent before applying the ownership filter. This prevents old failures from resurfacing on shared agents. - Keep unattributed historical failures in the local board's Mine view. Hide them from authenticated users with no matching owner. - Keep company health alerts outside the personal badge, consistent with the client. - Document the routing contract and add page, badge, and database regression coverage. ## Verification - Red: the new badge cases failed with three company failures instead of one personal failure; eight Mine page cases failed across both inbox versions. - Green: 113 focused tests pass in `ui/src/lib/inbox.test.ts`, `ui/src/pages/Inbox.test.tsx`, `server/src/__tests__/heartbeat-list.test.ts`, and `server/src/__tests__/inbox-dismissals.test.ts`. - `pnpm check:token-gates` passes. - Agent calls on behalf of a user have two additional red-to-green API regressions. - Full `pnpm -r typecheck` and `pnpm build` pass. Server typecheck also passes after the agent-call fix. - All CI test shards and browser tests pass on `243bfa681`. The duplicate local `pnpm test:run` was stopped after the CI test lanes completed; it did not finish locally. ## Risks - Authenticated users no longer receive unattributed legacy failures in Mine. Those failures remain visible in All. - The server badge no longer counts company health alerts, matching the existing client badge. - No migration, run state, retry behavior, or company access rules change. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, terminal execution, and browser tools. The exact deployment variant and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3ca196b0a6 |
feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - An agent needs personal files across tasks and sessions. > - AGENTS.md is one file in that directory. Supporting files need the same persistence. > - The Instructions Editor and agent runs must share one current directory. > - Concurrent runs should apply only the files they change. The last sync of the same file wins. > - This pull request uses existing file transport and removes temporary copies after sync. > - Old instruction-only sessions keep their restore contract. New saves do not create revision history. ## Linked Issues or Issue Description Refs #14325. This replaces its revision-oriented design with persistent agent files. Keep #14325 unmerged. Transport prerequisite #14416 merged first at `d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master and remains below 100 changed files. Related work: #4513 and #8798 cover instruction tooling. This change handles run synchronization, cross-task personal files, browser editing, and old-session restoration. ## What Changed - Keep one current directory per company and agent. Point AGENT_HOME at a temporary working copy for each active run. Keep task files and provider HOME separate. - Restore text, binary files, and nested folders through workspace transport. Exclude remote agent files from task Git snapshots with a self-ignoring file inside the reserved runtime directory; never write through repository-controlled Git metadata. - Collect after the provider and child processes have stopped. Keep resumable conversation state. - Apply changed and deleted files under the agent lock. The last sync wins for the same file. Unrelated concurrent changes survive. - Remove temporary copies after successful sync, rejected sync, and staging failure. Register ownership before copying so restart recovery can remove interrupted preparation. Retry transient synchronization up to three times. Preserve the original remote lease reference until deletion succeeds; restart cleanup never acquires a replacement sandbox. Do not create captured directories or a conflict-review queue for new runs. - Keep browser editing, stale-draft protection, and streaming binary downloads. Keep the instruction entry and text editor limited to 1 MiB. - Keep historical agent-folder sync failures on their affected runs instead of repeating them above current saved instructions. Preserve legacy candidate review and current browser-save errors. Avoid duplicate quota warnings while retaining separate sync failures when they describe a different problem. - Require target-scoped caller grants for peer instruction access, while preserving self edits, responsible-user checks, and protected-change consent. - Treat full storage as a nonblocking run warning, never an agent pause or run-admission failure. Restore already-over-quota saved folders so ordinary agent cleanup can recover; warn on each run until cleanup. The run detail view shows the warning. - Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash large files as streams. Check editor-save quotas with metadata instead of hashing unrelated files. - Preserve old native inputs, instruction-only copies, paths, digests, and pending legacy candidates. Adopt old revision heads once. New writes do not append history rows. - Add idempotent migration 0287 and verify upgrades from the preview tables and receipts. - Add nine interactive stories under **Agents / Persistent files**, including automatic incoming edits, stale browser drafts, and storage-limit diagnostics. ## Verification - Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after merging current master and the landed transport prerequisite. Integration required no manual conflict resolution; the feature remains 99 changed files. Full workspace typecheck, production build, token gates, and 715 focused tests passed on this merge candidate. Fresh Greptile review is 5/5 with no unresolved findings. All 55 checks passed, with four conditional skips, including the build, typecheck, browser E2E, and canary dry run. A single retry recovered four jobs interrupted by runner shutdowns; no source changes were required. - Historical-warning UI fix: all 6,834 UI tests across 640 files passed, including regression coverage for three old failures, legacy preserved edits, and warnings scoped to the affected run. Full workspace typecheck, production build, Storybook build, and token gates passed. Browser-verified Storybook playtests passed for Historical Failures After Successful Save, Storage Limit, and Full Storage Run Warning. - Review follow-ups at `4e20c9fb2`: all 18 focused tests passed, including external Git directories, linked worktrees, symlinks, hardlinks, and distinct I/O failures alongside storage warnings. Server and UI typechecks, token gates, and the production build passed. - Storage warning regressions at `0724f3012`: all 33 directory tests and all five heartbeat-list tests passed, with no skips in their successful runs. They cover repeated runs while full, an already-over-quota saved folder, cleanup, warnings retained after unrelated save failures, and bounded warnings in large result JSON. Server typecheck passed after the final warning fixes. - Full workspace typecheck, production build, and token gates passed during this follow-up. Product E2E harness: 631 tests passed across 52 files; harness typecheck passed. Earlier native session/context and directory/legacy collection suites passed 537 tests; Runner unit/transport suites passed 329 tests. - **Real E2E at `0724f3012` (before this follow-up):** legacy local Codex and native Daytona Codex each passed six tasks, one server restart, seven independent assertions, and cleanup verification. Both prove browser-to-agent edits, agent-to-browser edits, nested/binary restoration, per-file last-sync-wins, a successful run after an oversized save rejection, and cleanup clearing the warning. - Native local Codex also passed the six-task quota flow before the final warning-retention fixes. That pass began at `918d1ed02` while the bounded-result warning fix was being edited, so it is not claimed as exact-final-head evidence. Its final-head rerun failed during embedded PostgreSQL bootstrap before any provider run: the macOS host had 87,365 of 87,381 SysV semaphores occupied. No unrelated services or kernel limits were changed. - The final-source report intentionally records **2/3 cells passed**, preserving the blocked native-local attempt: `tests/runner-e2e/results/agent-files-quota-final-20260928-report/`. Earlier failed attempts and provenance notes remain under `tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and the original campaign directories. - Daytona used immutable image `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf` and its exact Linux runner binary. Controller source is `0724f3012`; image source is recorded separately. - Legacy-session compatibility and all three ACP Stop/resume browser regressions passed on the prior validated feature head `169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same provider session is retained and interrupted writes are not replayed. Migration upgrade tests also passed earlier. - Nine interactive stories are under **Agents / Persistent files**, including **Full Storage Run Warning**. Its playtest and visual browser inspection passed; the warning states that runs continue and the editor remains available. - Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs skipped, no failures or pending checks. All eight browser E2E shards and their aggregate passed. Fresh Greptile review is 5/5 with no findings; all review threads are resolved, the security scan passed, and GitHub reports no merge conflicts. - The broad local follow-up test run was interrupted after host semaphore exhaustion affected isolated PostgreSQL instances. It also encountered the existing macOS long-path fixture failure and two timeout failures. This is not a claim that the full local suite passed. Logs are retained; focused storage/warning tests passed. ## Risks - A later sync can overwrite an earlier edit to the same file, including a saved browser edit. There is no text merge or retained version. This is the intended last-sync-wins policy. - A save that exceeds a storage limit is rejected and its temporary copy is discarded. The run itself continues normally, and later runs restore the last saved files with a warning until cleanup. Transient sync failures get bounded retries. An I/O failure partway through a sync can leave some files updated; a failed receipt does not claim whole-folder success. - Larger folders increase copy time, network traffic, and temporary disk usage. Active runs still need working copies. Terminal runs do not accumulate archives. Operators must provision disk for agents and configured concurrency; these limits are not company-wide quotas. - A restored old native session remains instruction-only until a fresh session starts. Its original conflict fence and existing pending candidates remain compatible. - Provider processes close at the collection boundary. Conversation resume remains available, but warm process reuse is lost. - Backups must include the instance filesystem and database. External bundles keep their existing behavior until explicitly moved to managed storage. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and browser tools assisted this change. Real provider E2E uses `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing> |
||
|
|
992f720262 |
fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task descriptions, comments, continuation data, skills, and execution rules enter several agent adapters. > - The same source can be rendered by more than one automatic input carrier. > - Failed resumes can also rebuild input from stale or compact context. > - This pull request gives each Paperclip-owned source one delivery owner and preserves the required transport boundaries. > - It adds deterministic adapter, interaction, runner, and browser tests for these boundaries. > - The benefit is more predictable context delivery with explicit evidence for later live qualification. ## Linked Issues or Issue Description Related: #13144 removes a duplicate environment payload and bounds wake lists. Related: #11360 addresses Hermes resume behavior. This pull request preserves compatible active-session formats while repairing context ownership and stale question creation. **What happened?** Task descriptions and comments could enter more than one automatic context block. Native transports could wrap a complete model input in a second task envelope. Some legacy and gateway adapters could omit the owned assignment on ordinary tasks or rebuild a failed resume with stale compact context. A continuation could also request a question after newer human comments had arrived. **Expected behavior** Each task or comment source has one automatic model-facing owner. Distinct comment IDs and repeated wording remain distinct. Fresh fallback attempts rebuild the required full context. A question request is rejected when newer queued human direction makes it stale. Harness access policy remains owned by execution configuration. **Steps to reproduce** 1. Build a task with a description and current comments. 2. Capture the actual adapter or runner input. 3. Compare source ownership and task-envelope nesting. 4. Queue a human comment before a continuation requests a question. 5. Trigger a failed resume and inspect the fresh retry input. 6. Run the focused adapter, interaction, runner, and browser checks. ## What Changed - Add shared prompt-section selection at the provider-attempt boundary. - Deliver owned assignment context through native, legacy CLI, ACP, gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and Hermes paths. - Rebuild full or compact context after resume recovery changes the attempt. Add native and Claude ACP tests of actual recovery requests. - Preserve custom templates, loaded instruction files, execution policies, and older active-session formats. - Record continuation source metadata and reject stale question creation under the issue-row lock. - Add explicit Product E2E context-integrity profiles, prerequisite gates, credential-isolation checks, and report fixtures. - Bypass service-worker forwarding for same-origin Vite development modules. A real Chromium test fails with resource exhaustion before the repair and passes after it. Production asset caching keeps its existing policy. - Add browser diagnostics and service-worker module-loading regressions. - Add an explicit zero-retry eval option. The default retry behavior remains unchanged. Each campaign records its effective policy. - Remove the model-facing working-directory sentence from four prompt builders. Existing workspace, sandbox, permission, and custom-template configuration remains unchanged. - Align the everyday workflow assertion with the current 47-entry catalog. Compared with current upstream master, the branch carries the context-ownership implementation and its tests, the explicit context-integrity catalog and evidence harness, and the focused browser regression checks. ## Verification **Merge assessment:** focused regression evidence supports merge. This is not full completion of the original broad qualification matrix. The maintainer has authorized merge after fresh verification of the master integration. - Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`. All 14 conflicts are resolved. Cancellation checks, workspace finalization, native Grok support, and both sets of tests are retained. - Current-head Greptile: **5/5**, with no blocking findings. The review names this exact commit. All **59 reported checks are terminal: 55 successful, 4 skipped, zero pending or failing**. This includes the full root general and serialized suites, separate runner checks, typecheck, build, canary, browser E2E, Docker, and security checks. The successful legacy security status is included in that total. - After integration: workspace typecheck and full build passed. Separate runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust tests, and 39 preparation checks**. Other passing checks include 621 Product E2E harness units, 376 focused shared/adapter tests, 160 real-database/API tests, 86 Hermes tests, 18 browser-support checks, and Product E2E typechecking. The complete root suite passed in CI. The duplicate local monolithic root run was stopped after that CI result; it is not counted as a completed local pass. - New native recovery coverage retains full assignment, completion contract, and explicit skill selection after safe replacement, for old and prepared input formats. Full native session test file: **136/136 passed**. - New Claude ACP coverage captures actual fresh, resumed, and missing-session fallback requests. It verifies one assignment copy, comment order, identical text under distinct comment IDs, and full fallback context. Full file: **33/33 passed**. Both affected TypeScript checks passed. - Existing deterministic tests cover source revisions, approval and trust boundaries, completion validation, custom templates, compatible sessions, standalone driver wrapping, and maintained adapter transport requests. - Provider-free browser support: **17/17 passed** after the master merge. Service-worker unit tests: **33/33 passed**. The module-overload regression failed before the repair and passed after it in real Chromium. ### Fresh live comparisons The new batch ran exactly four Product E2E attempts. **All four passed on the first attempt; no retries.** Each has six terminal matchers plus the existing browser lifecycle and invariant checks. | Exact case ID | Control | Candidate | |---|---|---| | `core-compatibility.runner-codex.local.plan-revise-accept` | Passed | Passed | | `local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume` | Passed | Passed | The plan case checks a revised canonical plan and revision-bound approval before completion. The question case restarts the server before submitting the answer, then verifies the continuation completes. Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical frozen definitions and provider versions: Codex `0.156.0` with `gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with `claude-sonnet-5`. The September 24 head added master browser recovery and test-only changes. The September 28 head also integrates newer master changes, including cancellation, workspace finalization, and native Grok. These are frozen-source live results, not exact-head live runs. The candidate received one description copy where the control initially received three. The submitted initial plan envelopes were 7,969 versus 19,097 characters. Question envelopes were 7,592 versus 18,919. These are structural measurements, not whole-provider token or dollar savings. ### Earlier evidence and failed attempts - The preceding fresh batch has four effective passing pairs: OpenCode comment continuation and assigned skill, native Codex comment continuation, and native Claude comment continuation. It retains **11 attempts: eight passed and three failed**. - Original failures remain recorded: missing local PostgreSQL library links before task creation; host-sleep cleanup after task/page checks passed; and a Claude **control** session-open rejection before a model turn. Setup was repaired identically on both worktrees. The permitted unchanged infrastructure retries passed. The underlying Claude provider startup error was not retained and remains unknown. - Older R2 retains **17 passes and one failure** across 18 attempts, including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode blank-page failure led to the service-worker repair. R2 is historical evidence: master changed the native fixed prompt and removed duplicate wake environment data afterward. - The September 24 CI run initially failed one unrelated preview readiness test (`ECONNREFUSED` on its local fixture). Its test and production code match master. Isolated local verification passed **28 tests, 3 skipped**. One unchanged CI retry passed the full shard: **831 passed, 1 skipped**, including all **31 preview-exposure tests**. The aggregate CI gate passed afterward. The precise startup cause remains unknown; a port race is a hypothesis, not a proved cause. ### Limits The original wider profile/workflow matrix, repeated trials, and remote Daytona qualification are incomplete. These results support a focused merge recommendation, not statistical equivalence or universal harness qualification. Some usage receipts are missing in both variants, so no token or dollar savings are claimed. The $500 ceiling was preserved using conservative allowances; failed attempts and unknown charges remain in the ledger. Reproduce the focused additions with `pnpm exec vitest run packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/native-session-runtime.test.ts`. Full checks use `pnpm -r typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner checks. Paid evals require the frozen definitions, profiles, and credentials; do not use `--all` as a substitute for the selected cases. ## Risks - Context placement changes can affect model behavior. Deterministic checks cover the selected paths, but live qualification remains incomplete. - The stale-question guard can reject a request when queued human comments arrived during the run. This is intended. - New stored inputs and model envelopes retain compatibility readers for older active sessions. - Custom templates may intentionally repeat content. - Removing a model-facing working-directory sentence does not change filesystem, command, sandbox, or permission configuration. - The worker bypass applies only to same-origin development module paths. Cache-policy tests preserve private-response handling and production asset caching. Mounted HTTP fixture changes remain test-only. - This PR does not claim measured token savings or statistical equivalence across every harness. ## Model Used OpenAI Codex, exact model gpt-6-astra, with repository tools and code execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The serving context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR using the required issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused local checks and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect these changes - [x] I have considered and documented risks above - [x] All current-head Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups for the current head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1a394bd30 |
feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
14795136f5 |
fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider. Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3447609d22 |
fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.
## Linked Issues or Issue Description
**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.
**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.
**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.
Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.
## What Changed
- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.
## Verification
- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.
## Risks
- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.
## Model Used
OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
8751e2de46 |
fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task board shows which recovery actions are active. > - A native run can stop while a person must repair its workspace. > - The board previously called that state “Recovery in progress.” > - The label implied that work would continue without operator action. > - This pull request derives the label from the recovery owner and live continuation. > - Operators can distinguish scheduled recovery from a repair that needs attention. ## Linked Issues or Issue Description **What happened?** A blocked task showed “Recovery in progress” after native finalization stopped and no automatic continuation remained. **Expected behavior** Show “Recovery needed” for an idle board repair or when no live recovery path exists. Show “Recovery in progress” while the recorded continuation can run, including an explicitly admitted export whose exact callback is executing even if the old recovery action remains board-owned. **Steps to reproduce** 1. Complete a native run whose workspace export cannot be recovered automatically. 2. Inspect the task recovery action and badge. 3. Compare the board-owned action with the old “Recovery in progress” label. Related work: #14314 serializes native workspace finalization and fences stale recovery outcomes. This change reports the recovery action that currently owns the task. ## What Changed - Show “Recovery needed” for board-owned active-run recovery unless the exact native export callback is positively verified as executing. - Require the native continuation run and a live or future continuation before showing progress. - Project native activity from the exact company, source issue, and run. Include active workspace export while the original heartbeat remains failed. - Require an executing callback before using a running export row as evidence. Preserve activity for long exports and clear it when the callback joins. - Give native resume its own card explanation. Preserve ordinary watchdog observation behavior. - Remove the redundant ownership sentence from all six recovery-card explanations that used it. - Document the labels and add regression cases for stopped, scheduled, and active recovery. ## Verification - Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests and token gates pass. No UI occurrence of the removed sentence remains. `pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54 successful checks, two optional Storybook checks skipped). Greptile is 5/5 with no open review threads. The duplicate local `pnpm test:run` was stopped after the full CI suite passed; it did not complete locally. - Original regressions: seven failures before the change, then 52 focused cases pass. - Review regressions: seven UI failures and nine database failures before the follow-up. All 75 database/API recovery tests, 167 UI tests, and 18 workspace lifecycle/finalizer tests pass. A further three RED cases cover explicit board retry activity; one RED case rejects orphaned running export rows after controller loss. Wrong company, issue, run, service, phase, and completed-operation cases remain inactive. - Final recursive typecheck, production build, token gates, and complete local suite coverage pass. Embedded PostgreSQL startup/socket failures passed in isolated retries with the canonical test environment; no expected behavior was weakened. The recovery regression added during the earlier full run passed in its final complete 75-case file. - Before the copy-only follow-up, all 56 CI checks passed on `c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no unresolved review threads. The final native-activity staging repeat passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`. - Verified on a separate staging instance: a real failed Daytona workspace export retains its board-owned repair action and displays “Recovery needed” in the task list. The repair card remains actionable without starting another provider turn. - Real staged export-only repair: the actual task list showed “Recovery in progress” while the original run had a positively identified running export operation, then Done after exact copyback of all 20,000 nonce-bound files. The accepted result and full provider session/turn/terminal envelopes remained unchanged. All 17 independent final checks passed; the browser downloaded the exact 19-byte result. The separate fixture was cleaned up with independent provider-absence verification. A control transport process restarted during repair; it did not submit another provider turn. - Final deployed-source repeat on `d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral Daytona allocation, actual browser export repair, and a saved full activity projection referencing the exact executing export operation with no scheduled retry. The task list showed recovery in progress, then Done; all 21 final checks passed, including 20,000 exact host files, unchanged provider provenance, and provider deletion only after committed copyback. The downloaded 19-byte result matched independently. This integrates #14334; no source change was required here. ## Risks - The label depends on the persisted recovery action. A separate runtime defect can still stop work; this change makes that condition visible. - Future recovery kinds must supply a valid continuation path before they can display progress. ## Model Used OpenAI Codex, based on GPT-6, with code execution, browser testing, and subagent tool use. The runtime does not expose an exact serving model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
01d9a12185 |
fix: make keyboard shortcut enablement a personal preference (#14141)
Store keyboard shortcut enablement per user and expose it in Profile settings. Co-Authored-By: Codie <Codie@users.noreply.github.com> |
||
|
|
ffa32373bc |
fix(runtime): honor provider acquisition timeout defaults (#14097)
Declare provider acquisition budgets so slow Daytona creation does not hit the host’s 30-second fallback. Bound creation and setup to one deadline and preserve scoped cleanup ownership after timeout. Legacy drivers retain their original call shape. Validated by 238 provider/manifest tests, focused database and heartbeat regressions, root and standalone provider typecheck/build, and full CI. Local broad tests also expose recorded Mac baseline limitations. Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
96bf004a79 |
fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bd6caf51bb |
fix: preserve restore failure results and stop unsafe retries (#14035)
Preserve agent output and earlier execution errors when workspace restore fails. Report the restore phase and confirmed saved-plan links. Require verified repair before retrying unsafe archives, while preserving approval states and the retry budget. Verified with full CI, 506 focused regression tests, and Greptile 5/5 with all review threads resolved. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e2f1a66aa7 |
feat: support default-hidden experimental settings (#13980)
## Thinking Path > - Paperclip is the open source control plane for AI-agent companies. > - Operators can hide settings that their users must not change. > - The server shares the effective restrictions with the UI and settings API. > - An explicit list must change whenever a new experimental flag is added. > - This pull request adds a wildcard with named exceptions to the existing setting. > - Core applies the policy to its current catalog, so new flags stay hidden automatically. ## Linked Issues or Issue Description Related: #11823 introduced settings visibility. #13907 added workspace isolation visibility. I searched related PRs and issues and found no duplicate wildcard implementation. **What existing behavior does this improve?** Operator control of experimental setting visibility through `PAPERCLIP_HIDDEN_SETTINGS`. **Subsystem affected** Shared settings policy and its existing server health and mutation consumers. **Current behavior** Operators must name every hidden experimental toggle. A new Core flag can become visible until the operator updates that list. **Proposed behavior** `instance.experimental.*` hides current and future experimental toggles. Entries such as `!instance.experimental.enableEnvironments` leave named controls available. Explicit hidden keys and the hidden parent page take precedence over exceptions. **Reason and benefit** Operators can maintain a short list of allowed controls instead of a second copy of Core's full feature catalog. **Breaking changes** Existing explicit lists and unset configuration keep their behavior. The new syntax is opt-in. Older images ignore it, so operators must retain explicit restrictions until those images are upgraded. Visibility does not change feature values. ## What Changed - Expand the wildcard into concrete catalog keys in the shared parser. - Limit exceptions to known experimental controls and preserve explicit restrictions in either input order. - Test a synthetic future catalog addition, duplicate and invalid entries, API rejection, same-value echoes, and the effective health payload. - Document the syntax and the transition for deployments with mixed image versions. ## Verification - Targeted parser, future-catalog, health, and settings-route tests pass: 95 tests across four files. - `pnpm -r typecheck` passes, including Rust checks, with the installed Cargo directory on PATH. - `pnpm build` passes. - All current-head CI gates pass, including the full test shards, Rust, build, browser E2E, and canary dry run: https://github.com/paperclipai/paperclip/actions/runs/36083292578. One unchanged runtime-exposure cold-start test passed on its first retry. - The full local `pnpm test:run` did not pass on macOS/Node 25: the first server group reported 13,356 passed, 18 failed, and 99 skipped, with six failed files (including two failed suite setups). Failures were in unchanged runtime/company skill cache, chat/email connector fixtures, embedded-Postgres setup, and workspace cleanup tests. A standalone filesystem probe reproduced the read-only-directory rename permission failure. Missing connector fixture paths, database startup failures, and two integration assertions also occurred; the remaining local groups were not reached after this group failed. The corresponding CI lanes all pass. These local failures are not claimed as fixed by this PR. - Browser suites were not run because this changes the shared policy, not UI rendering or browser workflows. Health payload and route tests cover the shared UI/API contract. ## Risks A malformed exception remains hidden and is reported as unknown. Exceptions cannot override an explicit hidden toggle or parent page. Older images ignore wildcard syntax; keep their explicit list during a mixed-version rollout. No schema or feature-value changes are included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact runtime variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8781f06a87 |
feat(connections): enable MCP aggregators by default (#13964)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Connections let those agents use external services with explicit access rules. > - Zapier, Arcade, Composio Connect, and Executor already have setup and runtime support. > - Their experimental switch still blocks discovery and setup by default. > - This pull request removes those gates and the Settings toggle. > - Users can connect these providers without enabling an experiment. ## Linked Issues or Issue Description Refs #13755. Refs #13941. **What existing behavior does this improve?** Apps browsing, inline setup, and agent connection search for the four MCP aggregators. **Current behavior** An instance must enable the MCP aggregators experiment before users or agents can start setup. **Proposed behavior** All four providers are available by default on local and managed instances. Old stored and managed values still parse but cannot disable them. ## What Changed - Remove the aggregator gates from Apps, inline setup, server setup, and agent search. - Remove the Settings toggle and its UI hook. - Retain the old setting key only for upgrade compatibility. Normalize it to true and ignore managed overrides, as Apps already does. - Replace opt-in fixtures with default-on coverage. Test old false values, all four setup flows, provider choice, and the removed toggle. - Update current connector guidance and remove the opt-in from the runner acceptance fixture. ## Verification - 306 focused tests passed across eight files: shared remote MCP contracts; server remote MCP lifecycle, aggregator fallback, settings normalization, and managed overlay; UI Apps browsing, setup, and experimental settings. - Server and UI TypeScript checks passed. - UI token gates and `git diff --check` passed. - The full local suite was not run, per the maintainer's instruction. All 54 CI checks passed; two checks were skipped. One unrelated workspace-preview readiness timeout passed on one failed-shard retry. - The setup fixtures use simulated MCP responses. This change does not claim new live provider acceptance. ## Risks - Existing instances now show all four providers, even if the old flag was false. This is intentional. - External provider choice, credentials, company isolation, agent grants, and tool policies still apply. Showing a connector does not authorize an external account. - No data migration is required. The compatibility key keeps old managed configuration documents valid. - Historical Zapier live acceptance remains incomplete in the existing evidence report. The maintainer explicitly requested the default-on rollout for all four existing providers; the report records that scoped exception. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, repository tools, and test execution. The context window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
aa8fc86331 |
feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external services. > - Native connections should remain the first choice for a supported app. > - Other apps may be available through an external MCP provider. > - The user must know which external provider handles the connection and choose it before setup. > - This pull request adds ranked alternatives and server-authored instructions to connection search. > - Agents can follow the returned instructions while Paperclip validates saved choices and access. ## Linked Issues or Issue Description Related: #13879, which fixed inline MCP provider setup. This PR adds discovery and provider selection on top of that work. **Subsystem affected** Cross-cutting: shared connection contracts, server search and intent services, native runtime, CLI, inline setup UI, and evals. **Problem or motivation** An agent cannot offer a clear external-provider choice when Paperclip has no native connection for an app. Adding provider-specific branches to the core prompt would make those instructions harder to maintain. **Proposed solution** Prefer a native connection. Otherwise return verified alternatives in Composio, Arcade, Executor, Zapier order. Include an external-service disclosure, a question with None, and the next instruction in the search result. Validate the saved human choice before creating a selected fallback setup card. Reuse existing provider accounts and verify underlying app access separately. **Alternatives considered** Do not silently choose a provider. Do not claim that broad execution tools prove support for every app. Reuse existing questions and connection intents rather than add another connection model. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback capabilities. The MCP aggregators experiment remains the gate. No duplicate provider-routing PR was found in the public search. ## What Changed - Add a dated support index and authorized cached-tool evidence for external routes. - Return provider questions and next-step instructions from `connections_search`. - Preserve pending choices and declines across continuation. Validate company, task, agent, human, app, and current route eligibility. - Carry the selected app into new setup and account reuse, validate explicit provider requests against persisted human messages, and distinguish provider readiness from app authorization. - Sync native, MCP, REST, and CLI contracts. Keep core agent instructions provider-neutral. - Add production-component Storybooks, focused database tests, and three real-agent browser eval cases. - Record the plan, observed failures, fixes, passing evidence, and acceptance limits. ## Verification - Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and all review threads resolved. - After rebasing on master `18dac1e1e`: 64 focused shared, validator, route-contract, and database tests passed; server typecheck passed. - Embedded-browser test drive on the rebased head: native Jira card, HubSpot external-provider question, Arcade account reuse, one actual MCP read against a local synthetic fixture, reload persistence, and None preventing further calls. A real OpenAI-backed agent performed discovery and continuation. - UX observation: the agent initially combined mutually exclusive request fields; the server rejected it and the agent recovered without changing access. This extra retry remains visible in the transcript. - After rebase: 23 focused eval grader/catalog tests, affected TypeScript checks, token gates, production UI build, and Storybook build passed. Full local tests are intentionally excluded at the maintainer's request. - Before rebase: four browser/real-agent attempts passed: native Jira, None, and reuse of the second provider on two Codex profiles. - Browser evals used an isolated deterministic MCP fixture through the real Paperclip gateway. They do not prove production compatibility with all four providers. - Review `Apps / Connections / Provider choice` in Storybook. Choose Arcade, continue through Access, and verify the app name, external-service disclosure, and URL configuration. - The detailed verification report is `doc/connections/2026-09-23-aggregator-routing-verification.md`. ## Risks - The public support index is finite and can age. Account capability and app authorization still require verification after selection. - Existing installed-tool permissions remain in effect. Provider choice is not a new execution permission boundary. An early Mini attempt skipped search; clearer provider-neutral instructions made the targeted rerun pass. This is not a measured reliability rate. - Explicit requests skip provider confirmation only when a clear persisted human message or saved provider choice supports them. Other phrasing falls back to confirmation; the agent query alone is not consent. These routes do not add tool permissions. - No database migration or legacy Composio broker is added. Real-provider acceptance remains separate from fixture proof. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser tools. The exact deployment variant and context-window size are not exposed in this session. Product evals separately used the repository's primary Codex and Codex Mini profiles; those agents supplied test behavior, not independent provider compatibility proof. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
18dac1e1ef |
feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path > - Paperclip manages AI agents and their work. > - Connections give agents governed access to external tools. > - Agents need durable memory across tasks and execution environments. > - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory tools. > - This pull request adds their setup flows behind an experimental toggle. > - It also delivers assigned MCP tools through the native remote Codex runner. > - Operators can connect a provider once and use the same governed tools locally or in Daytona. ## Linked Issues or Issue Description **Problem or motivation** The Apps catalog lacks a complete set of memory providers. Remote native Codex agents also need access to assigned managed MCP tools without receiving provider credentials. **Proposed solution** Add five memory connectors behind the disabled-by-default Experimental memory connectors setting. Use the existing connection setup and permissions UI. Default their tools to Allowed. Preserve company boundaries, operator permission changes, provider scopes, and audit attribution. **Alternatives considered** Direct provider credentials in each sandbox would duplicate setup and bypass the managed gateway. The remote runner instead uses its existing protocol channel to call the gateway on the server. **Roadmap alignment** This maintainer-requested experiment supports the Memory / Knowledge and Connected Apps roadmap areas. It adds provider connections without introducing a separate memory UI. Related connector authoring documentation is tracked in #13692; no duplicate memory-provider implementation was found. ## What Changed - Add provider definitions, official branding, and the experimental setting for all five providers. - Use OAuth for Zep and Supermemory, API credentials for Mem0 and Honcho, and a bundled Cloud API bridge for Cognee with no runtime downloads or subprocesses. - Default memory tools to Allowed and classify destructive actions explicitly. - Fix personal remote credential resolution and propagate provider tool errors. - Relay assigned managed MCP tools to remote native Codex through the runner protocol. Recheck current authority for each call and rotate stale tool contracts. - Add 21 Storybook states and complete OAuth walkthrough fixtures. - Document provider research, sanitized tool inventory, and live local and Daytona proof. ## Verification - Passed workspace typecheck: `pnpm -r typecheck`. - Passed production build: `pnpm build`. - Full local `pnpm test:run`: 13,273 passed, with failures from process/readiness timeouts under parallel load. Reran all 16 affected suites with one worker: 456 passed, leaving two macOS `/var` versus `/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`. No test failures remain unverified. Latest-head remote CI passes all 54 checks (two optional Storybook jobs skipped). Greptile is 5/5 with no unresolved findings. - Latest Cognee gateway regression: 71 passed, including public deployment without a runtime host and immediate recovery after a provider error. Bundled bridge tests: 25 passed. - Browser setup and real agent tasks exercised all five providers. Mem0, Cognee, Zep, and Honcho have successful store/retrieve proof. - Supermemory now has scoped read/write consent. Local storage and real Daytona write, document read, and semantic recall passed; indexing completion was verified before claiming success. - All five providers were exercised through a real Daytona sandbox and its native runner MCP relay. Provider credentials remained on the server. The final bundled Cognee bridge also passed a fresh Daytona store/recall run. All disposable sandboxes were removed and verified absent after testing. - Storybook is rebuilt and contains the experimental toggle, catalog, setup, permissions, and error states. The Zep and Supermemory access steps advance correctly. ## Risks - Provider OAuth scopes and plan limits remain independent of Paperclip tool permissions. An Allowed tool can still be rejected by the provider. - Providers can queue memory indexing; save acceptance does not prove that semantic recall is ready. - Remote tool contracts must stay synchronized with current connection authority. Regression tests cover revocation and stale contracts. - This change has no database migration. Existing connections remain usable when the experimental catalog toggle is disabled. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository edits, code execution, and browser automation. The exact context window size is not exposed in this session. Live acceptance agents used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b0155a681a |
feat(slack): connect Paperclip conversations and scheduled messages (#13920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack conversations use the same tasks and agents as the Paperclip board. > - A board reply must reach that Slack conversation and let the agent continue the work. > - An assigned agent also needs its Slack tools during normal tasks and scheduled routines. > - Both paths must keep the linked user's authority, delivery rules, and conversation history. > - This pull request adds those paths and reduces setup friction for Slack bots. ## Linked Issues or Issue Description **Subsystem affected** Server orchestration, Slack connector tools, shared contracts, and chat setup UI. **Problem or motivation** Replies entered in Paperclip did not provide a complete round trip to the linked Slack thread. Slack and board wakeups could select different model sessions for the same task. Agents also lacked their assigned Slack tools outside Slack-origin work, which prevented a routine from sending its responsible user a briefing. Inviting a bot could leave the new channel disabled. **Proposed solution** Mirror human board messages with author attribution and route the agent result to the same thread. Use the same session key across both entry points. Supply Slack tools to the connection's assigned agent in normal tasks and routines, using the current responsible user's verified link. Enable newly invited channels while preserving explicit disabled choices. Add a browser-agent setup prompt to the Slack wizard. **Alternatives considered** A separate Slack scheduler or task dispatcher would duplicate existing Paperclip workflows. Reusing the connection owner's identity would grant the wrong authority. Replaying old channel history could start unintended work. This change uses ordinary task wakeups, routine dispatch, and Slack's original invitation mention event instead. **Roadmap alignment** Extends the shipped Scheduled Routines and governed Apps capabilities. It does not add a separate task lifecycle. Related work: #13828 and #13809. Related test stabilization: #13877. The existing plugin Slack-control proposals are separate from this built-in connector change. ## What Changed - Queue human Paperclip messages for the original Slack thread with display-name attribution and stable delivery identities. Require the author’s current linked Slack identity and recheck access before delivering messages or agent replies. - Apply pause, dependency, cancellation, and closed-workspace guards before explicit Board sends request work and again when the durable outbox dispatches it. - Route agent results back to Slack and preserve model-session continuity, including replies that reopen completed tasks. - Resolve assigned Slack connections for normal agent tasks and routines. Recheck the responsible user's link, membership, and permissions at execution. - Add `slack_open_dm` for the responsible user's bot DM and request the `im:write` scope. - Enable newly discovered invited channels. Keep explicit OFF choices and normal admission and deduplication rules. - Add a copyable Slack setup prompt for a computer-use agent, with Storybook coverage. Share the prompt-button component with GitHub. - Update Slack tool documentation and runtime instructions. - Stabilize the mobile project browser test by waiting for the final canonical route before editing, preserving all persistence assertions. ## Verification - Live staging: invited the bot after the first mention. The channel became enabled and the bot answered that original mention. - Live staging: a normal Paperclip reply appeared in Slack with author attribution. The agent completed the calculation and replied once in the original thread and in Paperclip. - Live staging: a codeword entered in Slack was recalled from Paperclip. A following Slack calculation used the result from the Paperclip turn. Run metadata confirmed the same model session for both entry points. - Live staging: a scheduled routine used `slack_open_dm` and `slack_post_message` to deliver one DM. The existing app was reinstalled with `im:write`. The test routine was paused after verification. - Before the master merge: 397 focused feature tests passed. The continuity fix passed all 76 issue comment/update route tests and six focused route/integration cases. Typecheck, build, and token gates passed. - Review fixes: 47 focused integration cases passed, covering link revocation/replacement, private membership removal, guarded outbox dispatch, concurrent workers, lost scheduler responses, a real one-connection pool, exact reply provenance, and attachment retries. All 99 issue-comment route tests and the Slack catalog browser test passed. - Full local typecheck, production build, token gates, and module-boundary checks passed. The full local test command passed 25,915 tests before a 15-second timeout in `issue-thread-interaction-routes.test.ts`; that entire suite passed on isolated rerun (81 tests). Remaining serialized coverage is provided by the current-head CI shards. - An unchanged Cursor adapter test hit its 10-second limit in CI; all five tests in that file passed on a local rerun in 3.11 seconds, and the failed CI shard passed on its single retry. - The preview-server readiness test passed a local rerun (28 tests). The mobile-project readiness fix passed three repetitions of both browser tests (6/6). - Final commit `64ac0d9897f4353375996f1b1b38e5040bdeb0a0`: all CI gates passed, including all eight browser shards, all server/chat suites, typecheck, build, runner checks, and security checks. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - Human messages on a Slack-linked task now publish to its Slack thread. The task banner states this behavior. Incoming Slack messages and internal agent bookkeeping must not echo back. - Normal tasks and routines can now use the assigned bot. Authority remains bound to the current responsible user's link; it does not fall back to the connection owner. Revocation, private-context limits, and queued-write checks still apply. - Existing Slack apps need `im:write` and a reinstall to open DMs. Other existing capabilities remain available without that scope. - New invited channels default to enabled. Explicit disabled choices remain disabled. Channels created by bot tools still require a person to enable responses. - No database migration or new provider credentials are required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, GitHub CLI, and browser tools. The runtime does not expose a more specific model build identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
94a0aa7726 |
fix(ui): hide workspace isolation controls for managed hosts (#13907)
## Thinking Path > - Paperclip manages AI agents and their work. > - Workspace isolation keeps task checkouts separate. > - Managed hosts can enable isolation and hide its experimental toggles. > - Project, task, and routine forms still expose choices that override that policy. > - This pull request adds an operator visibility key for those controls. > - Workspace access stays available, and execution keeps its existing policy. ## Linked Issues or Issue Description **What existing behavior does this improve?** Operator control over workspace isolation settings in the UI. This follows the settings list cleanup in #13905. **Current behavior** Hiding the experimental isolation toggles leaves project policy editors, task selectors, routine and pipeline overrides, recovery actions, and workspace configuration visible. **Proposed behavior** Set `PAPERCLIP_HIDDEN_SETTINGS=workspaces.isolation` to hide these controls. Keep workspace navigation, status, files, and runtime access. Hide experimental toggles separately. Instances that do not set this key keep their controls. **Reason and benefit** Users on managed hosts should use the host's isolation default. A hidden form must not submit a stale draft that overrides it. ## What Changed - Add the UI-only `workspaces.isolation` key to the shared visibility registry. - Hide project workspace policy, task and subtask selectors, routine and pipeline overrides, and isolated re-issue actions. - Hide the workspace Configuration tab and redirect direct links to workspace issues. - Omit hidden new-task and routine overrides. Keep explicit task/subtask workspace launch context, saved policies, and automatic branch values for workspace routine runs. - Wait for the health visibility policy before showing controls. Keep workspace access and all execution APIs available. - Document the key and test visibility, form payloads, deep links, and unchanged workspace access. ## Verification - All 52 CI checks pass on `50cc770d33` (two expected skips). The branch is mergeable. Greptile is 5/5 with no unresolved comments. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm check:token-gates` passed. - Targeted UI checks passed: 346 tests across 14 suites, including hidden project/task controls, stale task drafts, routine branch defaults, recovery actions, configuration deep links, and workspace access. - Shared settings-visibility tests passed: 11 tests. - `pnpm test:run` was run and stopped after reproducing five failures in unchanged server tests: two `chat-channels.integration` cases (linked-request provenance and direct external-chat finals) and three `company-skills-service` cases (runtime refresh, concurrent download, and explicit update). The earlier local run for #13905 showed the same failures. The full local suite is not claimed green. Targeted UI/shared checks pass. All PR CI shards, including the affected chat and skills suites, pass. - Reviewed the diff for secrets, private links, and run artifacts. ## Risks This key changes UI visibility only. It does not reject API calls or change feature values. Operators must enable isolation and its default through their existing policy mechanism. Older app versions ignore the new key until upgraded. Removing the key restores the controls. No schema changes. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally; targeted checks pass (full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
30182c704c |
feat(telemetry): propose connector telemetry events behind proposal markers (#13582)
Adds connection.created / connection.updated / connection.invoked as proposed-telemetry wrappers (unregistered; the TelemetryClient cannot queue or send them) plus the server-side emitter wiring and tests. No telemetry is transmitted by this change. Proposal: #13578. Coordinated release plan and review package: PAP-24 (Step A.1). |
||
|
|
24429024e7 |
feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps gives agents governed access to external tools through stored credentials. > - Routines start work when an external service sends an event. > - Fireflies provides meeting transcripts and summaries through an official hosted MCP server. > - This PR adds that connection and accepts signed meeting events through the shared app webhook flow. > - Agents can review completed meetings with the same permissions and audit records as other work. ## Linked Issues or Issue Description **Problem or motivation** Operators need agents to read Fireflies meetings and start follow-up work when a summary is ready. The Apps catalog lacks Fireflies. The shared app webhook flow needs to accept its signed deliveries. **Proposed solution** Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer API key. Extend the existing Another app or script flow with signed webhook support. Verify the raw-body signature and pass the JSON payload as external data. Select Meeting Summarized in Fireflies. Deduplicate identical signed deliveries, including setup deliveries. **Alternatives considered** A separate REST connector would duplicate the governed MCP path. Polling, legacy V1 payloads, and automatic provider-side webhook registration are outside this change. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps and Scheduled Routines surfaces. It adds a provider to those systems. It does not introduce a second integration framework. **Additional context** A GitHub search found no existing Fireflies issues or PRs. Provider references and verification limits are in `doc/connections/FIREFLIES.md`. ## What Changed - Add the official Fireflies catalog definition, generated registry, provider evidence, and branded artwork. - Reuse Access → Connect, dynamic discovery, Permissions, vault storage, policy, and audit behavior. - Classify Fireflies sharing, movement, and access revocation as writes. - Preserve Off and Ask first restrictions during OAuth reauthorization and API-key replacement. New actions retain normal defaults. - Add `app_webhook` authentication to the shared Another app or script flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier `fireflies_hmac` triggers and revision snapshots for compatibility. Existing text columns need no migration. - Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact request body. Preserve generic event payloads and deduplicate identical signed requests. - Keep the routine wizard generic. Show one webhook URL and secret in Another app or script. Keep all new app webhook event names provider-neutral. Keep provider setup instructions in the connector documentation. - Pass generic webhook JSON to the task in an explicit external-data block, capped at 16,384 characters. Keep strict meeting validation for existing legacy Fireflies triggers. ## Verification - Feature implementation commit `0882dc8a1`: all 54 CI checks passed; two conditional Storybook checks skipped. This includes full tests, typecheck, build, browser E2E, canary dry run, and security checks. Greptile rated this commit 5/5; all review threads are resolved. - Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed on the final code. Targeted connector, gateway, webhook, revision, and UI suites passed during implementation. After the provider-neutral follow-up, all 84 app-webhook and routine-service tests passed; the final payload-to-task assertion also passed in the 72-test routine suite and a clean-config rerun. - The long local `pnpm test:run` invocation started before the final edits and was stopped after the final-commit CI suites passed. It reported one generic webhook test failure while those files were changing; that test and the entire routine suite passed on the final source, including a clean-config reproduction. The interrupted local run is not counted as a full-suite pass. - In the embedded browser, completed official OAuth consent and discovered 20 live actions. Real meeting listing, transcript retrieval, and summary/action-item retrieval succeeded as the selected agent. Turning a live read Off blocked its test; catalog refresh preserved the restriction. - Embedded-browser Another app or script setup, back/save/resume, narrow layout, and a signed synthetic Fireflies delivery succeeded. The UI reported authentication passed without creating a task. Fixtures cover signature tampering, malformed requests, ordinary app event names, duplicate/setup deliveries, rotation, revisions, pause/archive, and company isolation. - Existing MCP browser suite: 8 passed and 2 provider-dependent cases skipped. Branding checks passed; connector artwork and webhook setup were checked at desktop/mobile widths and in light/dark modes. - An unauthenticated POST to a correctly formatted public webhook URL reached the staging tenant verifier through the existing Cloud gateway. - A real Fireflies webhook delivery remains unverified. A staging callback is available for the operator walkthrough. Live API-key authorization, credential expiry, and a new meeting's summary completion were not tested against the provider. Fixtures cover these protocol and lifecycle paths where applicable. - Storybook follow-up `c54174faa`: 27 production-component stories cover every UI change, with a source-to-story map in the connector documentation. Static Storybook build, UI typecheck, token gates, and Playwright checks for all stories and the mobile footer pass. All PR checks passed for this Storybook follow-up; Greptile reviewed `c54174faa` at 5/5. ## Risks - Fireflies may change its hosted MCP tools or OAuth behavior. Tool discovery stays dynamic. Experimental search/fetch tools are not required. - Public webhook setup requires HTTPS and a separate signing secret. Fireflies normally emits events for meetings owned by the configuring account. - Reauthorization touches shared MCP permission code. Regression tests cover existing restrictions, new actions, connection removal, and other gateway callers. - Webhook receipt grants no tool access. The routine agent still needs an authorized Fireflies connection. - New generic triggers rely on provider event subscriptions. Without a sender-supplied idempotency key, changed request bytes count as a new event. Existing legacy Fireflies triggers retain summary-only filtering and per-meeting deduplication. ## Model Used OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing, code execution, and embedded-browser testing. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7944ed3d97 |
fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Native runner agents need governed tools, durable runtime state, and useful execution evidence. > - A first activity trace waited 53.467 seconds even though tool activity took 6.274 seconds; provider input arrived before the server API call executed. > - Native agents also need a safe way to hire teammates without asking the model to rebuild runtime configuration. > - This pull request separates the observed ACP input-stream window from the actual server `tool.execute` span and adds a server-owned native hire contract. > - The benefit is clearer latency evidence and safer native teammates with existing approval, auth, and company boundaries preserved. ## Linked Issues or Issue Description Related Daytona provenance work is in [#13814](https://github.com/paperclipai/paperclip/pull/13814). No duplicate public PR was found for this combined timing and native-hire change. **What existing behavior does this improve?** Native runner agents can use governed tools and request hires. The server did not expose a safe native hire operation that reused the caller's validated runtime settings. First-activity traces also mixed provider input timing with server tool execution timing. **Current behavior** A native hire must construct a separate runner configuration. Full configuration copying could expose paths, instructions, secrets, or sessions. Timing evidence could make a provider or MCP identity join appear proven when the trace did not contain that join. **Proposed behavior** The native `hire_agent` operation accepts identity and persona inputs. The server sends `adapterType: "paperclip_runner"` with `inheritRuntimeFrom: "caller"`, then copies only validated provider, model, permission, lifecycle, and bounded execution settings. It inherits and validates the default environment, derives the managed AI binding through existing normalization, preserves approval and permissions, and creates fresh child instructions. Caller secrets, paths, prompts, and sessions are excluded. Provider events now include the optional boolean `inputUpdated`, with Rust forwarding support. Timing evidence separately records the ACP input-stream window and the actual server `tool.execute` activity. It does not claim a provider or MCP join without matching evidence. **Reason and benefit** Native agents can hire teammates that start with the caller's approved execution policy. Operators retain company boundaries, auth rules, approval gates, and requalification. Reviewers can distinguish provider streaming time from server API execution time when diagnosing first-activity delays. **Breaking changes** None for existing hires or tool calls. `inheritRuntimeFrom` is optional and only applies to same-company native agent callers. Conflicting explicit runtime settings are rejected. The provider event field is optional for existing producers. ## What Changed - Added the native `hire_agent` protocol action, catalog entry, API contract, and runner authority checks. - Added `inheritRuntimeFrom: "caller"` validation and a closed native runtime inheritance allowlist. - Preserved managed AI binding normalization, default-environment validation, approval snapshots, permissions, requalification, and fresh child instructions. - Added provider `inputUpdated` schema support and Rust forwarding. - Added first-activity and server tool timing evidence with conservative identity-join handling. - Added route, authority, provider-event, sidecar, API, catalog, and Rust-focused tests. - Kept private Honeycomb links, raw traces, and local result paths out of this description. ## Verification Focused checks passed: - 458 timing/session checks. - 61 native hire inheritance checks. - 20 hire authority checks. - 1,741 API checks. - 106 catalog checks. - 54 provider sidecar checks. - 12 Rust provider checks. Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2). R1 stopped at missing Docker image setup. The final trace is available at https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces. Latest-head CI passed all required build, typecheck, Rust, static, Vitest, serialized-server, workspace, chat, and E2E jobs. The focused local checks listed above passed; the broad local suite was not run before the live evaluation, while CI provides the full repository verification. ## Risks - Timing fields describe separate observed windows. They do not prove a provider or MCP owner without a valid trace join. - The inheritance allowlist must stay synchronized with native runner configuration fields. - Approval snapshots include resolved safe inherited settings and should be reviewed when native configuration fields change. - The focused local suite is narrower than the full repository suite; latest-head CI covers the broader repository checks. > Roadmap review: `ROADMAP.md` places this work within Paperclip's bring-your-own-agent direction. It extends existing native runner hiring and observability behavior. ## Model Used OpenAI GPT-6 (exact serving model ID is not exposed), with extended reasoning and repository tool use; GPT-5.6 Luna assisted with focused implementation and verification work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7badae6981 |
feat(plugins): add an optional organization switcher slot (#13832)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Plugins can add UI surfaces to the board. > - Organization navigation is still fixed in the host sidebar. > - A distribution needs a supported way to supply its own organization menu. > - This pull request adds one optional React slot with host-owned navigation controls. > - The built-in menu stays available when the optional contribution cannot render. ## Linked Issues or Issue Description **Subsystem affected** Plugin SDK, server capability validation, and sidebar UI. **Problem or motivation** An installed plugin cannot replace the organization switcher without editing the host menu. Existing sidebar and overlay slots do not provide this replacement surface. **Proposed solution** Add an `organizationSwitcher` slot that requires `ui.sidebar.register`. Pass display state, an icon renderer, and navigation/logout callbacks. Keep the built-in menu for absent, ambiguous, missing, failed, or unsupported contributions. **Roadmap alignment** This keeps distribution UI in plugins and adds a small host contract. It does not add an account system or change company authorization. The maintainer requested this extension. Related prior menu changes: #12788, #10917, and #10850. No duplicate replacement-slot PR was found. ## What Changed - Add the slot to shared validation, SDK types, and server capability checks. - Wrap both sidebar menu variants with the optional replacement. - Resolve selection against the current account query before mounting, and reset replacement state on account or company changes. - Add host-specific component props and a fallback to `PluginSlotMount`. - Document the React-only contract and its trust boundary. - Report runtime-supervisor fixture startup details when CI readiness fails. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - `pnpm check:token-gates` passed. - `pnpm exec vitest run --project @paperclipai/ui`: 6,580 tests passed. - Targeted manifest, replacement, and built-in menu tests: 26 passed, including the incoming-account selection regression. - Installed a local test contribution into an isolated server. CLI inspection reported `ready`. Browser checks covered the loaded production UI, keyboard dismissal, current-organization selection, and an expired remote session. Remote account responses were fixtures. - `pnpm test:run` was attempted. Its first server phase passed 8,323 tests but failed in 36 files due to embedded PostgreSQL startup and filesystem permission errors on this Mac. Later phases did not run. The current Linux CI run is green: 54 checks passed and two were skipped, including build, typecheck, and browser gates. See https://github.com/paperclipai/paperclip/actions/runs/35796722770. - The earlier runtime-supervisor readiness failure did not reproduce locally. The complete affected shard passed locally: 58 files and 843 tests. The six supervisor tests also passed on Node 24.21.0 with CI flags. Added fixture startup diagnostics for the selected Node executable and listener port. The affected Linux shard then passed all 843 tests. The original root cause remains unconfirmed; no production runtime behavior or timeout was changed. ## Risks - Plugin UI remains trusted same-origin code. Display props do not authorize account requests. - A replacement can change navigation behavior. The host retains the built-in menu when discovery or rendering fails and keeps logout/session cleanup host-owned. - No database migration. Existing menus and portfolio behavior remain available. - Local full-suite verification is limited by the environment failures listed above. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser testing. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a10702a878 |
feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people start and continue agent tasks from other services. > - A Slack conversation needs access to its surrounding discussion and Slack collaboration tools. > - The agent must use the linked requester's access and keep private material within its permitted audience. > - This pull request adds Slack tools through the existing connector contribution and approval framework. > - People can ask an invited bot to read a discussion, create follow-up tasks, and collaborate in Slack. ## Linked Issues or Issue Description **Subsystem affected** Chat connectors, connector runtime, tool gateway, and connection Settings/Access. **Problem or motivation** Slack-origin tasks can receive messages but cannot inspect the rest of a channel or act through the originating bot. People must paste context or configure a separate integration. **Proposed solution** Supply typed Slack tools and a bundled skill only to the originating task and assigned agent. Resolve the linked requester on the server. Check bot and requester access before reads and writes. Use existing durable actions and approvals. Retrieved messages remain source material. **Alternatives considered** Slack's user-OAuth MCP server does not replace the customer-created chat bot. An unrestricted Web API proxy would not provide suitable permission or publication boundaries. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps work with a provider contribution. It does not add a task dispatcher or a separate Slack task lifecycle. Related: #11144 covers generic per-user MCP grant execution; this change binds Slack bot operations to chat-origin tasks. ## What Changed - Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and shared native/HTTP execution. - Bind tools to company, endpoint, task, run, assigned agent, and admitted linked requester. Check membership and revocation on each call and before queued writes. - Add paginated reads, bounded history search, source links, messages, file uploads, reactions, pins, bookmarks, topics, canvases, lists, and approved channel operations. - Restrict private-source publication, including automatic replies and uploaded deliverables. Keep other people's bot DMs inaccessible. - Reuse action receipts, idempotency, approvals, and reconciliation. Suppress an identical explicit-send/final-reply duplicate. Return governed results through their verified originating conversation. - Add endpoint-bound personal search OAuth storage and lifecycle. Keep native real-time search disabled until a runtime meets Slack's transient-result requirements. Current runtimes use bounded history search. - Show capabilities, scope upgrades, and personal search authorization in Settings/Access and Storybook. Document provider and runtime limits. ## Verification - Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no unresolved findings. One unchanged rapid-callback timing test passed on a single CI retry. - Approval presentation regressions cover board-comment precedence and exact Slack publication; the expanded database assertion passed in CI. The local PostgreSQL startup probe later became unavailable, so that final assertion was verified in CI. Slack setup and failed-run retry browser tests also passed locally. - Full workspace typecheck and build passed. Server typecheck/build passed again after the approval routing fix. - Broad local suites passed in separate groups: server 12,958 tests, UI 6,555, shared 770, skills catalog 20, and other workspace packages 2,652. CLI and serialized server checks passed after environment/timeout retries. These are composite results, not one uninterrupted green full-suite invocation. - PostgreSQL authority regression covers admitted identity, cross-company/task/agent rejection, recovery, retained-session revocation, OAuth refresh/disconnect races, approval execution, exact publication lineage, retries, uncertain sends, and duplicate suppression. - Gateway/response regressions cover separate-origin approval batches and durable continuation. Focused provider, access, search, native runtime, route, and AgentMail regressions pass. - Storybook capability, missing-scope, OAuth configuration, authorization, and disconnect states were inspected in the browser. - Live staging: read a channel decision and full thread, create exactly two assigned backlog tasks, add a reaction, paginate discovery to exhaustion, and return bounded search matches with source links and coverage. - Live staging: create/edit/read a canvas and list, inspect the canvas in Slack, post/edit one message, and create a channel only after approval. New channels remain disabled for responses. - Live staging: read a response-disabled channel from the requester's DM; writes to that channel were denied. The test setting was restored. - Final live retest passed: explicit file upload and exact content read-back; approved deletion of only the disposable bot message; continuation confirmation returned to the original Slack thread without repeating the action. - Optional OAuth, private multi-user boundaries, native RTS, and CLI provider execution are not fully live-qualified. The staging agent initially supplied malformed tool arguments; valid arguments succeeded, and the tool/skill descriptions now emphasize UUID write keys. ## Risks - Existing Slack apps must add scopes and reinstall for new capabilities. Provider plans and document permissions can still restrict operations. - Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET` for governed tool actions. The staging instance was configured with explicit operator approval; fleet provisioning is a separate gap. - Native RTS is not exposed on current transcript-retaining runtimes. Bounded history scans are deliberately reported as incomplete. Inline file reads support text/canvas content up to 256 KiB; other types return metadata. - Private document edits fail closed when the full audience cannot be verified. Uncertain effects other than posts/uploads require inspection instead of blind retries. - Shared approval-delivery code now separates outcomes by source run to preserve origin boundaries. No database migration is required. - A separate completion-validator gap remains when the agent cites a prior run's registered artifact during finalization. It asked for registration again even though Slack delivery was confirmed. This change does not add a connector-specific task-completion policy. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and browser testing. The exact deployed model identifier and context-window size were not exposed in the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a959e47508 |
fix(apps): reduce Google Chat scopes and block unread filters (#13820)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents controlled access to external services. > - Google Chat uses OAuth profiles and reviewed MCP tools. > - Those profiles request membership and read-state access that the supported feature set does not need. > - Removing read-state access also requires us to block unread search filters, including on existing connections. > - This pull request reduces both OAuth methods and enforces the reduced search contract before dispatch. > - Users retain conversation lookup, message history, ordinary search, and approved message sending. ## Linked Issues or Issue Description **What happened?** Both Google Chat profiles request membership and read-state scopes. The supported tool set does not include membership listing or read-state updates. Message search still advertises an unread filter. Related work: Refs #12619. **Expected behavior** Managed and customer-owned OAuth request only the scopes needed for supported features. Unsupported unread filters fail clearly before any provider call. Existing cached catalogs and broader grants must not bypass that policy. **Steps to reproduce** Start Google Chat OAuth from either connection method and inspect the requested scopes. Inspect the message-search tool schema, then submit a search with `searchParameters.isUnread` set to true or false. **Paperclip version or commit** The scope change is based on master at `110d176fc`. **Deployment mode** Managed Cloud and self-hosted instances with Google Chat Apps enabled. ## What Changed - Remove `chat.memberships.readonly` and `chat.users.readstate.readonly` from shared profiles and all four Chat connection methods. - Hide unsupported read-state fields and instructions in agent and board Test tool schemas. - Reject explicit unread filters, including false, null, snake-case fields, and encoded filter objects, before provider dispatch. - Recheck previously approved calls and support existing profile-bound and URL-only Chat connections. - Add scope, signed broker request, OAuth URL, allowlist, schema, and dispatch regression tests. - Document coordinated app/broker rollout, existing-grant reconnects, and the remaining deployment checks. ## Verification - All seven focused OAuth and Chat gateway test files pass: 498 tests on the rebased branch. - `pnpm -r typecheck` and `pnpm build` pass, using pinned pnpm 9.15.4. - The complete sharded CI test matrix passes, including general server, Chat, workspace, serialized server, Runner, and all eight browser e2e shards. The duplicate unsharded local `pnpm test:run` was stopped after CI passed; it did not complete locally. - `git diff --check` passes. - `node scripts/ingest-app-definitions.mjs` succeeds and leaves the branch unchanged. Google Workspace JSON is the durable reviewed input used by the generator. - All current-head CI checks pass at `7ee755371714dc036fff1c7da844776fee3f2ec1`, including typecheck, build, and canary dry run. Greptile is 5/5 with zero unresolved threads. The generator concern was withdrawn after review of the source and regeneration evidence. - No production deployment or live Google consent test was performed. After coordinated deployment, verify reduced consent scopes, normal search/history, message sending, and rejection of unread filters. ## Risks - Coordinate deployment with the companion Cloud broker scope change. Mixed versions can reject exact-scope requests. - Existing tokens are not narrowed or revoked. Grants with old scopes need new consent. Do not revoke a shared Google client to migrate one profile. - Explicit unread filters now return an error instead of being sent to Google. Ordinary search and the approved send tool remain available. - No database, UI, lockfile, or workflow changes. ## Model Used OpenAI Codex, a GPT-5-based coding agent, with tool use and code execution. The exact runtime model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8725d6ce09 |
fix: make answered Slack conversations idle (#13809)
Settle published, successful Slack turns as Idle; resume the same conversation on an admitted message. Preserve unfinished work, delivery errors, and explicit dispositions. Verified through focused lifecycle/API/UI tests, full CI, and a real staging Slack conversation in the embedded browser. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6de50ba594 |
fix(sentry): carry the deployment environment to the browser (#13784)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators can enable Sentry for the server and the signed-in browser. > - The server SDK reads `SENTRY_ENVIRONMENT` from the process environment. > - The browser receives its DSN through the session response, but receives no environment. > - A browser in staging therefore reports errors under the SDK's production default. > - This pull request passes the configured environment through the existing session and monitoring gate. > - Browser errors then identify the deployment environment while preserving the existing privacy settings. ## Linked Issues or Issue Description **What happened?** With `SENTRY_ENVIRONMENT=staging`, browser exceptions are tagged `production`. This can send errors to the wrong environment's alerts and makes deployment follow-up unreliable. **Expected behavior** The browser uses the server's configured Sentry environment. A reused image works in either staging or production. A signed-out browser still sends no events. **Steps to reproduce** 1. Configure a frontend Sentry DSN and `SENTRY_ENVIRONMENT=staging`. 2. Sign in and capture a browser exception. 3. Inspect the event environment. Before this change, it is `production`. **Paperclip version or commit** Reproduced on `a3749aac4680a901fa0fe1cc898907887abc9908` with the real browser SDK and a local test transport. **Deployment mode** Authenticated server and browser with optional Sentry monitoring enabled. No duplicate environment-attribution issue or pull request was found in the targeted GitHub search. ## What Changed - Add `sentryEnvironment` to the authenticated session response and shared schema. The optional field supports a newer browser reading an older server response. - Pass the environment to the browser SDK. An environment change restarts the client through its existing serialized lifecycle. - Cover environment attribution with a real SDK event, session authorization, unchanged-session refetches, environment changes, and legacy responses. - Document configuration and compatibility. Keep the loaded bundle's release identity and existing privacy filters. ## Verification - The regression test emits `production` for a requested staging environment before the fix. - Focused route, schema, browser lifecycle and real-SDK tests: 69 pass. - UI and shared-package typechecks, direct server `tsc --noEmit`, and token gates pass. - Full `pnpm build` and `pnpm -r typecheck` were attempted. Both stop at the Runner Rust step because `cargo` is absent on this machine. - Complete UI suite: 6,540 tests pass in 626 files. - Full `pnpm test:run`: 8,210 passed, 14 failed, 4,753 skipped; 36 files fail due to embedded PostgreSQL startup/cleanup and macOS runtime-cache `EACCES`. These match the existing local baseline; none touch the changed behavior. - Greptile: 5/5, no unresolved review threads. Linux CI has passed Build, Typecheck + Release Registry, and the completed test jobs so far. Remaining jobs are running or queued: the AWS runner provisioner is retrying EC2 CreateFleet `InternalError` responses. Full results will be recorded before merge. ## Risks Low risk. This adds one optional session field and changes Sentry attribution only. No migration or new monitoring opt-in is introduced. Missing settings keep the browser SDK default. Agent and unauthenticated requests still receive 401 without monitoring settings. Existing loaded browser bundles keep their old behavior until refreshed. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, code editing, and test execution. The session does not expose an exact model snapshot or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass (focused and full UI suites pass; full-root environment failures documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8813a50105 |
feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path > - Paperclip manages agent work as tasks and runs. > - GitHub chat brings repository conversations into those tasks. > - A review bot needs the assigned agent, its authority, and governed provider tools. > - The existing channel connection did not supply that review workflow or a complete setup journey. > - This pull request adds GitHub App setup, account access, event prompts, task-bound review tools, and exact-commit checks. > - Operators can inspect each review through the same task, run, and activity systems. ## Linked Issues or Issue Description **Subsystem affected** GitHub chat, governed connection tools, task execution, shared/database contracts, and connector setup UI. **Problem or motivation** Operators need a GitHub review bot that runs their assigned Paperclip agent. Mentions and PR events must preserve task ownership and requester authority. Provider publication must use the bot App identity and enforce the configured permissions. **Proposed solution** Extend the existing GitHub chat connector with resumable App onboarding, linked-member and sponsored-guest access, editable event prompts, and governed review operations. Validate structured assessments on the server and compute a stable Paperclip Review check for the exact head commit. **Alternatives considered** A separate review scheduler would duplicate Paperclip execution and permissions. Reusing personal GitHub credentials would change the bot identity and credential boundary. **Roadmap alignment** This extends the existing Connected Apps and governed-tool infrastructure. The project owner requested and approved this design. Related PR #8645 imports external Codex review feedback; this change runs an assigned Paperclip agent and publishes its results through the existing chat connector. ## What Changed - Include the current Paperclip instance origin in the copied setup prompt. Storybook uses its configured Paperclip origin; callback parameters and URL credentials are excluded. - Add a Claude/Codex copy button in the real setup and Storybook opening step. Its detailed prompt asks four setup questions and guides embedded-browser setup, verification, and optional required checks. Clipboard failure exposes selectable instructions. - Add a tutorial that explains why App installation, review scheduling, and required checks are separate choices. - Add manifest registration, an existing-App path, separate installation and repository selection, repository refresh, and explicit account confirmation. - Add low-trust agent guidance, effective capability verification, member selection, and explicit restricted guests with a sponsor. - Add configurable PR events, prompts, repository overrides, rating thresholds, and separate formal-review permissions. - Give the assigned agent governed App tools to read PRs, comment, begin an assessment, submit findings, and optionally submit a formal review. - Bind review history, root PR events, and inline replies to ordinary tasks. Deduplicate deliveries/findings and reject stale publication. - Link check Details to the underlying task on the current trusted hostname, or to Reviews before task creation. - Add schema migration 0283, API contracts, production UI, and 49 interactive Storybook states. - Repair local lease recovery. Keep the Cloud Dockerfile identical to master; no provider-pack layer or runtime-default environment variable is added. - Retry only rolled-back wake-admission transactions after transient endpoint-lock contention. A deterministic held-lock regression proves one accepted wake. ## Verification - Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero diff against master. Final workspace typecheck and build passed. The new PostgreSQL migration regression passed and preserves existing relation and constraint identities after replay. - Greptile reviewed this exact head at 5/5. There are zero unresolved review threads and no merge conflicts. - All current-head checks are green: 54 passed and two conditional Storybook jobs skipped. This includes complete server/workspace test suites, build, typechecks, policy checks, Runner suites, browser suites, and security status. One timing-sensitive callback-ordering test passed in isolation and its CI shard passed one retry. The duplicate local full-suite run was stopped after CI completed; it is not counted as a local full-suite pass. - Before the final Slack rebase and migration renumbering, 186 focused GitHub tests, 14 native bootstrap cases, token gates, and Storybook build passed. The final rebase retained the new Slack communication guidance. - The embedded-browser setup test copied the full detailed prompt, including the configured Paperclip instance URL. Desktop and narrow layouts were checked. Component tests cover successful copying and clipboard failure with selectable text and retry. - Live local and hosted GitHub acceptance evidence refers to application revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks exercised issue mentions, automatic PR reviews, inline findings, repeated mentions, task continuation, and failing-to-passing checks after a push. The Storybook agent generated, built, and browser-rendered pages; missing acceptance text failed, matching text passed, and broken JSX produced an incomplete result. - Live cases also covered independently disabled push events, prompt injection, duplicate signed deliveries, rapid pushes, stale-result rejection, finding deduplication, and restart recovery. Formal reviews were denied while disabled and published only after explicit enablement. Check Details links pointed to the underlying task on the trusted hostname. - Those hosted native Claude runs used the provider-pack layer now removed from this PR. They do not prove native Claude works on the standard Cloud image. A replacement hosted native Codex run is not yet verified: the disposable QA tenant has only an Anthropic AI connection. No new staging or production deployment was made for the packaging removal. - Required-check merge enforcement could not be tested because the private disposable repository's GitHub plan rejected the rules configuration. Published success/failure/incomplete check states were verified directly. ## Risks - Latest master allocated migration 0282 to Slack. The GitHub migration is regenerated as 0283 with replay-safe table/index/constraint creation; a PostgreSQL regression verifies existing relations and constraints are preserved. Existing preview tenants remain subject to the fleet migration-history compatibility preflight; no bypass is introduced. - Migration 0283 adds company-scoped configuration, registration, review, and publication records. Existing connections retain their behavior until reviews/tools are enabled. - Signed webhooks and expiring registration state remain required. Hosted installations also need the companion narrow Cloud gateway exemptions. - Agent assessments can be incomplete or wrong. The server enforces coverage/result structure, current-head publication, rating policy, and separate formal-review permission; it does not replace code-review judgment. - No Cloud image packaging changes are included. Remote native ACPX/Claude and OpenCode retain their existing operator-supplied provider-pack prerequisite. Native Codex and Codex with managed MCP tools do not require that pack. Earlier staging deployment evidence refers to its stated revision, not this packaging-removal head. Production rollout and merging remain outside this change. ## Model Used OpenAI GPT-6 through Codex, with repository, code execution, API, and embedded-browser tools. The exact serving model ID and context-window size were not exposed by the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d9b3a5653e |
feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people use the same tasks and agent tools from external conversations. > - Agents need communication guidance that fits the conversation medium. > - That guidance belongs in the original task context, without repeated instructions on each turn. > - Connection owners also need clear settings and a consistent way to remove a connection. > - This pull request adds initial Slack guidance, optional connection instructions, and chat connection menus. > - The benefit is clearer Slack replies with the existing Paperclip workflow and permissions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent replies in Slack and chat connection management in the Apps catalog. **Current behavior** Slack tasks do not carry a saved communication profile. The catalog shows a separate Manage button and does not offer removal on every chat connection row. **Proposed behavior** Save Slack guidance when a new conversation creates a task. Restore that original guidance when a model session is rebuilt. Do not append it to ordinary follow-ups. Expose optional additional instructions in Slack Settings. Put Manage and Remove connection in a three-dot menu for all chat providers. Keep Finish setup visible for drafts. **Reason and benefit** Small answers fit in Slack. Substantial deliverables use ordinary document or artifact tools with a useful Slack summary. Connection settings apply to new tasks and cannot change permissions. Users can remove both active and unfinished chat connections from the catalog. **Breaking changes** Two additive database columns store endpoint preferences and the initial conversation snapshot. Existing endpoints default to empty preferences. Existing conversations keep their original behavior. Non-Slack guidance is unchanged. Related public context: https://github.com/paperclipai/paperclip/pull/13741 improves native chat recovery. This change adds communication context to those existing execution paths. A search found no duplicate communication-guidance PR. ## What Changed - Add a provider-guidance registry, enabled for Slack first. - Persist optional endpoint communication instructions and capture an immutable snapshot when a conversation creates a task. - Resolve guidance from the verified company-scoped connection. Restore it for fresh native and legacy sessions without per-turn reminders, extra model calls, or extra context queries. - Add the Slack Settings field, validation, audit coverage, and Storybook save/error states. - Add Manage and Remove connection menus for all seven chat providers. Keep the draft setup button. Require removal confirmation and allow retry after failure. - Add regression coverage, an active/draft menu story, and connector documentation. ## Verification All CI checks are green for |
||
|
|
b82661b561 |
refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path > - Paperclip manages agents and their access to external tools. > - Connectors expose these tools through a governed MCP gateway. > - PR #13755 added a direct Composio MCP connection behind the experimental MCP aggregators flag. > - The old project API-key broker still created toolkit child connections and showed a separate Services tab. > - Keeping both paths leaves obsolete setup and session code in the product. > - This change removes the broker and preserves direct MCP setup, credentials, permissions, and execution. > - Saved legacy records fail closed and remain available for explicit removal. ## Linked Issues or Issue Description Related: #13755. This retirement supersedes the legacy-path fixes proposed in #12630, #12632, #12634, and #12906. It does not close those PRs. **What existing behavior does this improve?** Composio connector setup, management, and runtime dispatch. **Current behavior** Composio offers both direct MCP and a project API-key broker. The broker mints sessions and creates one child connection per toolkit. **Proposed behavior** Offer only direct MCP. Remove the toolkit Services UI, REST routes, API client, and session broker. Block saved legacy parent and child records from discovery, execution, health checks, reconnect, and OAuth. Preserve their records and credentials until the operator removes each connection. **Reason and benefit** The direct MCP connector becomes the single supported Composio workflow. Provider accounts remain managed in Composio. ## What Changed - Remove the API-key catalog method and its generated-source definition. - Delete Composio broker clients, session creation, account synchronization, child lifecycle, and toolkit routes. - Remove the Services tab, service rows, child provenance, and cascade-removal controls. Keep Vercel provenance intact. - Retain a shared retirement guard for stored legacy records. Show Retired status and replacement/removal guidance in the connection list and details; hide obsolete runtime controls. - Preserve the experimental MCP aggregators flag and direct MCP infrastructure. - Replace broker fixtures with retirement tests and extend direct Composio catalog/reconnect coverage. ## Verification - Focused shared, server, and UI tests passed with one worker. Server retirement tests use a name filter; no full local test suite was run, as requested. - Server and UI TypeScript checks passed. - Token gates and UI build passed. - Real browser: opened the saved Composio connection, refreshed all 11 tools, and ran the provider's read-only GitHub account-list operation through the standard Test dialog as an agent. The provider returned success using the existing OAuth credentials. - See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live evidence. - Storybook build passed. A fresh real agent used `COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway audit records confirm both calls succeeded. - Browser retirement check: a credential-free legacy fixture showed the guidance, opened the direct MCP replacement flow, and was removed through the standard confirmation. - Focused regressions for the experimental settings copy and exact OpenAPI route coverage passed. All latest-head CI checks passed (54 successful, two intentionally skipped); Greptile scored 5/5 with no unresolved review threads. The PR has no merge conflicts. ## Risks This intentionally breaks the old Composio project API-key and child-connection workflow. Existing legacy records cannot run, even if their stored status is active. Operators must create a new direct MCP connection and choose access rules; credentials and grants are not migrated. Remove each old record separately to delete its credentials. No schema migration or data deletion runs automatically. Direct MCP connections keep their existing grants and secrets. ## Model Used OpenAI GPT-6 via Codex, with reasoning, code execution, and browser tools. The exact runtime variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e8c8ba3c19 |
feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its tool gateway applies company access rules and approval controls to connected apps. > - MCP aggregators expose many apps through one provider endpoint. > - Each aggregator needs its own credential, catalog, grants, and lifecycle in Paperclip. > - This pull request adds independent Zapier, Arcade, Composio Connect, and Executor setup with a common Access → Connect layout. > - A default-off MCP aggregators flag lets operators opt in while we complete provider acceptance tests. > - Agents use the normal Paperclip permissions, Test screen, and gateway after setup. ## Linked Issues or Issue Description **Subsystem affected** Apps, connection setup, shared contracts, and the remote MCP gateway. **Problem or motivation** Aggregator endpoints need clear provider setup and correct MCP sessions. Generic setup does not explain each provider's authentication or broad execution tools. Provider approval must preserve the original execution instead of replaying a write. **Proposed solution** Add four separate connectors behind Settings → Experimental → MCP aggregators. Start with human and agent access, then connect the endpoint and read its tools. Enable tools by default. Use the existing Permissions and Test screens after setup. Keep legacy Composio API-key and child connections intact. **Alternatives considered** A shared connection for all providers would mix credentials and access rules. Separate provider-specific permission and test screens would duplicate existing controls. Vercel Connect is outside this change. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps capability and the Connected Apps roadmap area. This work was requested and reviewed by the maintainer. Related work: #11894, #12630, #12632, #12634, and #12906 concern the legacy Composio broker. #13102 also covers remote MCP pagination. This change preserves the broker path and adds initialized sessions, response matching, and provider resume handling alongside pagination. ## What Changed - Add branded setup and interactive Storybooks for Zapier, Arcade, Composio Connect, and Executor. Use the existing access controls and normal action tests. Do not request a connection name or action choices during setup. - Add the default-off `enableMcpAggregators` flag to settings, managed feature metadata, the catalog, and setup guards. Hidden connections keep running. Legacy Composio connections remain unchanged. - Reuse the vault, grants, policy, and catalog models. Support OAuth discovery, bearer tokens, custom headers, and credential-bearing URLs. Add no database tables or migrations. - Initialize and retain Streamable HTTP sessions by connection and effective credentials. Read paginated catalogs and match streaming responses to request IDs. - Classify unfamiliar aggregator tools as writes despite upstream read-only hints; only exact reviewed read capabilities enter the read-only allowlist. Legacy Composio child behavior is preserved. - Preserve provider authorization links and execution IDs. Support Executor approve/resume, decline, and cancel without automatic replay of uncertain writes. - Preserve Off and Ask first choices during refresh and reconnect. Allow new tools and retire removed tools. Keep agent access updates atomic and preserve an empty agent selection. - Document connector UX rules, provider branding sources, and live acceptance results. - Stabilize the existing Sentry release fixture after its repeated CI failure by reusing one module mock; production Sentry behavior is unchanged. ## Verification - Final head `d11781970`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/35633534900) passed, including broad typecheck, test shards, build, and E2E. All 54 checks pass; 2 optional checks are skipped. Greptile is 5/5, Security Scan passes, and all review threads are resolved. - Passed 27 focused connector Vitest checks and 18 connector-only Storybook browser checks before the flag change. All 85 stories rendered at desktop and narrow widths. - Passed 5 connector lifecycle/server checks and 7 selected flag checks after adding the flag. The latter cover settings, managed defaults, cached catalog visibility, and all four setup routes. - Review fixes passed 13 risk/handoff/lifecycle checks, dedicated session-expiration and transport regressions, 13 selected connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A real Composio connection-list call also succeeded through the refreshed UI on `9ab115f71`. - UI and server TypeScript checks passed. UI build, Storybook build, token gates, and diff whitespace checks passed during implementation. - Real browser and real Paperclip agent tests passed for Arcade, Composio, and Executor. Tested action permissions, denied agent access, reconnect, disconnect, and isolation. Tested Arcade catalog additions/removal and Executor provider approve/resume, decline, and cancel. - Zapier live acceptance is incomplete. Its dedicated provider server is configured, but its credential-copy dialog returned an empty clipboard through browser automation. No live Zapier action is claimed. - The three isolated Sentry release cases pass after the CI fixture fix. - Local verification is deliberately narrow at the maintainer's request. The full local suite, recursive typecheck, and repository-wide build were not run. CI provides the broader checks. ## Risks - Shared MCP transport changes affect other remote MCP servers. Protocol fixtures cover initialized sessions, streaming response matching, pagination, and isolation. - Broad execution tools remain broad permissions. The provider governs actions inside those tools. - Provider handoff links are retained briefly in memory. After a server restart, a one-time link may require reopening the provider dashboard. Paperclip does not replay the original call. - Zapier remains unproven live. Custom-header imports and self-hosted endpoints have fixture coverage rather than a separate live account for every variant. - Turning the experimental flag off hides setup; it does not revoke existing credentials or stop existing connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, shell execution, and browser automation. The exact runtime model ID and context-window size are not exposed in this session. A separate Anthropic-backed Paperclip agent performed live gateway acceptance tasks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b70641f23f |
feat(plugins): support image catalogs and persistent application overlays (#13646)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Plugins extend the application without adding each integration to
Core.
> - A downstream image needs a way to supply prebuilt plugins.
> - Some plugin UI must stay mounted as users move between pages.
> - This change adds an image catalog and a persistent application slot.
> - Operators can upgrade or remove these plugins through their image
and configuration.
## Linked Issues or Issue Description
**Subsystem affected**
Plugin packaging, activation and application UI.
**Problem or motivation**
The built-in plugin catalog is fixed in Core source. Downstream images
cannot add entries through an explicit catalog. Existing page slots also
cannot preserve a small application overlay across route changes.
**Proposed solution**
Read a bounded catalog of prebuilt plugins from the image. Verify its
files before importing manifests. Use the existing managed selection and
plugin lifecycle. Add an `appShellOverlay` slot with account and company
cleanup.
**Alternatives considered**
A downstream fork adds merge work. Script injection provides no plugin
lifecycle. A separate runtime download system adds a second distribution
channel.
**Roadmap alignment**
This extends the existing plugin system. Related PR #9006 covers runtime
install replication; this change covers immutable image contents. PR
#12555 covers CLI scaffolding. Neither provides this catalog or
application slot. The maintainer requested this work directly.
## What Changed
- Validate catalog identities, confined paths, package versions and
bundle hashes before importing code.
- Apply image selection to persisted plugin installs, including removal
and rollback. Adopt the verified image path from existing npm/local
installs and bind runtime worker/UI entrypoints to verified package
declarations.
- Mount application overlays in both UI shells. Preserve route state and
clear it on account, company and onboarding changes.
- Restrict service-worker offline storage/fallback to hashed public
assets in a separate cache namespace; exclude application HTML and
extension/API data, including after worker restart.
- Document the packaging contract, trust model and rollback
requirements.
## Verification
- Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`. Affected server/UI typechecks and builds, plus token
gates, passed again after rebasing onto current master; the 124 focused
tests also passed after rebase.
- Latest focused verification: 124 tests in nine files passed for
catalog/reconciliation/loader, overlay lifecycle, Layout and
service-worker policy. The broader UI/shared/SDK run passed 7,204 tests
in 690 files with canonical `TMPDIR`.
- Real disposable Core/PostgreSQL: catalog install, selection removal,
0.1.0→0.1.1→0.1.0, same-version npm/legacy-path adoption, and
preservation of disabled status passed. Added permissions entered
`upgrade_pending`, withheld UI across restart, and activated only after
explicit operator enable.
- Real Chromium: desktop/mobile layout, route draft retention and
Escape/focus passed with mocked extension responses. A persistent
browser restart retained public hashed-asset offline fallback while
refusing seeded legacy/current private entries and legacy HTML.
- Full `pnpm test:run`: 12,539 passed; 17 failed across six existing
files, stopping later phases. macOS read-only directory renames fail in
runtime-skill-cache and company-skills-service; email tests require an
absent local AgentMail fixture. Native runner/comment-redaction passed
in isolation after temporary Rust setup; agent-conversations also passed
in isolation. No unrelated source was changed to hide failures.
- After rebase, two unchanged chat timing tests failed in CI and passed
locally in isolation. Their CI shard passed on its single retry. All
other current-head CI jobs passed on the initial run; review is 5/5 with
no unresolved threads.
- No live deployment or external plugin service was used.
## Risks
- Plugins are trusted code. The catalog detects packaging errors; it
does not authenticate an untrusted image builder.
- Invalid catalogs fail startup. Images must contain the catalog and
bundles together, with stable directories.
- A host older than this contract lacks the activation guard. Disable
added plugins and remove their configuration keys before reverting to
it.
- Offline navigation now returns 503 instead of replaying cached
application HTML. Only public build assets have offline fallback.
- Rolling back an unapproved permission change retains the approval
gate; review the current manifest and explicitly enable it. A reduced
permission set cannot establish prior approval or prior enabled status.
- Plugin data migrations need their own rollback policy. This change
retains installed records and does not reverse migrations.
## Model Used
- OpenAI GPT-6 (Codex), model ID `gpt-6`, with repository inspection,
code execution and browser verification. The runtime does not expose an
exact context-window size.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites; broad
macOS server-run exceptions are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (fresh run on 488b3754ae; chat
shard passed its single retry)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(fresh review on
|