mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
5ee9e751fbb38787c055ee793c3412ecfc89da2e
1908
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8813a50105 |
feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path > - Paperclip manages agent work as tasks and runs. > - GitHub chat brings repository conversations into those tasks. > - A review bot needs the assigned agent, its authority, and governed provider tools. > - The existing channel connection did not supply that review workflow or a complete setup journey. > - This pull request adds GitHub App setup, account access, event prompts, task-bound review tools, and exact-commit checks. > - Operators can inspect each review through the same task, run, and activity systems. ## Linked Issues or Issue Description **Subsystem affected** GitHub chat, governed connection tools, task execution, shared/database contracts, and connector setup UI. **Problem or motivation** Operators need a GitHub review bot that runs their assigned Paperclip agent. Mentions and PR events must preserve task ownership and requester authority. Provider publication must use the bot App identity and enforce the configured permissions. **Proposed solution** Extend the existing GitHub chat connector with resumable App onboarding, linked-member and sponsored-guest access, editable event prompts, and governed review operations. Validate structured assessments on the server and compute a stable Paperclip Review check for the exact head commit. **Alternatives considered** A separate review scheduler would duplicate Paperclip execution and permissions. Reusing personal GitHub credentials would change the bot identity and credential boundary. **Roadmap alignment** This extends the existing Connected Apps and governed-tool infrastructure. The project owner requested and approved this design. Related PR #8645 imports external Codex review feedback; this change runs an assigned Paperclip agent and publishes its results through the existing chat connector. ## What Changed - Include the current Paperclip instance origin in the copied setup prompt. Storybook uses its configured Paperclip origin; callback parameters and URL credentials are excluded. - Add a Claude/Codex copy button in the real setup and Storybook opening step. Its detailed prompt asks four setup questions and guides embedded-browser setup, verification, and optional required checks. Clipboard failure exposes selectable instructions. - Add a tutorial that explains why App installation, review scheduling, and required checks are separate choices. - Add manifest registration, an existing-App path, separate installation and repository selection, repository refresh, and explicit account confirmation. - Add low-trust agent guidance, effective capability verification, member selection, and explicit restricted guests with a sponsor. - Add configurable PR events, prompts, repository overrides, rating thresholds, and separate formal-review permissions. - Give the assigned agent governed App tools to read PRs, comment, begin an assessment, submit findings, and optionally submit a formal review. - Bind review history, root PR events, and inline replies to ordinary tasks. Deduplicate deliveries/findings and reject stale publication. - Link check Details to the underlying task on the current trusted hostname, or to Reviews before task creation. - Add schema migration 0283, API contracts, production UI, and 49 interactive Storybook states. - Repair local lease recovery. Keep the Cloud Dockerfile identical to master; no provider-pack layer or runtime-default environment variable is added. - Retry only rolled-back wake-admission transactions after transient endpoint-lock contention. A deterministic held-lock regression proves one accepted wake. ## Verification - Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero diff against master. Final workspace typecheck and build passed. The new PostgreSQL migration regression passed and preserves existing relation and constraint identities after replay. - Greptile reviewed this exact head at 5/5. There are zero unresolved review threads and no merge conflicts. - All current-head checks are green: 54 passed and two conditional Storybook jobs skipped. This includes complete server/workspace test suites, build, typechecks, policy checks, Runner suites, browser suites, and security status. One timing-sensitive callback-ordering test passed in isolation and its CI shard passed one retry. The duplicate local full-suite run was stopped after CI completed; it is not counted as a local full-suite pass. - Before the final Slack rebase and migration renumbering, 186 focused GitHub tests, 14 native bootstrap cases, token gates, and Storybook build passed. The final rebase retained the new Slack communication guidance. - The embedded-browser setup test copied the full detailed prompt, including the configured Paperclip instance URL. Desktop and narrow layouts were checked. Component tests cover successful copying and clipboard failure with selectable text and retry. - Live local and hosted GitHub acceptance evidence refers to application revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks exercised issue mentions, automatic PR reviews, inline findings, repeated mentions, task continuation, and failing-to-passing checks after a push. The Storybook agent generated, built, and browser-rendered pages; missing acceptance text failed, matching text passed, and broken JSX produced an incomplete result. - Live cases also covered independently disabled push events, prompt injection, duplicate signed deliveries, rapid pushes, stale-result rejection, finding deduplication, and restart recovery. Formal reviews were denied while disabled and published only after explicit enablement. Check Details links pointed to the underlying task on the trusted hostname. - Those hosted native Claude runs used the provider-pack layer now removed from this PR. They do not prove native Claude works on the standard Cloud image. A replacement hosted native Codex run is not yet verified: the disposable QA tenant has only an Anthropic AI connection. No new staging or production deployment was made for the packaging removal. - Required-check merge enforcement could not be tested because the private disposable repository's GitHub plan rejected the rules configuration. Published success/failure/incomplete check states were verified directly. ## Risks - Latest master allocated migration 0282 to Slack. The GitHub migration is regenerated as 0283 with replay-safe table/index/constraint creation; a PostgreSQL regression verifies existing relations and constraints are preserved. Existing preview tenants remain subject to the fleet migration-history compatibility preflight; no bypass is introduced. - Migration 0283 adds company-scoped configuration, registration, review, and publication records. Existing connections retain their behavior until reviews/tools are enabled. - Signed webhooks and expiring registration state remain required. Hosted installations also need the companion narrow Cloud gateway exemptions. - Agent assessments can be incomplete or wrong. The server enforces coverage/result structure, current-head publication, rating policy, and separate formal-review permission; it does not replace code-review judgment. - No Cloud image packaging changes are included. Remote native ACPX/Claude and OpenCode retain their existing operator-supplied provider-pack prerequisite. Native Codex and Codex with managed MCP tools do not require that pack. Earlier staging deployment evidence refers to its stated revision, not this packaging-removal head. Production rollout and merging remain outside this change. ## Model Used OpenAI GPT-6 through Codex, with repository, code execution, API, and embedded-browser tools. The exact serving model ID and context-window size were not exposed by the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d9b3a5653e |
feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people use the same tasks and agent tools from external conversations. > - Agents need communication guidance that fits the conversation medium. > - That guidance belongs in the original task context, without repeated instructions on each turn. > - Connection owners also need clear settings and a consistent way to remove a connection. > - This pull request adds initial Slack guidance, optional connection instructions, and chat connection menus. > - The benefit is clearer Slack replies with the existing Paperclip workflow and permissions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent replies in Slack and chat connection management in the Apps catalog. **Current behavior** Slack tasks do not carry a saved communication profile. The catalog shows a separate Manage button and does not offer removal on every chat connection row. **Proposed behavior** Save Slack guidance when a new conversation creates a task. Restore that original guidance when a model session is rebuilt. Do not append it to ordinary follow-ups. Expose optional additional instructions in Slack Settings. Put Manage and Remove connection in a three-dot menu for all chat providers. Keep Finish setup visible for drafts. **Reason and benefit** Small answers fit in Slack. Substantial deliverables use ordinary document or artifact tools with a useful Slack summary. Connection settings apply to new tasks and cannot change permissions. Users can remove both active and unfinished chat connections from the catalog. **Breaking changes** Two additive database columns store endpoint preferences and the initial conversation snapshot. Existing endpoints default to empty preferences. Existing conversations keep their original behavior. Non-Slack guidance is unchanged. Related public context: https://github.com/paperclipai/paperclip/pull/13741 improves native chat recovery. This change adds communication context to those existing execution paths. A search found no duplicate communication-guidance PR. ## What Changed - Add a provider-guidance registry, enabled for Slack first. - Persist optional endpoint communication instructions and capture an immutable snapshot when a conversation creates a task. - Resolve guidance from the verified company-scoped connection. Restore it for fresh native and legacy sessions without per-turn reminders, extra model calls, or extra context queries. - Add the Slack Settings field, validation, audit coverage, and Storybook save/error states. - Add Manage and Remove connection menus for all seven chat providers. Keep the draft setup button. Require removal confirmation and allow retry after failure. - Add regression coverage, an active/draft menu story, and connector documentation. ## Verification All CI checks are green for |
||
|
|
b82661b561 |
refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path > - Paperclip manages agents and their access to external tools. > - Connectors expose these tools through a governed MCP gateway. > - PR #13755 added a direct Composio MCP connection behind the experimental MCP aggregators flag. > - The old project API-key broker still created toolkit child connections and showed a separate Services tab. > - Keeping both paths leaves obsolete setup and session code in the product. > - This change removes the broker and preserves direct MCP setup, credentials, permissions, and execution. > - Saved legacy records fail closed and remain available for explicit removal. ## Linked Issues or Issue Description Related: #13755. This retirement supersedes the legacy-path fixes proposed in #12630, #12632, #12634, and #12906. It does not close those PRs. **What existing behavior does this improve?** Composio connector setup, management, and runtime dispatch. **Current behavior** Composio offers both direct MCP and a project API-key broker. The broker mints sessions and creates one child connection per toolkit. **Proposed behavior** Offer only direct MCP. Remove the toolkit Services UI, REST routes, API client, and session broker. Block saved legacy parent and child records from discovery, execution, health checks, reconnect, and OAuth. Preserve their records and credentials until the operator removes each connection. **Reason and benefit** The direct MCP connector becomes the single supported Composio workflow. Provider accounts remain managed in Composio. ## What Changed - Remove the API-key catalog method and its generated-source definition. - Delete Composio broker clients, session creation, account synchronization, child lifecycle, and toolkit routes. - Remove the Services tab, service rows, child provenance, and cascade-removal controls. Keep Vercel provenance intact. - Retain a shared retirement guard for stored legacy records. Show Retired status and replacement/removal guidance in the connection list and details; hide obsolete runtime controls. - Preserve the experimental MCP aggregators flag and direct MCP infrastructure. - Replace broker fixtures with retirement tests and extend direct Composio catalog/reconnect coverage. ## Verification - Focused shared, server, and UI tests passed with one worker. Server retirement tests use a name filter; no full local test suite was run, as requested. - Server and UI TypeScript checks passed. - Token gates and UI build passed. - Real browser: opened the saved Composio connection, refreshed all 11 tools, and ran the provider's read-only GitHub account-list operation through the standard Test dialog as an agent. The provider returned success using the existing OAuth credentials. - See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live evidence. - Storybook build passed. A fresh real agent used `COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway audit records confirm both calls succeeded. - Browser retirement check: a credential-free legacy fixture showed the guidance, opened the direct MCP replacement flow, and was removed through the standard confirmation. - Focused regressions for the experimental settings copy and exact OpenAPI route coverage passed. All latest-head CI checks passed (54 successful, two intentionally skipped); Greptile scored 5/5 with no unresolved review threads. The PR has no merge conflicts. ## Risks This intentionally breaks the old Composio project API-key and child-connection workflow. Existing legacy records cannot run, even if their stored status is active. Operators must create a new direct MCP connection and choose access rules; credentials and grants are not migrated. Remove each old record separately to delete its credentials. No schema migration or data deletion runs automatically. Direct MCP connections keep their existing grants and secrets. ## Model Used OpenAI GPT-6 via Codex, with reasoning, code execution, and browser tools. The exact runtime variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1483bb8bcf |
fix: pass plugin workers to issue tree resume (#13757)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue tree controls can resume tasks and wake their assigned agents. > - A sandbox-backed run needs the shared plugin worker manager to acquire its lease. > - The issue tree route constructed a heartbeat service without that manager. > - A ready sandbox provider therefore appeared offline and the resumed run failed setup. > - This pull request passes the existing manager to the route and distinguishes missing wiring from a stopped worker. ## Linked Issues or Issue Description Related implementation: #10262 identified this wiring defect in several dispatch paths. This PR applies the issue-tree resume fix to current master and adds coverage in the current route and runtime suites. The other dispatch paths remain in the scope of that earlier PR. **What happened?** Resuming a paused issue subtree can fail to acquire a sandbox lease even while its provider plugin is ready and its worker is running. The error says the worker is not running because this route's heartbeat service never received the process's worker manager. **Expected behavior** Resumed work uses the same worker manager as normal issue dispatch. Missing wiring and an actually stopped worker produce distinct diagnostic errors. **Steps to reproduce** 1. Configure an agent with a plugin-backed sandbox environment and start the provider worker. 2. Pause a subtree assigned to that agent. 3. Release the hold with `metadata.wakeAgents: true`. 4. The old route constructs an unwired heartbeat service and the resumed run fails during setup. **Paperclip version or commit** Reproduced by source inspection and regression coverage against `9d19f98b50`. **Deployment mode** Server with a plugin-backed sandbox environment. ## What Changed - Pass the process's plugin worker manager from `createApp` through issue tree controls to heartbeat dispatch. - Check for a missing manager before reporting the sandbox worker as stopped. - Exercise the resume request with a shared manager and cover both runtime failure cases. Model a stopped worker explicitly in the existing infrastructure-retry fixture. Existing board and company access checks remain covered. ## Verification - Targeted Vitest route/runtime/recovery suites: 17 passed; 384 native database tests skipped because embedded Postgres cannot start on this host. - Direct server `tsc --noEmit` passed. - `pnpm -r typecheck`, server typecheck wrapper, and `pnpm build` were attempted. Their runner dependency requires `cargo`, which is absent on this host. CI must pass these checks before merge. - Full `pnpm test:run` was attempted. The general server stage reported 8,146 passed, 4,747 skipped, and 14 failed tests across 35 failed suites. Failures were embedded Postgres startup errors (with related teardown errors) and 10 cache-directory rename failures on this macOS host. The local run stopped there; Linux CI must pass the complete suite before merge. - All CI gates are green, including the native database recovery suite, workspace typecheck/build, runner checks, and browser tests. Greptile reviewed the current commit at 5/5 with no unresolved threads. - No live provider operations were performed by these tests. ## Risks Low risk: dependency forwarding and error classification only. The route retains board authorization, company boundaries, hold semantics, and existing cancellation/replay guards. The change does not replace a sandbox or discard a retained lease. No schema or API shape changes. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tooling, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] This fixes existing behavior and does not add planned core features - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the bug issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no private ticket or instance data - [x] Targeted local tests pass; full-suite and toolchain limits are recorded above - [x] I have added or updated tests where applicable - [x] No documentation changes are needed for dependency forwarding - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e8c8ba3c19 |
feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its tool gateway applies company access rules and approval controls to connected apps. > - MCP aggregators expose many apps through one provider endpoint. > - Each aggregator needs its own credential, catalog, grants, and lifecycle in Paperclip. > - This pull request adds independent Zapier, Arcade, Composio Connect, and Executor setup with a common Access → Connect layout. > - A default-off MCP aggregators flag lets operators opt in while we complete provider acceptance tests. > - Agents use the normal Paperclip permissions, Test screen, and gateway after setup. ## Linked Issues or Issue Description **Subsystem affected** Apps, connection setup, shared contracts, and the remote MCP gateway. **Problem or motivation** Aggregator endpoints need clear provider setup and correct MCP sessions. Generic setup does not explain each provider's authentication or broad execution tools. Provider approval must preserve the original execution instead of replaying a write. **Proposed solution** Add four separate connectors behind Settings → Experimental → MCP aggregators. Start with human and agent access, then connect the endpoint and read its tools. Enable tools by default. Use the existing Permissions and Test screens after setup. Keep legacy Composio API-key and child connections intact. **Alternatives considered** A shared connection for all providers would mix credentials and access rules. Separate provider-specific permission and test screens would duplicate existing controls. Vercel Connect is outside this change. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps capability and the Connected Apps roadmap area. This work was requested and reviewed by the maintainer. Related work: #11894, #12630, #12632, #12634, and #12906 concern the legacy Composio broker. #13102 also covers remote MCP pagination. This change preserves the broker path and adds initialized sessions, response matching, and provider resume handling alongside pagination. ## What Changed - Add branded setup and interactive Storybooks for Zapier, Arcade, Composio Connect, and Executor. Use the existing access controls and normal action tests. Do not request a connection name or action choices during setup. - Add the default-off `enableMcpAggregators` flag to settings, managed feature metadata, the catalog, and setup guards. Hidden connections keep running. Legacy Composio connections remain unchanged. - Reuse the vault, grants, policy, and catalog models. Support OAuth discovery, bearer tokens, custom headers, and credential-bearing URLs. Add no database tables or migrations. - Initialize and retain Streamable HTTP sessions by connection and effective credentials. Read paginated catalogs and match streaming responses to request IDs. - Classify unfamiliar aggregator tools as writes despite upstream read-only hints; only exact reviewed read capabilities enter the read-only allowlist. Legacy Composio child behavior is preserved. - Preserve provider authorization links and execution IDs. Support Executor approve/resume, decline, and cancel without automatic replay of uncertain writes. - Preserve Off and Ask first choices during refresh and reconnect. Allow new tools and retire removed tools. Keep agent access updates atomic and preserve an empty agent selection. - Document connector UX rules, provider branding sources, and live acceptance results. - Stabilize the existing Sentry release fixture after its repeated CI failure by reusing one module mock; production Sentry behavior is unchanged. ## Verification - Final head `d11781970`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/35633534900) passed, including broad typecheck, test shards, build, and E2E. All 54 checks pass; 2 optional checks are skipped. Greptile is 5/5, Security Scan passes, and all review threads are resolved. - Passed 27 focused connector Vitest checks and 18 connector-only Storybook browser checks before the flag change. All 85 stories rendered at desktop and narrow widths. - Passed 5 connector lifecycle/server checks and 7 selected flag checks after adding the flag. The latter cover settings, managed defaults, cached catalog visibility, and all four setup routes. - Review fixes passed 13 risk/handoff/lifecycle checks, dedicated session-expiration and transport regressions, 13 selected connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A real Composio connection-list call also succeeded through the refreshed UI on `9ab115f71`. - UI and server TypeScript checks passed. UI build, Storybook build, token gates, and diff whitespace checks passed during implementation. - Real browser and real Paperclip agent tests passed for Arcade, Composio, and Executor. Tested action permissions, denied agent access, reconnect, disconnect, and isolation. Tested Arcade catalog additions/removal and Executor provider approve/resume, decline, and cancel. - Zapier live acceptance is incomplete. Its dedicated provider server is configured, but its credential-copy dialog returned an empty clipboard through browser automation. No live Zapier action is claimed. - The three isolated Sentry release cases pass after the CI fixture fix. - Local verification is deliberately narrow at the maintainer's request. The full local suite, recursive typecheck, and repository-wide build were not run. CI provides the broader checks. ## Risks - Shared MCP transport changes affect other remote MCP servers. Protocol fixtures cover initialized sessions, streaming response matching, pagination, and isolation. - Broad execution tools remain broad permissions. The provider governs actions inside those tools. - Provider handoff links are retained briefly in memory. After a server restart, a one-time link may require reopening the provider dashboard. Paperclip does not replay the original call. - Zapier remains unproven live. Custom-header imports and self-hosted endpoints have fixture coverage rather than a separate live account for every variant. - Turning the experimental flag off hides setup; it does not revoke existing credentials or stop existing connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, shell execution, and browser automation. The exact runtime model ID and context-window size are not exposed in this session. A separate Anthropic-backed Paperclip agent performed live gateway acceptance tasks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3790ca2f13 |
fix(runner): repair approval and Stop races and eval infrastructure (#13750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Runner tasks must continue after approval and stop when the user presses Stop. > - Live evals found races at approval delivery and provider startup. > - Browser readiness and CI setup errors also hid the actual task results. > - This pull request fixes those races and the related test infrastructure. > - Regression tests and saved live reports show which cases now pass. ## Linked Issues or Issue Description Companion eval definitions PR: https://github.com/paperclipai/paperclip-evals/pull/25 (AgentCore paused and provider/environment infrastructure). Related: #13741 now supplies the late-startup Stop fence and warm-attachment recovery; this PR retains that fence and extends startup tracking and regression coverage to both native backend paths. #13539 introduced queued approvals during active runs. #13738 fixes child assignment, task replies, and warm process continuity and is already in the base. #13291 concerns automatic continuation of interrupted legacy sandbox runs; this PR fixes native startup cancellation and does not change that recovery policy. **What happened?** An accepted service approval could wait after its source run stopped. Stop could return success before the provider handle existed. Work could then start after Stop, or a cancelled run could be recorded as failed. Some E2E tests also failed on unloaded browser content or irrelevant reply wording. Runner CI could fail before model work because of dependency or sandbox setup. **Expected behavior** Deliver each settled approval once after its source run stops. Do not start work after an acknowledged Stop. Preserve the audited cancellation. Test the intended product behavior with a ready browser and verified runtime dependencies. **Steps to reproduce** 1. Approve a service request while its source run is active. Let the run finish. Check that its result starts one continuation. 2. Delay provider startup. Press Stop before its handle is available. Check cancellation, then submit `/new`. 3. Run the browser, warm-workspace, and Stop-and-redirect cases from the linked report. **Paperclip version or commit** The branch includes master at `9d19f98b5`. The report records the original source for each focused attempt. **Deployment mode** Isolated local development instances and disposable Daytona sandboxes. ## What Changed - Deliver settled tool-action results for the exact company and source run during final cleanup. Keep the existing idempotent receipt and periodic recovery sweep. - Wait for startup to hand off its provider handle before acknowledging Stop. Reject first-turn admission after cancellation. Preserve a matching audited pending or acknowledged cancellation. - Wait for mounted task history and connector controls in browser tests. Record failure evidence. Grade workspace contents and process continuity separately from exact reply wording. Require each warm-turn marker once and in order, allowing surrounding prose. - Stop-and-redirect now checks that the source file exists and work is active before Stop. - Resolve target dependency locks in an uncredentialed CI job. Verify the lock artifact hash. Keep orchestration and publication on the trusted workflow revision. - Materialize the pinned OpenCode executable and configure the exact Codex executable's user-namespace profile before provider credentials are available. - Compress Daytona directory uploads with gzip. Preserve files, executable modes, symlinks, empty directories, and confinement checks. - Classify file-transfer RPC deadlines as infrastructure. Keep unrelated runner RPC failures visible. ## Verification - [Focused live report with screenshots and original attempts](https://pages.paperclip.ing/runner-reliability-20260921/): 14 of 15 selected Product E2E cases pass across the recorded revisions. Claude and Codex Stop → `/new`, Claude service approval, delegation, both hiring/reuse cases, and native Daytona warm continuity pass. - Two credentialed Runner smoke cases pass. These are not full protocol coverage. - E2E harness after the master merge: 429 tests pass. E2E and server TypeScript checks pass. - Daytona plugin: 239 tests pass, 6 skipped. Plugin TypeScript build passes. The compression test fails against the old code and passes with the change. - Runner backend/runtime regression group: 161 tests pass. Cancellation/startup selection: 26 tests pass. Approval delivery: 34 real-database tests pass. - Workflow security: 7 tests pass. Both edited workflows pass actionlint. Runner TypeScript and Rust builds pass. - After merging master, all 389 native executor tests pass, including both native backend paths and late startup after the Stop deadline. - Post-merge `pnpm -r typecheck` and `pnpm build` pass. The monolithic local `pnpm test:run` was interrupted to integrate master and is inconclusive. The [hosted CI test partitions](https://github.com/paperclipai/paperclip/actions/runs/35620461738) pass on `50a3e43822bcba1e0d07b1b45b0be91cbf9312da`. An unchanged sandbox callback schema test initially received HTTP 503. It passed five isolated local runs, its full local test file, and one failed-job CI retry. No assertion was weakened. ## Risks - Stop can wait for the bounded startup handoff. If it cannot settle, the existing pending-recovery state remains instead of a false acknowledgement. - Immediate approval delivery must remain idempotent across cleanup and recovery sweeps. Tests cover duplicate delivery and company/run boundaries. - The workflow changes still need hosted Linux verification. They retain the trusted workflow and credential boundaries. - Gzip reduces the observed provider upload from about 1.8 GB to 663 MB. It does not yet fix the remaining Claude Daytona transfer timeout. That recovery test never reached Claude, so recovery remains unverified. Use a matching image with the verified provider package preinstalled for the next recovery test; retain cold-upload coverage separately. - The report preserves diagnostic runs with missing source metadata and marks them as such. It does not claim a new full-suite pass. - This PR adds no new prompt policy or historical status reconciliation. ## Model Used OpenAI GPT-6 through Codex performed the primary implementation and review. The exact primary backend model ID is not exposed in this session. OpenAI `gpt-5.6-luna` assisted with bounded infrastructure work and verification. The agents used repository tools, code execution, and browser tests. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 <noreply@openai.com> |
||
|
|
9d19f98b50 |
fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent chat uses native runner sessions to plan, delegate, and track that work. > - A user can press Stop while the native session is still starting. > - The server can acknowledge that Stop without dispatching it, then let the session submit a turn. > - This leaves chat recovery waiting for an execution that the user expected to stop. > - This PR waits for the startup handle, dispatches cancellation, and prevents a late startup from submitting a turn. > - New full-stack evals check the resulting records and outputs across Claude and Codex. > - Those evals also exposed missing ACPX readiness fields, unbounded polling, and an old-run identity check that rejected valid warm handoffs. ## Linked Issues or Issue Description **What happened?** Stop during native startup could record an acknowledged cancellation with `dispatched: false`. The provider could then begin work. A subsequent `/new` stayed queued. A remote Claude follow-up also exhausted the command journal while probing warm-session readiness: ACPX never returned the readiness fields required by the shared transport. Once readiness worked, attachment incorrectly compared the next run descriptor against the old run ID. The 25 ms polling loop could issue 4,800 commands during its two-minute wait, beyond the 500-command bound. The existing chat eval treated lifecycle logs as proof of an active provider turn, so it did not distinguish startup cancellation from active-turn cancellation. **Expected behavior** A Stop during startup must reach the pending session. A late session must not submit a prompt after Stop. Recovery must retain control when startup exceeds the bounded wait. Chat evals must check saved task state, document contents, worker identity, account binding, and duplicate effects. **Steps to reproduce** 1. Start a native Claude or Codex chat turn. 2. Press Stop after process startup is requested but before the provider turn starts. 3. Send `/new`, then send a fresh message. 4. On the affected base, cancellation can be acknowledged without dispatch and the reset stays queued. **Paperclip version or commit** The live Claude baseline reproduced this on `29d6b3509`. The branch also includes master commit `0f5fafe16`. Related work: #13678, #13686, #13693, #13291, #13738. A separate runner reliability branch also contains a startup-wait fix. Its overlap must be reconciled before merging; this branch additionally prevents prompt submission after a late startup. ## What Changed - Wait for a pending native startup before acknowledging a run-scoped Stop. Preserve the existing recovery error when that wait expires. - Keep a Stop guard on startup. Cancel a late handle before it can submit a provider turn. - Add regression tests for normal handle publication and publication after the Stop deadline. - Back off blocked warm-attachment probes. Keep the fast two-snapshot barrier, fail closed, and record changed blockers. - Add red/green tests for delayed readiness, persistent blockers, alternating readiness, and readiness near the deadline. - Publish ACPX readiness and blockers. Preserve the old authority’s event acknowledgement barrier; only settled sessions can proceed to attachment. - Bind warm ACPX descriptors to the validated next authority while retaining old-run event correlation until activation. Preserve session identity and provider profile checks. - Exercise two consecutive run rotations through a qualified fake sidecar, verifying checkpointing, provider identity, pre-activation rejection, and new-run work admission. - Separate startup and active-turn cancellation checkpoints in the browser eval. - Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells across Claude and Codex. - Cover hiring and reuse through managed AI accounts, source-based review, current blocked-task status, request replay after a lost HTTP acknowledgement, server restart continuity, and Stop/reset continuity. - Use ordinary production agent instructions. Enable API tools only for the two coordination cases that need them. - Calibrate the matchers with invalid records and outputs. Require remembered context after restart and a structured status snapshot that distinguishes the current blocker from history and task status from active execution. Compare the public issue mutation contract and relationships during read-only reporting. Preserve before/after source records in failed eval evidence. - Fix the lost-ack browser harness and verify it against a real HTTP server. Check the chat composer after restart instead of waiting for an unrelated document lifecycle event. - Document the scope and limits of each case. ## Verification - The startup regression failed on the unfixed executor and passed after the fix. - `pnpm test:e2e:runner:typecheck` passed. - `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files. - `pnpm exec vitest run server/src/services/native-runtime/native-session-executor.test.ts` passed: 385 tests. - [Baseline live campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868): Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed. Codex delegation was blocked by provider capacity. - [Eval-only startup campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786): both providers failed as expected. Both persisted `dispatched: false` and left `/new` queued. - [First fixed campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706) on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both providers. Failed cases exposed eval harness defects and remote continuity failures. All attempts remain available. - [Original workflows and stronger memory checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649) on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and startup Stop passed for both providers; Codex remote restart passed. Claude remote restart exposed the missing readiness contract. Two Codex planning cells hit provider capacity. - [Unchanged-model retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548): Codex planning and backlog creation both passed. - [18-cell campaign with ACPX readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963) on `6a98ef743`: 16/18 passed, including all local/remote Stop and committed-send cases. Claude remote continuity exposed the next-authority check, now fixed. Codex hiring produced its checklist, but the runner redacted the requested marker after it appeared as “Tracking token: …”. That content-redaction policy is unchanged and remains an explicit limitation. - [Structured status grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725) on `50448c228`: both providers passed on their first attempt, including cleanup. - [Complete read-only state grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011) on `551e13892`: both providers passed. - [Final ACPX handoff and hiring retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456) on `cbd637587`: all three Claude Daytona cases passed (restart continuity, active Stop/reset, and lost-ack replay). Codex hiring reproduced the content-redaction failure: the saved checklist contained `Tracking token: [REDACTED]` instead of the required business marker. All four cases completed cleanup successfully. [Published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/). The only subsequent commit adds the qualified-sidecar integration test; production code is identical to this live proof. - `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without paid models. - Runner TypeScript typecheck passed. All 5 warm-readiness tests pass; two failed with the prior fixed-rate loop, and the late-readiness test failed before the pacing correction. - ACPX readiness and warm-identity regressions each failed before their fixes. All 292 runner-core Rust library tests passed. The qualified-sidecar integration test passes. Rust formatting is checked. - Status-grader regressions for misleading historical mentions and previously unchecked mutations each failed before tightening the oracle and pass now. - [Latest-head CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307) passed on `a4093c8f1`: full build, type checks, test partitions, browser E2E, and native runner checks. Two unrelated tests initially failed (Sentry fixture release attribution and local-service fixture readiness); both passed locally together (35 passed, 5 optional SDK tests skipped) and on the failed-job retry. No changes were made to those tests. - Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed and all review threads are resolved. - The paid live suite is not fully green: the reproducible content-redaction case remains red. This is separate from the passing PR merge checks. No production content-redaction, prompt, model, or completion-policy change is included. - Managed-account hiring and review cases explicitly enable API tools; these do not qualify default new-user onboarding. ## Risks - Stop can wait up to 30 seconds for startup, then use the existing pending-recovery path. This does not prove that remote cleanup has finished. - Blocked warm readiness adds up to 750 ms between later probes with the two-minute remote budget, or about 32 ms with the default five-second budget. Ready sessions retain the short second barrier. - Paid evals can fail because of provider capacity or agent decisions. Each failure needs evidence-based classification. - The HTTP request replay case checks comment idempotency and duplicate effects. It does not prove replay safety for an ambiguous provider tool call. - The new suite is opt-in. It does not increase the default paid campaign. - No production prompts or model selection change. Review-handoff behavior and content-redaction policy remain separate product decisions. The latter can remove harmless business content that looks like credential syntax; the failing attempt is retained. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fd071748ee |
fix: stop repeated notifications for finalized run failures (#13739)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native run reconciliation repairs saved execution outcomes after
interruptions.
> - The sweep also visits runs whose final results are already
committed.
> - An unchanged failed run still received a new status delivery ID on
every sweep.
> - The browser treated each delivery as a new failure after its short
duplicate window expired.
> - This pull request makes unchanged projections a no-op and suppresses
repeated or historical run toasts.
> - Operators receive fresh failure alerts without repeated alerts for
old work.
## Linked Issues or Issue Description
**What happened?**
An old failed run repeatedly produced failure toasts while the browser
remained open. The task could already be cancelled. Reconciliation
rewrote the same failed outcome and queued another status broadcast.
**Expected behavior**
An unchanged committed run must not queue a new status notification.
Repeated deliveries must still refresh cached state without another
toast.
**Steps to reproduce**
1. Finalize a native run with a failed result and successful workspace
finalization.
2. Deliver its pending execution status and cancel its task.
3. Replay finalization and status delivery on each periodic sweep.
4. Observe another failure broadcast for every sweep before this fix.
**Paperclip version or commit**
Reproduced against source commit
|
||
|
|
0f5fafe16b |
fix(runner): preserve task replies and warm process continuity (#13738)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - The Runner connects provider sessions to task state, replies, and delegated work. > - Full-stack tests found lost final replies, rejected helper calls that stopped the parent, and unnecessary process restarts. > - A completed child could also receive a new assignment wake that the scheduler then cancelled. > - This change fixes those boundaries and gives agents clearer teammate instructions. > - The tests retain strict completion and process-continuity requirements. ## Linked Issues or Issue Description **What happened?** A generated attachment comment could suppress an agent's final reply. A known Codex helper could stop its parent when it requested a Paperclip tool. Native Daytona processes restarted between turns because Paperclip minted an unused GitHub broker token. Reassigning a completed child queued a run that immediately cancelled. Revision instructions also allowed agents to do work assigned to a named teammate themselves. **Expected behavior** Keep the final reply. Reject helper tool requests without borrowing parent authority or stopping the parent. Keep an unconfigured sandbox process alive between turns. Treat assignment-only changes to completed tasks as metadata changes. Preserve explicit teammate assignments during revisions. **Steps to reproduce** Run the retained Runner E2E cases for file handoff, teammate reuse, Daytona warm continuity, and Legacy Claude interview/plan acceptance. The focused regression tests reproduce the reply, helper, process-lifetime, and assignment-wake defects without provider calls. **Paperclip version or commit** The live lifetime and completion campaign used `db3857807`. This PR replays the changes on master `c65fc9e3c`. See Verification for the limits of that evidence. **Deployment mode** Isolated local instances and native Runner sessions in Daytona sandboxes. Related work: #13546 handles a different queued-run issue after an issue-lock compare-and-set failure. This PR prevents the unnecessary assignment wake earlier. #13410 covers retained user services; this PR covers the provider process. No duplicate fix was found. ## What Changed - Exclude generated deliverable-binding comments from final-reply deduplication. Preserve the attachment and explicit user-facing replies. - Reject Paperclip tool and input requests from known Codex helper threads without terminating the parent. Keep unknown-thread rejection intact. - Explain how to hire or reuse a persistent teammate and preserve named delegation on revisions. Update generated protocol fixtures. - Use stable, token-free GitHub wrappers for unconfigured native sandboxes. Preserve credential isolation, configured-account rotation, and cleanup after partial staging failures. - Do not queue assignment-only wakes for done or cancelled tasks. Keep explicit reopening behavior. - Make warm-continuity fixtures create real review cards. Read the persisted final response selected by production presentation logic. Missing selected evidence still fails. ## Verification - Before rebase: 560 focused route, native-executor, and launcher tests passed. The new regressions were reproduced before their fixes. - Live E2E: Legacy Claude interview/plan acceptance passed 3/3 repetitions. Daytona warm continuity passed 2/3 full repetitions. Each successful run retained one process and provider session for all three turns. - The remaining Daytona repetition stopped after a same-URL browser reload left the page blank. Both completed turns retained the same process. Its failed verdict remains unchanged; this PR does not claim the blank-page cause is fixed. - Reports: https://pages.paperclip.ing/runner-e2e-lifetime-race-20260920/investigation.html and https://pages.paperclip.ing/runner-e2e-behavior-followups-20260919-results/investigation.html - Post-rebase `pnpm build` and `pnpm -r typecheck` passed. All 414 Runner E2E harness unit tests and its typecheck passed. Codex protocol tests: 88 passed, 2 ignored. - Latest-head CI: 55 successful checks and 2 intentional skips. Greptile: 5/5 with no review threads. The unchanged workspace exposure tests hit a fixed-port collision on the first CI attempt; their local suite passed (25 tests, 3 platform skips), and the CI shard passed on one retry. - The duplicate local `pnpm test:run` was stopped after the full hosted general and serialized test shards passed. It did not finish locally and is not counted as a local full-suite pass. ## Risks Configured GitHub accounts retain run-scoped credential rotation and can still restart warm processes. That limitation requires a separate design. Known provider helpers cannot use Paperclip coordination tools directly; they must return findings to the parent. The delegation prompt is an instruction, not an enforced guarantee; Codex Mini hiring/reuse failures remain open. No schema or workflow changes are included. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and parallel coding agents. The exact deployment suffix and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub issues and shared reports) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; full hosted test shards also pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c65fc9e3c8 |
fix: recover authentication and browser connection failures (#13724)
fix: recover authentication and browser connection failures Include connect timeouts in the bounded retry policy for idempotent actor synchronization. Handle WebSocket constructor failures through existing reconnect paths and preserve HTTP polling while realtime is unavailable. Refresh visible company queries until the socket recovers and clear all fallback timers on hiding or unmount. Verify 172 focused tests, server/UI typechecks, UI build, and design token gates. Full workspace build/typecheck require the unavailable Rust toolchain; the full test run is tracked separately. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9f30eb10dd |
fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path > - Paperclip manages AI agents and keeps their work attached to tasks. > - Chat connectors carry user messages and agent replies between a provider and those tasks. > - Each extra startup and context reset delays a reply. > - Managed account metadata was lost during adapter decoding, so compatible follow-ups started fresh. > - This branch fixes the reset and measures the remaining preparation, execution, and delivery costs. > - The changes must preserve account isolation, authorization, durable output, and recovery ownership. ## Linked Issues or Issue Description Refs #13699. The related service lifecycle work in #13410 and #13408 is separate; this branch focuses on task-bound chat response latency. **What happened?** Managed AI follow-ups started new provider sessions even after their configuration fingerprint stayed stable. The Codex codec removes unknown fields. The resume check then read the removed credential identity and treated it as a credential change. **Expected behavior** Compatible follow-ups resume the correct provider session. Changes to credentials, responsible users, permissions, or task configuration retain their reset behavior. **Steps to reproduce** 1. Use a Slack connector with a managed AI connection. 2. Send a message, then send a same-thread follow-up. 3. Inspect the configuration reset reason and the provider session identity. **Paperclip version or commit** Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`. **Deployment mode** Cloud staging with a native Codex runner. ## What Changed - Read saved credential identity before adapter decoding discards it. - Remove the internal credential identity from adapter-facing session params. - Test the real Codex codec and missing, changed, or unmanaged identity cases. - Preserve configured warm Codex runners and flush refreshed credentials after every turn. - Fence detached or closing session handles from successor credential ownership. - Stage current Codex launch credentials after restoring durable session history, uploading launch assets only once. - Reuse a runner binary already in the retained sandbox only when its SHA-256 matches the controller-owned artifact; still verify required capabilities before launch. - Lock the task before the run when saving results, preventing deadlocks with task updates. - Scope reusable projectless sandboxes to the company, environment, task, agent, and runtime configuration; verify Daytona sentinels for that scope. - Admit a new authorized chat message after a fully committed failed run and verified process cleanup. - Send compact deltas for verified plain-text Slack continuations. Match the actual prior run and current comment identity/body; exclude edited historical comments and prior agent output, preserve genuine brief edits and the full bootstrap fallback. - Keep attachments, omitted input, questions, approvals, recovery, and other providers on their existing framing. - Document managed session compatibility, credential lifecycle, and compact continuation boundaries. ## Verification - Workspace/session coverage: 156 tests passed. - Native session and credential ownership coverage: 390 tests passed, including exact artifact reuse, mismatches, failed probes, timeouts, and explicit artifact overrides. - Explicit continuation and durable chat authorization coverage: 172 tests passed. - Session resume and launch preparation coverage: 416 tests passed. - Result persistence coverage: 15 tests passed. The new concurrency test reproduced a PostgreSQL deadlock before the lock-order fix. - Environment lifecycle coverage: 92 tests passed, including projectless reuse and task/agent isolation at both selection and atomic handoff. - Daytona plugin coverage: 237 tests passed; 6 gated tests skipped. Standalone plugin build passed. - Compact Slack continuation and native resume coverage: 69 tests passed, including full-bootstrap retention, matching message authors/bodies, current-delivery selection, rejection of duplicate identities and historical comments, brief edits, and attachment/recovery fallbacks. - Final frozen-head `pnpm test:run` on repository-supported Node 26: 668 suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped. The sole failure was a local `socket hang up` in `issue-recovery-actions.test.ts`, not an authorization assertion mismatch. All 57 tests in that suite passed three fresh reruns, and the suite passed latest-head CI. The full local invocation is therefore not claimed green. - An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup failure; that 11-test catalog suite passes on Node 26 and in CI. No test behavior or timeout was relaxed. - Full local typecheck and build passed. Latest-head CI is green; Greptile is 5/5 with no unresolved review threads. - Two real Slack baseline replies took 25.1 and 24.6 seconds (24.9-second mean). Three same-thread signed probes on this head took 23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample and a modest wall-clock improvement, not a large or statistically established speedup. - In that same thread, uncached provider input fell from 8,514 tokens before compact input to 694–765 tokens afterward. The current delivery uses a 362-character delta; the full 19–21k-character bootstrap remains available for failed resume. Verified runner artifact preparation fell from about 1.2 seconds to 0.6 seconds. - A fresh thread created a separate task, sandbox, and provider session with full bootstrap (24.4 seconds). Its follow-up reused its own sandbox/session and compact input (28.3 seconds, including 16 seconds of model execution). Model variability and process startup remain substantial. - A signed duplicate webhook produced exactly one user comment, one successful run, and one final Slack reply. Slack's API independently confirmed the actual replies and a public task URL without an internal or pool hostname. - Earlier signed probes verified recovery after a failed run and reuse across a server deployment. The final idle test observed Daytona report the sandbox as stopped, then delivered a new reply in 19.9 seconds using the same sandbox/provider-session identity and compact input. Slack’s API confirmed that reply. - Live probes use signed synthetic inbound webhooks and real outbound Slack delivery, read back through Slack’s API. The final browser recheck found the Mac locked and the Slack tab blocked by another extension, so this is not claimed as full UI E2E proof. - This is a review branch. Do not merge until the maintainer reviews it. ## Risks - Incorrect session reuse could mix account or task context. Missing or changed identities continue to reset, and existing authorization checks remain in place. - Warm mode remains opt-in. Remote warm mode requires a reusable sandbox lease. Retained processes keep credentials until they close, so idle expiry and ownership fences are required. - A fresh user message may continue after a committed provider failure. Approval, current authorization, process termination, and prior-result checks remain required. - Projectless sandbox reuse is task- and agent-scoped. Missing or mismatched ownership cannot replace an existing lease; existing workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox per task/agent, so provider auto-stop and deletion policies still determine idle compute and storage costs. Fleet defaults are unchanged. - Compact prompts apply only after proven resume and a matching prior-run delta. Missing or specialized context falls back to full input; fresh sessions always receive the full bootstrap. - No schema or migration changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, tool use, and test execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
600e552d7b |
fix: attribute Sentry errors to the loaded source release (#13719)
Attribute optional server and browser Sentry events to their source build. Use validated build commits for Docker and source/npm artifacts, preserve explicit server release overrides, and keep cached browser bundles tied to the commit they loaded. Verify 127 focused tests, server/UI typechecks, Docker and source build stamps, all 53 CI checks, and Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2a99de80ec |
fix: supply public task links and preserve managed AI sessions (#13699)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors deliver agent replies to external conversations.
> - Agents need a public task link when a user asks to open the task.
> - The prompt and task tools lacked that link, so an agent could invent
an internal address.
> - Managed AI credential directories also changed the session
fingerprint on each run.
> - This change supplies public task URLs and excludes only those
temporary directory values from the fingerprint.
> - Follow-up messages can reuse compatible sessions while real
configuration changes still reset them.
## Linked Issues or Issue Description
Refs #13680 and #13694 for the related Cloud-origin fixes. No duplicate
open PR was found.
**What happened?**
An external chat reply could contain an invented internal task URL. The
publication filter then removed the link. Follow-up runs also lost their
saved provider session because each managed credential home used a
different temporary path.
**Expected behavior**
Agents receive the current public board URL for a task. Temporary
credential directories do not reset an otherwise compatible session.
Account, credential, model, permission, and custom environment changes
still invalidate it.
**Steps to reproduce**
1. Use a chat connector with a managed AI connection.
2. Ask for the current task link.
3. Send a follow-up message with the same agent configuration.
4. Inspect the task URL and the session reset reason.
**Paperclip version or commit**
Reproduced on the source at
|
||
|
|
b193077582 |
fix: report Slack callback health correctly behind Cloud proxies (#13694)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack connections turn messages into governed agent runs and return replies to Slack. > - Remote runs must replace an incompatible sandbox runner with the controller's packaged binary. > - The fallback used a package-relative path that does not match the vendored server layout. > - Slack callback health also compared the internal proxy address with the public callback address. > - Master now contains the runner fallback fix; this pull request fixes callback health and extends missing-runner regression coverage. ## Linked Issues or Issue Description Refs #13677, #13680, #13691, and #13686. **What happened?** A packaged controller stopped a Slack-triggered remote run with `runner_remote_artifact_unavailable` when the sandbox runner needed replacement. Working Slack callbacks also showed a stale URL warning behind the Cloud gateway. **Expected behavior** The controller stages its packaged runner when needed. Callback health uses the observed public address and still detects real address changes. **Steps to reproduce** 1. Run a packaged server with a sandbox that has an older runner or no runner. 2. Send a Slack mention to an agent that uses that sandbox. 3. Route signed Slack callbacks through a claimed Cloud gateway that rewrites the upstream host. 4. Check the run and the Slack callback health panel. **Paperclip version or commit** Reproduced on master at `aeef493f4a7603b7b1254421b80fb00212982390`. The fix branch also includes #13691. **Deployment mode** Packaged server with a Cloud gateway and a Daytona sandbox. ## What Changed - Extend the controller-owned runner fallback tests with a missing sandbox binary case; retain the fix now merged in #13686. - Prefer dedicated gateway diagnostic headers that survive provider rewrites of standard forwarded headers. Use validated host hints only for callback-health evidence on claimed Cloud instances after provider acceptance. Preserve request bodies, routing, authentication, and configured callback URLs. - Cover current, stale, and missing sandbox runners, all Slack callback surfaces, rejected callbacks, malformed proxy hints, real host and port changes, and the existing self-hosted behavior. - Document the artifact lookup and callback-health boundaries. ## Verification - Native session executor and binary resolver suites: 379 tests passed after merging current master (`45c99a0d0`). - Targeted callback integration suite with disposable PostgreSQL: 4 tests passed before rebase. - Full `pnpm -r typecheck` and `pnpm build` passed after merging current master. The full local suite passed 12,668 tests; one suite failed to start its disposable PostgreSQL. Rerunning that suite alone passed all 31 tests. - The built server resolver selected the executable under `server/dist/vendor/paperclip-runner/bin/`. - Final callback regression: all 4 targeted integration tests pass, covering provider header rewrites, default ports, uppercase/trailing-dot hosts, and ignored self-hosted hints. - Live staging proof of the runner fix: a previously failed Slack thread recovered, a new mention received its requested response, and the account-connect command succeeded. Unsigned callbacks returned 401. - Live browser and Slack acceptance passed: generated callback URLs, account linking, a new mention, an interactive question and answer, and all three callback-health indicators. The connector was activated through the onboarding UI. - All 54 latest-head checks pass, two optional checks are skipped, and Greptile is 5/5 with no unresolved threads. ## Risks - Runner fallback must select a binary for the remote platform. This preserves explicit remote artifact overrides and the existing capability checks. - Proxy headers are not identity proof. They are used only for diagnostics on claimed Cloud instances after the provider accepts the request. Self-hosted instances ignore them. Wrong public hosts and ports still warn. - No database migration or new public API contract. ## Model Used OpenAI GPT-6 via Codex. Used reasoning, repository tools, code execution, and browser/native-app testing. Exact model variant and context window were not exposed by the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
45c99a0d06 |
fix(adapters): default legacy harnesses and connected tools to full auto (#13693)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Legacy adapters launch provider CLIs and expose connected tools. > - Existing defaults did not consistently grant full automatic permission. > - Remote Claude used a fixed tool list that omitted MCP tools and future tools. > - Direct Codex launches and OpenCode configuration also used narrower defaults. > - This change gives all these paths the same full-auto default as native runners. > - Explicit restrictive settings continue to work. ## Linked Issues or Issue Description Refs #13686. This PR is stacked on that native-runner and task-reassignment PR. Merge #13686 first. Related: #831 (constructed Claude agents), #1935 (adapter-switching permission defaults). ## What Changed - Use actual Claude permission bypass for local and remote runs and probes. Remove the fixed tool list so MCP and future provider tools are included. - Identify actual managed sandbox targets to Claude with `IS_SANDBOX=1`. Do not mark ordinary host execution as a sandbox. - Default direct Codex execution to approval and sandbox bypass, matching agent creation. Preserve explicit false, CLI profiles, sandbox modes, approval policy, and network restrictions. - Set OpenCode's full-auto runtime permission to `allow` for every tool and connection. Preserve the existing explicit opt-out. - Default Gemini probes to the same YOLO mode as execution. Make the legacy ACP `default` alias use `approve-all` for fresh and resumed sessions. - Add default, opt-out, remote, probe, connected-tool, and resume regression tests. Update adapter configuration documentation. - Other adapter paths already request full automatic permission or have no provider approval gate. ## Verification - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server and legacy ACP tests passed, including defaults, explicit opt-outs, remote launches, connected tools, and fresh/resumed sessions. - Greptile reviewed current head `8ca135eaffcf9cfdba6f1368e896a781a0891d50` at **5/5**. The security reviewer acknowledged the documented full-auto requirement. Acknowledged discussions are resolved. - Current head has **54 passing checks**. [PR checks](https://github.com/paperclipai/paperclip/pull/13693/checks). The process-adapter signoff browser shard passed on one retry after its first attempt exceeded a three-second issue-run wait. - **Six native Claude/Codex real-provider cases passed on their first attempt, with cleanup passing**, against the combined branch: plans, reassignment, and backlog creation/status. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). This does not claim a real-provider run of every legacy adapter. - The live-tested revision is `a37881c824dcd7170380fc4b788732fc743e5da7`. The current head differs only in the corrected heartbeat test expectation; application code is identical. - The campaign result-enforcement job passed. The separate report publisher failed during frozen dependency installation because the trusted workflow's patched-dependency configuration does not match its lockfile. Passing case evidence remains downloadable from the workflow. - Full-suite coverage comes from CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. ## Risks - Missing permission settings now grant all provider operations, including connected tools. OpenCode full-auto also overrides ambient provider permission rules. An explicit Paperclip permission opt-out preserves restrictive behavior. - Claude refuses full bypass as root outside an identified sandbox. Ordinary host deployments must run Claude as a non-root user. Managed sandbox launches include the required marker. - These defaults do not grant additional Paperclip roles, connections, or company access. Existing controller authorization and governance still apply. - This PR depends on #13686. Retarget it to master after that PR merges. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
04546c82d5 |
fix(runner): reconnect Daytona sessions after controller restart (#13691)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner can execute a task inside a Daytona sandbox. > - The sandbox can keep running when the Paperclip controller restarts. > - Recovery treated sandbox process IDs as local process IDs and selected the wrong recovery path. > - Live verification also found races between startup, shutdown, and queued task cleanup. > - This pull request verifies the existing remote owner and orders those transitions. > - Users can continue the same task and provider session after a controller restart. ## Linked Issues or Issue Description **What happened?** The Daytona `recover-controller` cases failed with `runner_state_identity_mismatch`. Remote process IDs can be absent on the controller or collide with unrelated local processes. Recovery then looked for remote state in the local runner directory. Later turns could also start before the previous executor released its sandbox resources. **Expected behavior** Reconnect to the original sandbox and authenticated runner. Preserve the task, provider session, and queued comments. Reject a replacement sandbox or mismatched identity. Do not start another provider during reattachment. **Steps to reproduce** Run the `everyday-workflows` `recover-controller` case for `runner-codex` or `runner-acpx-claude` in Daytona. The browser creates a Python tool, requests a revision, restarts the controller during execution, and queues another revision. It then downloads and tests the final ZIP. Related: #13682 is the preceding operational fix. #13291 addresses legacy sandbox conversation recovery, a different execution path. #13666 includes broader run-capacity work; this change guards cleanup of an existing native task executor. ## What Changed - Add remote runner recovery without interpreting sandbox PIDs on the controller. - Verify the original provider lease, remote workspace, durable state, process marker, and authenticated PRP authority before adoption. - Compare the process marker with live Linux boot identity and start ticks to reject PID reuse. Read virtual proc files through the guaranteed Node runtime; unavailable proof blocks adoption without blocking a fresh launch. - Make the E2E supervisor own the actual server process so forced restart cannot leave a late database closer behind. - Scope the chat delivery lease test to its own fixture instead of draining other tests’ pending deliveries. - Preserve provider-attempt counts and recorded evidence during reattachment. - Serialize an idle-session checkpoint with admission of the next native turn. - Wait for an in-progress startup to acknowledge restart detachment. Fail after a bounded deadline if it cannot. - Keep a queued comment waiting until the previous native task executor releases its resources. Allow unrelated tasks to continue. - Update the Daytona image's resolved lock digest to match current dependency manifests. - Add classifier, ownership, process, startup, checkpoint, and queued-admission regression tests. Document recovery behavior. ## Verification - 415 focused tests passed across native execution, restart recovery, workspace synchronization, queued admission, and real-process restart tests. The final Node-based fingerprint change passed all 375 native-session tests. - Runner harness unit tests: 394 passed. Chat integration shard 2: 335 passed after fixture isolation. - The exact fingerprint command succeeded twice in a disposable Daytona sandbox and returned the same identity; the sandbox was deleted. - 11 real-process restart integration tests passed, including absent and colliding remote PIDs. - Repository typecheck and final build passed. Broad local checks found machine-dependent database startup and timing failures; focused retries passed. The final-revision PR pipeline is green. One unrelated browser shard hit a five-second blank-page timeout on the first run and passed its targeted retry. - Final-revision local headed browser E2E: `everyday-workflows.runner-acpx-claude.daytona.recover-controller` passed on attempt 1 in 4.7 minutes, **40/40 checks**. Manual browser inspection confirmed Done, all three ZIPs, and delivery of the queued follow-up. All three runs succeeded using the same provider session. The harness downloaded and independently tested the final artifact. - Final-revision Daytona campaign: https://github.com/paperclipai/paperclip/actions/runs/35463999611 — **Codex passed first attempt (4.8 minutes); ACPX Claude passed first attempt (6.1 minutes)**. Campaign aggregation/publication is finishing; both test jobs succeeded. - Greptile reviewed `beb08d8493b3286f5bb988dead369ff8c96a395d`: **5/5**, no open findings. - Staging browser verification is pending selection of a disposable staging instance and removal of a Chrome extension UI block. ## Risks - Recovery now depends on the original sandbox remaining available. A replacement or mismatched identity still blocks adoption. - Shutdown waits up to 30 seconds for a native startup to reach a safe detach point. An unfinished startup returns a clear failure instead of a false detach receipt. - Queued native work on the same task waits for cleanup. Unrelated tasks remain eligible. - The image digest update rebuilds the Daytona runtime image. No database migration or public API change is included. ## Model Used OpenAI Codex, GPT-6, with repository inspection, code execution, and browser tools. The runtime does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b70641f23f |
feat(plugins): support image catalogs and persistent application overlays (#13646)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Plugins extend the application without adding each integration to
Core.
> - A downstream image needs a way to supply prebuilt plugins.
> - Some plugin UI must stay mounted as users move between pages.
> - This change adds an image catalog and a persistent application slot.
> - Operators can upgrade or remove these plugins through their image
and configuration.
## Linked Issues or Issue Description
**Subsystem affected**
Plugin packaging, activation and application UI.
**Problem or motivation**
The built-in plugin catalog is fixed in Core source. Downstream images
cannot add entries through an explicit catalog. Existing page slots also
cannot preserve a small application overlay across route changes.
**Proposed solution**
Read a bounded catalog of prebuilt plugins from the image. Verify its
files before importing manifests. Use the existing managed selection and
plugin lifecycle. Add an `appShellOverlay` slot with account and company
cleanup.
**Alternatives considered**
A downstream fork adds merge work. Script injection provides no plugin
lifecycle. A separate runtime download system adds a second distribution
channel.
**Roadmap alignment**
This extends the existing plugin system. Related PR #9006 covers runtime
install replication; this change covers immutable image contents. PR
#12555 covers CLI scaffolding. Neither provides this catalog or
application slot. The maintainer requested this work directly.
## What Changed
- Validate catalog identities, confined paths, package versions and
bundle hashes before importing code.
- Apply image selection to persisted plugin installs, including removal
and rollback. Adopt the verified image path from existing npm/local
installs and bind runtime worker/UI entrypoints to verified package
declarations.
- Mount application overlays in both UI shells. Preserve route state and
clear it on account, company and onboarding changes.
- Restrict service-worker offline storage/fallback to hashed public
assets in a separate cache namespace; exclude application HTML and
extension/API data, including after worker restart.
- Document the packaging contract, trust model and rollback
requirements.
## Verification
- Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`. Affected server/UI typechecks and builds, plus token
gates, passed again after rebasing onto current master; the 124 focused
tests also passed after rebase.
- Latest focused verification: 124 tests in nine files passed for
catalog/reconciliation/loader, overlay lifecycle, Layout and
service-worker policy. The broader UI/shared/SDK run passed 7,204 tests
in 690 files with canonical `TMPDIR`.
- Real disposable Core/PostgreSQL: catalog install, selection removal,
0.1.0→0.1.1→0.1.0, same-version npm/legacy-path adoption, and
preservation of disabled status passed. Added permissions entered
`upgrade_pending`, withheld UI across restart, and activated only after
explicit operator enable.
- Real Chromium: desktop/mobile layout, route draft retention and
Escape/focus passed with mocked extension responses. A persistent
browser restart retained public hashed-asset offline fallback while
refusing seeded legacy/current private entries and legacy HTML.
- Full `pnpm test:run`: 12,539 passed; 17 failed across six existing
files, stopping later phases. macOS read-only directory renames fail in
runtime-skill-cache and company-skills-service; email tests require an
absent local AgentMail fixture. Native runner/comment-redaction passed
in isolation after temporary Rust setup; agent-conversations also passed
in isolation. No unrelated source was changed to hide failures.
- After rebase, two unchanged chat timing tests failed in CI and passed
locally in isolation. Their CI shard passed on its single retry. All
other current-head CI jobs passed on the initial run; review is 5/5 with
no unresolved threads.
- No live deployment or external plugin service was used.
## Risks
- Plugins are trusted code. The catalog detects packaging errors; it
does not authenticate an untrusted image builder.
- Invalid catalogs fail startup. Images must contain the catalog and
bundles together, with stable directories.
- A host older than this contract lacks the activation guard. Disable
added plugins and remove their configuration keys before reverting to
it.
- Offline navigation now returns 503 instead of replaying cached
application HTML. Only public build assets have offline fallback.
- Rolling back an unapproved permission change retains the approval
gate; review the current manifest and explicitly enable it. A reduced
permission set cannot establish prior approval or prior enabled status.
- Plugin data migrations need their own rollback policy. This change
retains installed records and does not reverse migrations.
## Model Used
- OpenAI GPT-6 (Codex), model ID `gpt-6`, with repository inspection,
code execution and browser verification. The runtime does not expose an
exact context-window size.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites; broad
macOS server-run exceptions are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (fresh run on 488b3754ae; chat
shard passed its single retry)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(fresh review on
|
||
|
|
f589660ec0 |
feat(routines): add safe webhook setup and in-routine run management (#13637)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Routines turn scheduled work and external events into tasks for an assigned agent. > - Webhook setup was disabled, and actor authentication rejected valid webhook bearer keys. > - Operators need to connect and test a sending app before events can start work. > - This pull request adds a guided setup with durable connection tests that cannot dispatch a task. > - It keeps trigger management, execution tasks, and activity within the routine. > - The benefit is a webhook that can be configured, verified, and operated from one place. ## Linked Issues or Issue Description Fixes #11937. Related: #13216 adds provider-specific Sentry support. This PR addresses general routine setup and ingress. #6841 addresses legacy secret bindings; this PR retains the existing secret service. **Current behavior** Webhook creation is disabled. Bearer deliveries can fail in agent authentication before the routine checks its key. Setup has no safe connection test. Runs and Activity send the operator away from the routine. **Proposed behavior** Choose a schedule or a webhook. Follow the setup steps, copy credentials or complete agent instructions, and test delivery without creating work. Finish setup to allow future events to start tasks. Edit or remove compact trigger cards, undo removal, and inspect tasks and activity inside the routine. **Reason and benefit** An operator can verify credentials and delivery before enabling automatic work. Durable setup state survives refreshes and restarts. Retry receipts prevent an old test event from starting work after activation. ## What Changed - Add a production trigger wizard using reusable Slack setup navigation and footer components. - Add schedule and webhook choices, one-time credentials, agent instructions, and live connection feedback. - Persist pending setup, test delivery receipts, connection status, and reversible trigger removal. - Keep setup checks free of routine runs, tasks, and agent wakeups. Preserve delivery idempotency after activation. - Add compact trigger cards, inline editing, key rotation, pause controls, removal, and Undo. - Keep Runs and Activity in the routine. Use the shared task list and compact activity rows. - Permit only exact public delivery POSTs through actor authentication. Retain webhook authentication, JSON-object validation, and log redaction. - Add production-backed Storybook states and focused server, database, and UI coverage. - Document signing modes, setup checks, retries, rotation, HTTPS ingress, and navigation. ## Verification - Full workspace typecheck, build, and token gates passed on the rebased branch. Storybook also builds. - Focused routine, middleware, logging, shared wizard, and UI coverage passes on the rebased branch: 195 tests across 14 files. The migration passed on a fresh PostgreSQL database and on two repeated applications. - Browser testing used the real app, database, and a deterministic process worker through Tailscale HTTPS and the current Cloud proxy code. - Verified rejected keys, safe setup deliveries, persisted state after restart, activation, retry deduplication, key rotation, schedule editing, removal, and Undo. - Fresh bearer and GitHub-signed deliveries created tasks that the worker checked out and completed. Runs and Activity stayed within the routine. - Current Cloud ingress tests passed. Public delivery POSTs passed through without a browser session; management routes remained gated. - All 54 current-head PR checks pass, including general and serialized tests, all eight browser E2E shards, typecheck, build, runner checks, security checks, and the canary dry run. Two optional Storybook jobs are skipped by workflow conditions. - Greptile is 5/5 on commit `7ea63a61e`, with no unresolved review threads. The stale connection-status finding is fixed and covered by a regression test. - No production deployment was performed. ## Risks - Migration 0281 adds three trigger columns and a test-receipt table. It is additive and safe to reapply. Apply it before running the new server. Existing triggers remain live by default. - Requests without delivery IDs are new events after activation. Senders must reuse an event's delivery ID for retries. - Completed webhooks keep normal dispatch behavior. Their management connection check can start work; the UI states this. - Removing a trigger archives it. Undo restores the URL and credentials. Permanent deletion remains available through the existing API. - Public ingress must remain restricted to the delivery POST route. The tenant verifies credentials. Cloud sleeping-stack behavior is unchanged. - Shared setup components also serve Slack. Existing setup contracts and navigation tests cover that integration. - Senders must use application/json with an object. Other media types receive 415. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
20d26117c9 |
fix(chat): use the claimed Cloud origin for connector URLs (#13680)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors publish callback URLs and links to the board. > - Cloud can assign a warm instance its final origin after the server starts. > - The signed runtime identity already tracks that change. > - The chat service kept a copy of the startup origin and continued to publish it. > - This pull request resolves the trusted origin when it creates each URL. > - New connector setup uses the claimed hostname without a server restart. ## Linked Issues or Issue Description **What happened?** Chat setup in a claimed warm instance used its old pool hostname in provider callbacks. Account confirmation and task links could also use the old hostname. **Expected behavior** Chat URLs follow the signed canonical origin after the claim. An explicit webhook ingress override still applies only to provider callbacks. Self-hosted URL precedence stays the same. **Steps to reproduce** 1. Construct the chat service with a pool origin. 2. Apply the Cloud claim without restarting the service. 3. Open Slack setup or create an account-linking intent. 4. Observe the startup hostname in the returned URL. **Paperclip version or commit** Reproduced on master at `9335b7db1`. **Deployment mode** Paperclip Cloud warm-instance claim. Related: #12766 introduced the signed canonical runtime identity. ## What Changed - Resolve the signed Cloud origin when building chat setup, account confirmation, and task URLs. - Use the same callback origin for Telegram registration and GitHub webhook recovery. - Preserve explicit webhook ingress and self-hosted configuration precedence. - Add regression coverage for existing and new endpoints across Slack, GitHub, Teams, and Telegram. - Document the origin precedence and the need to update callbacks already saved at a provider. ## Verification - Reproduced both new regression cases against the original code. - Full chat integration and signed Cloud identity suites: 1,015 tests passed after the production-code correction. - Five origin and ingress cases passed after review additions, including Telegram registration and GitHub webhook repair after a live claim. - Focused origin, ingress, callback, task-link safety, and tenant-isolation checks: 67 passed. - `pnpm -r typecheck` and `pnpm build` passed. Server typecheck and compilation passed again after the task-link validation correction. - A broad local `pnpm test:run` started before the correction was stopped after the final-commit CI suite passed. It is not counted as a passing local run. - Final-commit CI: 54 successful checks; two optional Storybook checks skipped. Greptile: 5/5 with all review threads resolved. - No live deployment or Slack app mutation was performed. ## Risks - Cloud chat URLs now follow the signed runtime identity. Request host headers cannot set this value. - An explicit webhook ingress override still takes precedence for callbacks. - Existing Slack app settings are external state. Operators must replace an old callback URL in Slack. - This change does not deploy the app or change gateway ingress policy. No schema migration is required. ## Model Used OpenAI Codex (GPT-6), with repository search, code execution, and automated tests. The exact model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c1f6c3310a |
fix(runner): repair catalog runtime and grading boundaries (#13676)
## Thinking Path > - Paperclip manages tasks across persistent agent sessions. > - The full Runner E2E catalog exposed failures in session restoration, tool validation, and test controls. > - These failures prevented valid work from resuming or made a valid interaction fail the test. > - Invalid completion reports also reached finalization before the provider received useful feedback. > - This pull request repairs those boundaries without changing production prompts or approval policy. > - Focused regressions and fresh paid cases verify each fix. ## Linked Issues or Issue Description Follow-up to #13655. Stacked on the trusted worker prerequisite fix in #13674. **What happened?** Read-only skill uploads failed in resumed Daytona sandboxes. Invalid criterion IDs escaped tool validation. A progress event could park a run before its tool response settled. Partial question forms hid required answers. Two test assumptions rejected valid plan keys or failed to navigate an optional question page. **What did you expect to happen?** Resume identical skill bundles, give repairable feedback for malformed completion calls, preserve in-flight tool responses, show all required questions, and test the rendered workflow accurately. **Steps to reproduce** Inspect the failed cases in https://github.com/paperclipai/paperclip/actions/runs/35417932353. Fresh campaigns: https://github.com/paperclipai/paperclip/actions/runs/35444497313 and https://github.com/paperclipai/paperclip/actions/runs/35445327618. The later backup cleanup is tested in https://github.com/paperclipai/paperclip/actions/runs/35446477285. Combined report: https://pages.paperclip.ing/runner-e2e-operational-35444497313/investigation.html. ## What Changed - Compare immutable archives before reusing read-only Daytona bundles. Reject corrupted content and preserve unrelated files. - Validate exact criterion IDs before accepting completion. OpenCode returns a tool error instead of emitting a result that terminates runnerd. - Complete the activity item for rejected OpenCode calls. - Remove retired read-only harness backups without altering live files or following symlinks. A fresh paid rerun exposed this later checkpoint-cleanup failure. - Exclude progress messages from the governed-wait completion boundary. - Reject newly created question forms that omit questions or contradict their stored answer semantics. Keep historical rows readable. - Navigate all rendered question pages and recognize revision-bound descriptive plan keys in the continuation suite. ## Verification - Harness unit suite: 383 tests pass. Harness typecheck passes. - Native session executor and status corpus: 381 tests pass. - Shared question and interaction-service tests: 42 pass; native question bridge and executor: 360 pass. Daytona sync: 21 pass, including foreign-owner archives and corrupted immutable content. - OpenCode driver: 29 tests pass, including wrong, missing, and duplicate criterion IDs followed by a valid retry. - Repository typecheck and build pass. The later OpenCode activity fix also passes its package build. - The latest commit passes all 52 PR checks and Greptile 5/5. The backup-cleanup fix also passes 351 related local tests and server typecheck. Local full-suite coverage completed across runs. adapter-auth-signal-routes and pipelines-routes encountered transient socket resets; both pass on retry, and all remaining 24 serialized files pass. Paid reruns are complete: 27 of 29 unique cases pass using the latest recording per case. Both Daytona controller-restart cases still fail with runner_state_identity_mismatch; the report describes this remaining runtime issue. Eight affected cells need #13674 on master before their rerun. ## Risks Creation rejects inconsistent dual question representations but does not change historical records. Immutable bundle comparison must verify bytes before skipping extraction. Completion feedback must use the contract bound to the current run. Durable suspension and approval checks remain enforced. Production prompts are unchanged. ## Model Used OpenAI GPT-6 via Codex, with repository inspection, code editing, and test execution. The exact API model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
36dbb7ed1c |
fix: harden agent chat runner tools and recovery (#13678)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat turns discussion into plans, tasks, reviews, and hires. > - These workflows need reliable tool results and task context on the native runner. > - Live Claude and Codex tests exposed lost retry requests, invalid project inputs, and a child startup crash. > - Recovery also exposed a misleading retry action and missing child task context. > - This pull request fixes those paths and adds regression coverage. > - Agents can continue the original request and operators can inspect a stopped run. ## Linked Issues or Issue Description **What happened?** A failed Agent Chat retry could lose the user's question. Project creation accepted unsupported icons in its tool schema. Codex could stop when a helper's MCP startup event arrived before its thread lineage. A stopped task offered Retry even when the server required execution reconciliation. Resumed agents could miss existing delegated tasks. Hiring and review instructions did not describe the native runner's available tools and source requirements. **Expected behavior** Retries retain the selected request. Tool schemas match the API. Child startup information does not gain authority over the parent or stop it. Recovery actions match the server's requirements. Task context exposes existing child work. Handoffs contain the material the assignee needs. **Steps to reproduce** 1. Enable experimental Agent Chat in an isolated development instance. 2. Configure native Codex and ACPX Claude agents on Paperclip Runner. 3. Ask for a plan, revise it, approve task creation, and request a hire and status report. 4. Retry a failed chat turn and check that it answers the original request. 5. Start a Codex helper before its thread lineage arrives. 6. Resume a delegated task and inspect its existing children and saved output. **Paperclip version or commit** The live failures were found at `f2c5e54dc`. This branch is rebased onto `86b7ee992`. **Deployment mode** Isolated local development instance with native Codex and ACPX Claude. No database migration or default permission change. Related work: Refs #13284 for Agent Chat. Refs #13438 for the server-side API receipt fix, which this branch preserves. The transport also accepts the earlier HTTP receipt format. Refs #13655 for the current Codex continuation and helper lineage handling, which this branch also preserves. ## What Changed - Preserve failed Agent Chat wake-comment IDs and session generation from the authorized source run. Reject pre-reset retries. - Wrap API receipts with the correct semantic call identity. Test current and earlier receipt formats through real HTTP and runnerd. - Classify early child MCP startup notifications as information. Keep foreign completion and result events rejected. - Constrain project icons on both tool surfaces and regenerate the protocol contracts. - Include bounded, company-scoped visible direct child tasks in task context. Filter hidden tasks before applying the limit. - Replace the rejected Retry action with Inspect run for native continuation reconciliation. - Update hiring, review handoff, status reporting, and development guidance. ## Verification - Live tests covered Claude and Codex questions, plan revisions, approval, task creation, hiring, status, chat reset, failures, and recovery. - The recovered task produced its saved checklist and example. A later follow-up read the existing child tasks and document without creating more work. - Full build, repository type checks, token gates, 142 focused tests, 188 runner TypeScript tests, and the Rust notification/descendant regressions passed after rebase. The separate local full-suite run was stopped after the complete CI suite passed. - Review fixes passed the updated route, tool-authority, and icon regression tests plus server type checking. - Required commands: `pnpm build`, `pnpm -r typecheck`, `PAPERCLIP_IN_WORKTREE=false pnpm test:run`, and `pnpm check:token-gates`. - At `4ce8047b0`, all 55 applicable GitHub checks pass (two Storybook checks are intentionally skipped), including the complete general/serialized test matrix, runner tests, browser tests, build, type checks, Docker checks, and canary dry run. - Fresh Greptile review is 5/5 on `4ce8047b0`; all three findings were fixed with regressions and there are no unresolved review threads. - Two initial CI service-startup timeouts passed unchanged in local reproductions and in the latest CI run. ## Risks - The new event classification is limited to MCP startup information. It does not authorize foreign task completion, results, or tool requests. - Task context returns at most 100 direct child tasks and reports truncation. It excludes hidden tasks and other companies. This improves delegation context but does not enforce semantic duplicate detection. - Native reconciliation still requires an operator to inspect and record prior outcomes. The new link does not replace the recovery API. - API tools remain opt-in. Claude permission choices remain explicit. No default permission, schema, or workflow changes. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository editing, code execution, API tools, and browser testing. The exact deployment identifier and context-window size are not exposed in this session. Live acceptance agents used OpenAI `gpt-5.6-sol` and Anthropic `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9335b7db10 |
fix(runner): validate inherited environments and replace stale sandbox binaries (#13677)
Resolve the effective environment for account adoption and adapter tests. Preserve saved-agent overrides when the request omits environmentId, and treat explicit null as inheritance from the instance. Reject sandbox runners that lack unlimited-runtime and connection-lease-renewal capabilities. Stage the bundled runner before launch when the image binary is stale. Add regression coverage for environment precedence, fail-closed validation, adapter switches, API-key reverification, and runner artifact fallback. Document the operational workaround for older controllers. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
86b7ee992c |
feat(onboarding): ClipLab sleepy-to-wake hero and step hand-offs (#13629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents have a persistent visual identity (#13171): a ClipLab character in one of 17 palettes, rendered as cached PNGs in lists and as a live character in larger placements. > - The onboarding wizard is where a person meets that identity first, and it showed a stock ClipLab expression on the previous engine while the rest of the app would show a different character on a newer one. > - The wizard's steps also cut from one screen to the next, so the arc read as separate pages rather than one walk. > - This pull request puts one character on one engine everywhere, gives the wizard's hero the studio's sleepy → wink → idle sequence on Review, and hands the steps over inside one presence. > - The benefit is that what wakes on Review is exactly what the agent looks like on the dashboard afterwards, and the walk to it reads as one screen changing. ## Linked Issues or Issue Description Refs #13171, now merged into master. This PR contains the onboarding and ClipLab update on top of that foundation. Original feature work by @tonio-alucema; merge preparation preserves the original commits. **Problem or motivation** The onboarding hero and the app's avatars were two different characters on two different ClipLab engines. Steps 1 → 4 of the wizard cut between screens, and the wizard mounted cold when a cloud-managed workspace arrived from Cloud's naming screen. **Proposed solution** Vendor ClipLab v0.2.0 as the shared engine and render one studio-exported character from it in every palette, for every pose and size. Play the export's one-shot wake on Review with the palette fading in over the gray dormant loop. Hand steps over inside one presence so the footer slides instead of jumping, and play the arrival half of that hand-off when the wizard opens directly on the agent step. **Alternatives considered** Exporting mp4/webm loops per size: no cursor following, no clean alpha, and the palette "colour in" is a runtime blend. Minting a `cap-v2` character version: nothing had shipped `cap-v1`, so the artwork is regenerated in place instead of migrated. Keeping the separately vendored runtime bundle for the hero: two engines and two characters in one app. ## What Changed - `packages/shared/src/cliplab`: re-vendored from ClipLab v0.2.0 (`987b6db0`) with the Paperclip adaptations replayed (optional graphics backend for the Node SVG snapshot path, supersampled live textures, character framing, deterministic SVG id prefixes); new upstream `particles.ts`. - `packages/shared/src/cliplab/character.ts`: the studio export, mirrored from `ui/src/assets/cliplab/onboarding.character.json` by `scripts/sync-cliplab-character.mjs` (drift caught by `check:token-gates`). `characterDefinition` builds every palette from it; the resting portrait is its idle beat. - `OnboardingCharacter`: gray `sleepy` loop through the agent and connect steps; on Review the one-shot sleepy → wink → idle plays on two lock-step canvases while the palette fades in, then the `idle` loop. Body-follows the pointer, page-scoped. 160px in the wizard. - `OnboardingWizard`: steps 1 → 2 → 3 → 4 hand over inside one `AnimatePresence` (departing content fades and gives its room back; arriving content opens its room then fills); the hero has a room that opens on the walk into the agent step; opening directly on the agent step plays the arrival half; the self-hosted naming step uses the arc's label and field. - Motion vocabulary in `onboarding-motion.ts` (`stepContentMotion`, `ledeMotion`, `heroRoomMotion`, `heroRoomArrival`, `titleSwapMotion`). - Storybook: `Onboarding / Character` (Wake Up), `Onboarding / Agent arc` walkable from the naming step plus `Arrive From Cloud`; the companies fixture answers the wizard's create call with a company. - Uses the shared runtime for onboarding; `doc/agent-personas.md` documents the shared character. - Releases both onboarding canvases after partial startup or transition failure. Registers each canvas before seeking so synchronous render errors can release it. Six component tests cover these failures and palette changes before or during wake. - Refreshes both sleeping canvases when the palette changes, including a palette change in the same render as wake. - Moves choreography values into the CSS token layer and preserves the shared motion catalog drift check across the imported stylesheet. - Repairs the static Storybook avatar route and uses accessible heading names/current button labels in the wizard play functions. - Closes the lazy avatar worker pool during application shutdown. ## Verification - Merge-preparation checks: `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm check:token-gates` pass. The final UI typecheck and 123 focused onboarding, lifecycle, and token catalog tests pass. All 55 checks on final head `b4f5e201a1564083d163abc6f93f5b3da06ccefd` pass, including the full sharded test suite, runner verification, and all eight browser shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/35445430535)). The duplicate monolithic local `pnpm test:run` was stopped after CI completed; it is not claimed as a separate completed local run. - Chromium walkthrough: palette change, wake, return to sleep, WebGL failure fallback, Review step hand-offs and cloud arrival pass with normal and reduced motion; no browser errors. The signoff happy-path browser test also passes against a disposable instance. - The final CI run confirms the catalog fix and a passing signoff browser shard. The earlier signoff failure was a heartbeat-run availability timeout; the focused local reproduction and final CI passed without signoff code changes. - Original author verification: - `pnpm check:token-gates` (includes the new character sync check); shared, server avatar/persona (17) and UI onboarding/persona (137) suites pass; `pnpm build-storybook` packages all 3,564 avatar PNGs through the worker pipeline. - Storybook: `Agents / Personas` Sizes, Expressions and Palettes render the studio character at every size and pose; `Onboarding / Character → Wake Up` plays the wake on the shared engine; `Onboarding / Agent arc` walks 1 → 4 with the hand-offs, and `Arrive From Cloud` plays the arrival (measured: content room 6 → 65px over 320ms, fade to 1.0 by ~560ms, footer travel continuous). - The original author walked the agent → connect → review flow and wake after a real sign-in on staging. - Not done here: the Linux Storybook visual baselines (`tests/storybook-visual/agent-personas.spec.ts`) need re-baselining for the new engine, hero size and naming-step changes. ## Risks - Every avatar's pixels change (new engine, new character) under the unchanged `cap-v1` name. Stacks that rendered avatars on the previous engine keep those PNGs in their cache (`generated-agent-avatars/cap-v1/...`, served immutable) until cleared; only the two pinned staging stacks ever did. - The one-shot handoff to the idle loop is timed from the sequence's authored duration (the engine reports completion by continuing into idle itself); presentation only, nothing in the wizard's state waits on it. - Reduced motion skips the wake and the hand-offs; jsdom is treated the same way, so the wizard tests see the next step's content immediately. - The committed export differs from the studio by one animation (Loop off, leading idle step removed); a re-export without that fix would play a 5.6s idle before the wake. ## Model Used Original feature: Anthropic Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with shell, browser, and file tools. The original context window was not recorded. Merge preparation and lifecycle regression fixes: OpenAI GPT-6 in Codex, with reasoning, shell execution, file editing, GitHub CLI, and automated tests. The session does not expose an exact runtime model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1ef3b08714 |
feat(ui): integrate agent personas across the app (#13171)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A stable agent persona is useful only when the same identity appears across the app. > - Lists, task messages, selectors, and activity feeds need inexpensive static avatars. > - Onboarding and agent headers need a larger character with expressions and pointer tracking. > - This pull request connects the persona foundation to those existing views and preserves onboarding draft assignments. > - Full-page stories and Linux checks make the placements and performance contract reviewable. ## Linked Issues or Issue Description **Problem or motivation** Agents need a stable visual identity in lists, tasks, onboarding, and configuration. External tools also need an image URL for that identity. **Proposed solution** Assign each agent a permanent palette from a fixed ClipLab character library. Store the assignment on the agent. Render and cache preset PNG URLs on demand. Use static images in dense views and one animated character in larger placements. **Alternatives considered** A generated image bundle requires a separate asset build. A live renderer in every avatar adds unnecessary work in large lists. Arbitrary uploaded images do not provide the requested shared character system. **Roadmap alignment** This improves agent identity across existing control-plane views. It preserves agent permissions, company boundaries, and status labels. ROADMAP.md has no separate ClipLab persona milestone. Related approaches: #2422 adds configurable image URLs and DiceBear generation; #5578 adds optional uploaded avatars. This work uses a fixed, versioned character library and preset URLs. ## What Changed - Replace agent icons with static persona images across lists, the sidebar, org charts, tasks, comments, selectors, activity, and dashboard views. - Put one animated character in the agent header. Let it follow the pointer across the page, with reduced-motion and touch fallbacks. - Add larger padded characters to agent creation. Keep the palette stable across draft refreshes and connection retries, then reveal it after success. - Pass appearance through shared projections rather than fetching each agent separately. - Add real full-page Storybook examples for the agent list, overview, task, dashboard, new-agent dialog, and connection page. - Add Linux screenshot, clipping, density, and 500-avatar performance checks. ## Verification - `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased tree. Persona lifecycle tests pass. - The rebased feature passes 38 Linux screenshot/performance checks, including both display densities, corner pointer positions, and the no-WebGL/no-live-download contract for 500 avatars. - The final Linux persona suite passes all 38 visual, lifecycle, density, and full-page checks using the standard Storybook configuration and real on-demand avatar endpoint. - Final local focused verification: 45 avatar/native-recovery tests pass; UI identity/routine tests, typecheck/build, token gates, and Storybook build pass. - Current-head CI passes: full workspace/server tests, all serialized server groups, typecheck/release checks, build, canary validation, and end-to-end shards. The build passed after retrying a native-runner concurrency-test failure; its three targeted cases also pass locally. - Manual inspection covered stable identities in the app, header placement, full-page mouse tracking, onboarding size, and task/dashboard placements. ### Screenshots Linux captures use synthetic Storybook fixtures. Full-page captures use reduced motion. The live character, mouse tracking, and disposal are checked separately. <details> <summary>Agent overview with the character in its header</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png" width="900" alt="Agent overview with the character in its header" /> </details> <details> <summary>Task messages and assignee identity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png" width="900" alt="Task messages and assignee identity" /> </details> <details> <summary>Larger onboarding character with room for expressions</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png" width="900" alt="Larger onboarding character with room for expressions" /> </details> <details> <summary>Dashboard agent activity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png" width="900" alt="Dashboard agent activity" /> </details> ## Risks - This PR depends on #13170, the persona foundation. Merge the foundation first, then retarget this PR to master. - Many placements change from icons to character silhouettes. Human avatars and authoritative agent status labels retain their existing behavior. - Only one character can render live per view. Reduced motion, hidden/offscreen content, touch input, and renderer failures use the defined fallbacks. - The full-page stories use fixture data. They do not contact a real company or complete real provider sign-in. ## Model Used OpenAI Codex, GPT-6 family. The exact model identifier and context window are not exposed in this session. Used code editing, shell execution, browser inspection, and Linux visual testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Tonio <tonework@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
43acbcc398 |
fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task state to provider sessions. > - Follow-up turns must retain provider memory and carry new user direction. > - Lost session IDs caused repeated context and extra input tokens. > - Native question answers and approval races could leave valid work blocked. > - This pull request repairs those paths and adds regression coverage. > - Agents can continue accepted work without repeating the conversation or losing the user's answer. ## Linked Issues or Issue Description Refs #13574. That merged PR shortened continuation prompts and moved question instructions into tool documentation. This change preserves sessions and fixes failures exposed by broader testing. Related runtime work: #13408 and #13410. **What happened?** Native follow-up turns could lose the provider session ID. Completion guidance could replace the original task with its latest comment. Claude native questions could remain pending after the user answered. Approval during a running tool call could suspend the run before the tool response arrived. Onboarding and chat handoff instructions also caused repeated planning or missing plan documents. **Expected behavior** Reuse a valid provider session. Send only new events when that session already has the history. Preserve the task requirements and apply later user direction. Store the question answer and deliver it to the waiting run. Finish governed tool responses before suspending. Execute the accepted plan without asking for the same approval again. **Steps to reproduce** Run the continuation, local-session-integrity, first-task, and agent-chat suites with native Codex and Claude. Include provider-question-bridge, accept-while-running, and plan-handoff. **Paperclip version or commit** This branch is based on master |
||
|
|
3ff3b34e15 |
Return safe client errors for malformed JSON requests (#13660)
Malformed JSON requests currently reach generic crash handling and return 500 before any route handler runs. Classify the specific Express parser error as a 400 with a constant response, preserving unrelated server error reporting. Carry forward the original three commits from #7410, preserve contributor credit, and add request privacy and negative regression checks. Document the API response. Fixes #7390. Validation: 80 focused tests, direct server typecheck, and complete Linux CI passed. Greptile 5/5. Full local typecheck/build require missing Rust tooling; local test environment failures are documented in the PR. Co-Authored-By: developers-universe-1 <madelynreyes2026@gmail.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e18ed02a56 |
Validate heartbeat run IDs before database lookups (#13657)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators inspect heartbeat runs, logs, and provider traces through the API. > - These routes use UUID database keys. > - A malformed path value such as `undefined` reaches the database and causes a server error. > - This pull request validates run IDs before those lookups. > - Valid requests keep the existing company and permission checks. ## Linked Issues or Issue Description Related: Refs #8135. That proposal guards actor run headers and activity writes; this fix covers heartbeat run path parameters. **What happened?** A request such as `GET /api/heartbeat-runs/undefined` passes a non-UUID value to a UUID lookup and returns a server error. Run logs and the other heartbeat run endpoints have the same unchecked path input. **Expected behavior** Reject malformed run IDs with HTTP 400 before any run lookup. Preserve valid run reads, company isolation, and the existing board and instance-admin checks. **Steps to reproduce** 1. Request a heartbeat run endpoint with `undefined`, `null`, another malformed ID, or a UUID with surrounding whitespace. 2. Observe the database UUID error. 3. Run the route regressions before and after this change. **Paperclip version or commit** Confirmed on master at `6d0342868`. **Deployment mode** The defect affects API deployments backed by PostgreSQL. Regression tests exercise the actual Express routes and authorization code with stubbed services. The historical request does not identify the originating client, so this change does not alter a guessed UI caller. ## What Changed - Share strict run-ID validation across the 12 heartbeat run endpoints in the agent router. - Keep existing board and instance-admin gates ahead of validation. Keep valid-run company and telemetry checks intact. - Reject surrounding whitespace, which the shared UUID helper accepts but PostgreSQL rejects. - Encode the UUID constraint and document the 400 response in OpenAPI. Test the generated parameter pattern on all 12 endpoints. - Cover malformed IDs on every affected endpoint, uppercase UUIDs, missing and cross-company runs, and permission precedence. Use UUID-shaped run fixtures in existing route tests. ## Verification - Before the fix: four malformed-ID regression cases fail; three access-control cases pass. - Focused agent route, permission, cross-company, and OpenAPI suites: 154 tests passed, including the final uppercase-UUID case. - Direct server typecheck passed: `pnpm --filter @paperclipai/server exec tsc --noEmit`. - The full local test attempt is still running; the complete Linux test suite passed in CI. Local results will be recorded when it finishes. - Complete Linux CI passed on the exact head: 53 checks passed, two non-applicable checks skipped. One untouched preview-service readiness test failed initially; its three targeted cases passed locally and the failed-jobs-only CI rerun passed. Greptile scored the final head 5/5 with no unresolved comments. - Full local `pnpm -r typecheck` and `pnpm build` reach the native runner step and stop because `cargo` is absent. The complete Linux CI checks passed. ## Risks Low risk: this changes malformed route inputs to HTTP 400. Valid UUID requests keep their existing lookup and authorization paths. There is no migration, dependency, provider operation, or configuration change. The separate activity router and actor run headers are outside this change. ## Model Used OpenAI GPT-6 through Codex, with code editing, shell execution, and test tools. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; full-check limits are recorded above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6d03428682 |
Skip task-only connector reads for agent chat views (#13654)
Agent chat views reuse the task surface with synthetic chat-prefixed IDs. Skip their task-only email and external chat-binding queries, and reject invalid UUIDs after existing authentication checks on both read routes. Preserve normal task reads and company isolation. Verified failing regressions before the fix, all 6,377 UI tests, route and OpenAPI regressions, server/UI TypeScript checks, all Linux PR CI gates, and Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d54b750111 |
Preserve Claude ACP quota classification and reset time (#13651)
Typed Claude ACP quota failures lost their recovery classification and reset time when the runtime reduced provider metadata to a generic category error. Inspect terminal metadata in memory and retain only safe recovery labels and a parsed reset timestamp. Preserve the existing handling of other limits. Verified real child processes on both pinned ACPX runtimes, adapter and server recovery regressions, all PR CI gates, and Greptile 5/5. Also isolate a pre-existing chat regression from unrelated fixtures’ retry work. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
924f07be8c |
feat(chat): simplify Slack onboarding and account linking (#13638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connections let people start and continue that work from Slack. > - Setup mixed app creation, credentials, URL verification, account linking, and testing on the same screens. > - People also needed a safe way to link their own Slack identity after the first operator finished setup. > - This pull request gives each step a clear place and keeps membership approval separate from identity linking. > - It also makes connection details easier to use and fixes misleading callback health behind HTTPS proxies. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: chat routes and services, shared contracts, and the Apps board UI. **Problem or motivation** Slack onboarding made users find settings without enough guidance. A second user needed operator help to link their account. Activity stopped at 100 records, and TLS termination could mark working callbacks as stale. **Proposed solution** Use six setup steps with editable app names, a generated manifest, credential guidance, URL verification, account linking, and an optional message test. Send each Slack user a private, expiring confirmation link. Require company membership or an approved access request before linking. Add cursor pagination and tolerate the internal HTTP hop in callback diagnostics. **Roadmap alignment** This improves the existing connected-app surface and supports CEO Chat without changing the task-and-comments model. The maintainer requested and reviewed the flow during a live Slack test drive. **Additional context** Related work: #7, #3349, #13000, and #13620. Those cover broader chat capabilities, older webhook paths, or plugins. This PR improves the existing native connector's setup and account-linking flow. HTTPS documentation was published separately in paperclipai/paperclip-docs#128. ## What Changed - Split Slack onboarding into six clickable sidebar steps. Keep secondary and primary actions on one row. - Generate the Slack creation link and read-only manifest from editable app, bot, and command names. Add credential prefix validation and direct instructions. - Add live account-link status and an optional mention-based message test. - Add private, single-use Slack account invitations and membership access requests. Retain cloud authentication/bootstrap checks and enforce the chat rollout flag in all identity APIs. Default new Slack connections to linked users only. - Put Settings, Access, Conversations, and Activity in the sidebar. Simplify conversation rows and remove active header badges. - Add 25-item activity pages, stable timestamp/ID cursors, and replay safety across pages. Preserve the legacy array API for clients without pagination parameters. - Fix false callback warnings when HTTPS terminates at a proxy. Keep host, port, and path drift detection. - Document the setup flow, pagination, callback diagnostics, and shared wizard footer rule. ## Verification - Passed: `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. - Passed: focused Slack callback and pagination integration tests; UI clipboard, wizard, pagination, and activity tests; OpenAPI route tests. The final access-gate fix also passes 27 focused tests covering cloud authentication/bootstrap, nonmember invitations, token validity, and the server-enforced rollout flag. - Passed: all 1,002 chat integration tests, 6,356 UI tests, and all 11 provider browser scenarios (including mobile light/dark navigation). After rebase, the identity route, sidebar, and 25 clipboard tests pass. - The full local `pnpm test:run` was attempted. The first run found 14 Slack fixtures that needed explicit guest access; those are fixed and the complete chat suite passes. Unrelated embedded PostgreSQL startup/resource failures and timeouts prevented a clean full local run. All CI checks pass on `2d858b036`, including the full chat, server, workspace, build, typecheck, and browser suites. - Live test drive: Slack app creation, credential setup, URL verification, private account confirmation, mention messages, and thread replies. Verified the callback warning clears for the existing proxied connection. - Review: create a Slack connection, follow the six steps, link a second user's account, and browse older activity with Next and Previous. ## Risks - Identity invitations carry a temporary capability. Tokens are hashed, expire after 15 minutes, work once, and require explicit confirmation by a company member. Access requests do not grant membership. - New Slack connections reject unlinked people by default. Existing connection settings remain intact. - Activity is a live ledger. Updated action rows can move forward in time. Older pages do not poll. - Proxy tolerance affects health display only. Slack signature checks and proxy authentication settings remain unchanged. - No database migration or package-lock changes. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, code execution, and browser verification. The runtime does not expose an exact model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted suites; full local-run limitations documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b8dc416da |
Fix default isolation for projects without workspace configuration (#13636)
Require a company-scoped configured workspace before applying the operator default for Git worktree isolation. Projects that only have a plain managed directory retain their existing behavior. Explicit isolation requests still require a valid checkout. Add policy and heartbeat integration regressions and document the default. The regression fails before the fix. An isolated checkout passes 541 relevant tests, and the server TypeScript check passes. Local repo-wide typecheck and build require the missing Rust toolchain; all CI lanes passed, and Greptile reviewed the refreshed head at 5/5 with no comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
de9414dd8b |
fix: return empty read instead of past-EOF range when log reader is caught up (#13592)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The server streams agent run logs to the UI. It reads the log in
byte ranges from a local file, or from an S3 mirror after a pod restart.
> - The range math in `readS3Range()` clamps the range end up to the
range start. A fully caught-up reader then asks S3 for the range
`bytes=total-total`.
> - S3 rejects a range that starts at the end of the object. It returns
a 416 `InvalidRange` error. The API turns this into a 500 error, and the
log poller repeats it.
> - This pull request removes the clamp. A caught-up or past-EOF reader
now gets an empty read, and the server does not send an invalid range to
S3.
> - The benefit is that log polling after a pod restart does not cause
repeated 500 errors.
## Linked Issues or Issue Description
No public GitHub issue exists for this bug. The description follows the
bug report template.
**What happened?**
The run-log API returns a 500 error when a client polls a run log that
lives on the S3 mirror and the client is fully caught up (`offset ===
total`). The cause is in `server/src/services/run-log-store.ts`. The
function `readS3Range()` computes `end = Math.max(start, Math.min(start
+ limitBytes - 1, total - 1))`. When `offset === total`, the
`Math.max(start, …)` clamp forces `end` up to `start`. The `start > end`
empty-read guard does not operate, and the code sends `Range:
bytes=total-total` to S3. S3 rejects a range that starts at or past the
end of the object with a 416 `InvalidRange` error. The error monitor
records this error many times, only in the staging environment, because
only the S3 fallback path is sensitive to it. The local-file path has
the same math, but Node file streams accept past-EOF reads. The function
`readFileRange()` in
`server/src/services/workspace-operation-log-store.ts` has the same
latent math.
**Expected behavior**
A caught-up reader gets an empty read: `{ content: "", nextOffset:
undefined }`. The server does not send an invalid range request to S3.
The poller sees no contract change.
**Steps to reproduce**
1. Start a run and let it write a run log.
2. Let the log upload to the S3 mirror, and remove the local file (this
occurs when the pod restarts).
3. Poll the run-log read endpoint until the client offset is equal to
the log size.
4. Poll one more time. The server sends `bytes=total-total` to S3, S3
returns 416 `InvalidRange`, and the API returns a 500 error.
**Relevant logs or output**
```
InvalidRange: Invalid range
at readS3Range (server/src/services/run-log-store.ts)
```
## What Changed
- `server/src/services/run-log-store.ts` — remove the up-clamp in
`readLocalRange()` and `readS3Range()`. A caught-up or past-EOF reader
gets an empty read.
- `server/src/services/workspace-operation-log-store.ts` — apply the
same fix to the shared math in `readFileRange()`.
- `server/src/services/run-log-store.test.ts` — the in-memory S3 mock
now rejects past-EOF ranges with `InvalidRange`, the same as real S3.
Add two regression tests for caught-up readers on the S3 path and on the
local path.
## Verification
- Run `pnpm vitest run server/src/services/run-log-store.test.ts`. All
17 tests pass.
- Revert only the source fix, and the new regression test fails with the
exact caught-up scenario. This shows the test covers the bug.
- Run the suites that use the workspace operation log store
(`workspace-runtime-control-recovery`,
`workspace-operations-reconciliation`). All 15 tests pass.
- Run `tsc --noEmit` on the server package. It reports no errors.
## Risks
- Low risk. The change only affects the empty and caught-up boundary of
range reads. Normal in-range reads give byte-identical results.
- Behavior change: a read with `limitBytes <= 0` now returns an empty
chunk instead of one byte. No caller passes a non-positive limit (the
default is 256000).
- Caught-up local reads keep the `nextOffset: undefined` semantics, so
pollers see no contract change.
## Model Used
- Claude Fable 5 (Anthropic, model ID `claude-fable-5`), with extended
thinking and tool use, run through the Claude Agent SDK.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Bender (Fable) <noreply@paperclip.ing>
|
||
|
|
84fe89906d |
fix: complete native agent review handoffs (#13581)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Native execution uses durable runs, issue locks, wake requests, and typed tool authority > - A child can finish with a native agent review request while its original assignee stays responsible for the work > - The reviewer then needs a bounded execution path that can inspect the child, record one decision, and finish safely > - Before this change, assignee-only gates rejected the reviewer or left the parent waiting after the child review ended > - This pull request adds typed reviewer admission, scoped reviewer tools, durable wake and recovery handling, and parent continuation evidence > - The benefit is that native review handoffs complete without changing child ownership or granting broad mutation access ## Linked Issues or Issue Description Refs: #13314 Refs: #13574 **What happened?** A native child run could report `needs_review` for an agent reviewer. The reviewer wake then failed assignee and execution-lock checks. The child remained in review and the parent remained waiting. **Expected behavior** The named reviewer should receive one durable wake. The reviewer should inspect the child and resolve the exact review card. The child assignee should stay unchanged. The parent should receive the recorded review outcome after the child reaches its terminal state. **Steps to reproduce** 1. Run a native task with a different named agent reviewer. 2. Keep the child assigned to its original worker. 3. Let the worker finish with a native completion review request. 4. Start the durable reviewer wake. 5. Resolve the review and finish the reviewer run. 6. Observe the child and parent state. **Paperclip version or commit** Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification source: `eea171aae`. **Deployment mode** Built from source. **Installation method** Built from source (pnpm build). **Agent adapter(s) involved** Not adapter-specific (core bug). **Access context** Both. **Database mode** Embedded PostgreSQL in the isolated live test fixtures. ## What Changed - Add server-validated native review assignment facts. - Admit only the exact company, issue, source run, decision, revision, addressee, and resolver policy. - Give reviewer runs a narrow set of Paperclip read and resolve tools. File and shell access follow the configured agent and environment policy, so reviewers can run tests. - Separate server-owned reviewer instructions from untrusted persisted review data. Escape the data boundary; retain server-enforced authorization. - Keep the child assignee unchanged. Atomically claim the reviewer run, wake request, and issue execution lock. A competing lock prevents provider startup. - Require the exact running reviewer session and current issue lock to resolve its assigned card. Reject missing, unrelated, or terminal reviewer runs. - Add durable reviewer wake, lock, stale-card, and abandoned-run recovery handling. - Prevent duplicate native wake dispatches during deferred admission and recovery. - Carry accepted or rejected child review outcomes into parent task context and continuation evidence. - Add focused server, runner, and native protocol coverage. - Preserve upstream continuation rules. Add child review decisions as separate evidence, while keeping real human answers in their own field. - Return actionable completion validation feedback to both providers. Permit a corrected completion after rejection. Keep strict terminal acknowledgment validation. - Apply exclusive shared-workspace locks to sandbox environments. Local and SSH folders can run concurrently, including when old settings request serialization. - Repair test timing, native event parsing, and the review artifact assertion. Allow a valid reject, correct, and accept review sequence. Check the accepted card against its reviewer run and decision. Keep polling within the existing deadline when review acceptance precedes the parent wake projection; report a specific missing-continuation error at timeout. - Apply the ACPX pending-call limit to reserved finish/block calls, with capacity-release and cancellation tests. ## Verification - `pnpm build`: passed on `eea171aae`. - `pnpm -r typecheck`: passed on `eea171aae`. - `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on `b31ad9ab8`; runner E2E typecheck also passed. - `pnpm check:token-gates`: passed. - Focused DB review, reviewer authority, and prompt-boundary checks: 31 tests passed. They cover invalid reviewer runs, competing locks, atomic admission, duplicate claims, and valid resolution. - Heartbeat, workspace, and recovery checks: 30 tests passed. - ACPX sidecar suite: 27 tests passed. Moving the capacity guard back below reserved handling makes both new regression cases fail. - Four focused live continuation checks passed on their first attempt at `f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and question tool guidance (12/12 each). These cases do not use the reviewer prompt path changed afterward. - Fresh Codex and Claude review-handoff checks passed all 29 native checks each on their first attempt at `eea171aae`. Both runs received the expected fixed prompt and completed cleanup. Only the six selected live flows were tested; no full paid provider catalog run. - The final commit only extracts the existing test-harness timeout diagnostic into a shared helper and adds positive and negative coverage. Removing the accepted-review guard makes two regression assertions fail; restoring it passes all six timeout tests. Production runtime code, prompts, deadlines, and grading criteria are unchanged by this final commit. - Deadline regressions: a valid continuation delayed 20 seconds succeeds within its 30-second unit-test deadline; an absent wake returns a specific candidate-failure diagnostic at that same deadline. Both assertions failed before the fix. Production E2E deadlines remain unchanged. - Historical native failures remain recorded: Docker availability failures; a valid reject/correct/accept sequence that the first-card grader misread; and a test that rejected the gap between accepted child review and parent wake projection. No failed result was regraded. The latest tests use a protected reference to the pinned Docker image and the unchanged artifact oracle and time limits. - Full repository verification runs in GitHub CI. Local verification uses the focused suites above, full build, and full typecheck. An unchanged Codex shutdown timing test failed once in CI, passed in isolation, and its full shard passed on the final commit without changes to that test or its causal code path. The original failure is retained in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no outstanding actionable findings. All review threads are resolved. All current-head CI gates passed, including the isolated native runner Docker build (55 successful checks; two skipped by the workflow). ## Risks - Reviewer admission depends on exact persisted decision and interaction bindings. A stale or changed card is rejected. - Paperclip control-plane tools are limited to inspection and review resolution. This is not a filesystem permission boundary; provider file and shell access retain the configured policy. - Deferred wake recovery changes dispatch receipt coalescing. A scheduler regression could delay a continuation if the receipt state is wrong. - Parent review outcomes are evidence for the model. They do not grant tool authority or change issue ownership. - This change does not address legacy lease-hold handoff behavior. > Roadmap review: native execution, review gates, and durable recovery are existing roadmap capabilities. This PR completes a narrow reliability path for those capabilities. ## Model Used OpenAI `gpt-6-astra` with reasoning, tool use, and code execution. OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and journal work. Context window size is not exposed by this session. Live test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the PR authors. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e926b13017 |
fix: keep sandbox termination progressing after bridge loss (#13287)
## Thinking Path > - Paperclip must stop remote execution after losing its controller. > - A bridge can remain blocked while the sandbox still incurs costs or performs actions. > - Waiting forever for that bridge prevents provider termination. > - A temporary provider outage can also exhaust cleanup attempts permanently. > - This pull request bounds bridge drain and persists cleanup retries with backoff. > - Cleanup ends only after provider confirmation, without a user accepting uncertain side effects. ## Linked Issues or Issue Description Builds on merged #13285. Related #13254 added exact provider termination receipts; merged #13272 adds explicit user retry. Merged #13352 stops active sandbox startup before waiting for setup. This PR preserves that immediate cancellation path and extends bounded teardown to ordinary release and destroy. Cleanup continues automatically after repeated provider failures. Refs #12953 for provider failures blocking execution. **What happened?** Daytona release waits for in-flight bridge activity before stop/delete. A dead bridge can prevent that wait from finishing. The host also stops cleanup after five failed attempts. **Expected behavior** Provider termination proceeds after a bounded bridge drain. Cleanup retries survive service restarts and provider outages. **Steps to reproduce** Start a sandbox command whose bridge promise never resolves, then release its lease. Separately, persist a pending-cleanup lease with five failed attempts and recover the provider. **Deployment mode** Hosted Paperclip with a Daytona provider; rebased onto master at `728f7185f` on September 14. ## What Changed - Bound bridge drain and provider lifecycle calls. Prefer stop for reusable sandboxes, with delete fallback. - Persist cleanup attempt identity, renewable in-flight deadline, and cooldown. Fence completion writes against superseded attempts. - Preserve scoped explicit Retry and its activity log. Explicit Retry can skip cooldown, but cannot take over a live cleanup attempt. - Continue cleanup after five failures with slower retries and an operator warning. - Exclude leases in cooldown before paging so they do not starve due work. - Add hung-bridge, restart, provider-recovery, and concurrent-cleanup regressions. ## Verification - Rebased onto master at `728f7185f`. The outstanding diff contains only cleanup changes; the merged controller-ownership prerequisite is excluded. - Daytona plugin suite: 160 passed, including immediate startup cancellation, graceful release, hung activity, and teardown regressions. - `pnpm exec vitest run server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts`: 31 passed. Two added integration cases verify explicit Retry during cooldown and while another cleanup owns the lease. They also verify run scoping and the activity log. - Targeted cleanup and cancellation cases in `environment-runtime.test.ts`: 20 passed. - Earlier live disposable Daytona test: provider stop ended background work, resume preserved files without restarting the old process, a new command succeeded, and the sandbox was deleted. This verifies provider behavior; it was not repeated for this rebase. - Latest-head CI and automated review are pending. Broad local tests, typecheck, and build were not rerun for this focused rebase; CI supplies those checks. ## Risks - Timing out bridge drain permits provider termination; it never supplies a stop receipt. - A crashed cleanup attempt remains protected for 15 minutes, then becomes eligible again. Repeated failures retry every 30 minutes after escalation. - The existing counter saturates at the escalation threshold; the new attempt identity and deadline prevent overlapping claims. - No schema, UI, telemetry, lockfile, or workflow change. ## Model Used OpenAI GPT-6 through Codex, using reasoning, repository inspection, code execution, and test tools. The precise backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fcdb3f2499 |
feat: add optional you.com search integration (#13555)
<!-- Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents that do research work need live information from the web > - Paperclip reaches external systems through governed, catalog-based MCP connections > - The Apps catalog is data-driven: a researched provider with a hosted remote MCP server becomes a connectable app with no runtime code change > - You.com operates a hosted remote MCP server for web search, content extraction, and research tools > - The server supports OAuth 2.1 with dynamic client registration, an API key in a bearer header, and a keyless free profile at a separate endpoint > - This pull request adds You.com to the self-serve MCP research ledger and generates its catalog entry with three connection methods: browser sign-in, API key, and the keyless free profile > - The benefit is that an operator can give agents live web search through the normal connection governance, and the free profile needs no account at all ## Linked Issues or Issue Description No public issue exists for this provider. The problem description follows the new-adapter issue template. **Agent or provider** You.com — web search and research tools over a hosted remote MCP server. **Why this adapter is useful** Agents that do research, monitoring, or fact-finding tasks need current web results. You.com exposes web search (`you-search`), live page extraction (`you-contents`), citation-backed research (`you-research`), and finance research (`you-finance`) as MCP tools. Any Paperclip company can connect it in a few clicks. The free profile offers `you-search` without an account, so a new company can try agent web search at zero cost and zero setup. **How the agent is invoked** Hosted remote MCP server (Streamable HTTP) at `https://api.you.com/mcp`. Three supported access paths, verified against the live server on 2026-09-16: - OAuth 2.1 browser sign-in. The server returns a `WWW-Authenticate` challenge with RFC 9728 protected-resource metadata and advertises a dynamic client registration endpoint, so Paperclip's automatic DCR path applies. - API key. Sent as an `Authorization: Bearer` header per the provider's official server manifest and docs. Keys come from you.com/platform and unlock higher rate limits plus the full tool set. - Keyless free profile at `https://api.you.com/mcp?profile=free`. Provides a reduced, read-only tool set. Official docs: https://you.com/docs/build-with-agents/mcp-server **Are you willing to implement it?** Yes. Implemented in this pull request. **Additional context** Research evidence collected 2026-09-16, from live protocol probes and official provider sources only: - Unauthenticated `POST https://api.you.com/mcp` returns HTTP 401 with `WWW-Authenticate: Bearer resource_metadata="https://api.you.com/mcp/.well-known/oauth-protected-resource" scope="Tools offline_access"`. - RFC 9728 metadata lists one authorization server with scopes `Tools` and `offline_access`. - The authorization-server metadata (RFC 8414) publishes authorization, token, and revocation endpoints, and advertises a `registration_endpoint`, so DCR is available. No registration was performed during research, per the runbook's non-registering preflight rule. - The keyless free profile answers `initialize` (server `You.com`, version `4.0.1`), lists the tools `you-search` and `you-discover`, and executed both tools successfully during the probe. - The API-key placement matches the provider's official `server.json` in the youdotcom-oss/mcp repository: header `Authorization`, value `Bearer <key>`. ## What Changed - Added You.com (slug `youcom`, wave 4, risk tier S2) to the self-serve MCP research ledger in `packages/shared/src/self-serve-mcp-research.json`, and refreshed the ledger verification date. - Added the You.com category (`ai`) and API-key header spec to `scripts/ingest-app-definitions.mjs`. - Added a You.com case to `specialMethodsFor` that emits three methods: browser sign-in (`mcp-oauth`, DCR), API key (`mcp-api-key`, bearer header), and keyless free profile (`mcp-free`, no auth). - Regenerated `packages/shared/src/app-definitions/youcom.json` and the generated registry via the ingestion script (`--definitions-only` mode; no unrelated provider churn). - Added the official You.com wordmark artwork (light and dark theme variants, taken from the provider's docs site) under `ui/public/brands/apps/`, with a manifest entry. - Updated `packages/shared/src/app-definitions.test.ts`: ledger counts (47 providers, 44 candidates), store count (48), verification date, and assertions for the three You.com methods and their endpoints. ## Verification - `node scripts/ingest-app-definitions.mjs --definitions-only` — passed. Generated the new definition and registry import only; no other provider JSON changed. - `node scripts/check-app-brand-assets.mjs` — passed (71 identities). - `node --test scripts/app-brand-validation.test.mjs` — passed. - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx` — passed (39 tests). - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts server/src/__tests__/tool-access-service.test.ts server/src/__tests__/generic-mcp-connection.test.ts server/src/__tests__/tool-connection-removal.test.ts ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx` — passed (181 tests). Two server suites that require embedded Postgres skipped on this machine by their own environment gate; the gate is unrelated to this change. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps` — passed; builds `@paperclipai/shared` with the new definition. - `pnpm test:run` (full Vitest suite) — 7,930 passed, 18 failed, 4,592 skipped. Every failure is environmental on this container: the embedded-Postgres suites refuse to start because the machine runs as root, the native runtime suites need the Rust runner binary that this container cannot build, and one media suite needs a native HEIC binary. No failure touches the app-catalog, connection, branding, or shared-package surface; those suites pass locally. CI is the authoritative gate for the full suite. - `pnpm --filter @paperclipai/server typecheck` — not completed: the script's `prepare:runner-vendor` prelude builds the Rust runner, which cannot build on this container. A direct `tsc --noEmit` reports only pre-existing errors from the missing vendored runner types; no error touches this change. No server code is changed. - Live You.com proof on 2026-09-16 (keyless free profile, real network calls): preflight 401 challenge with RFC 9728/8414 metadata and DCR endpoint ✓, `initialize` ✓, `tools/list` ✓, `you-search` call returned results ✓, `you-discover` call returned results ✓. - Live proof NOT run: an authenticated OAuth connect and an API-key call against the full server. This environment has no You.com account or API key. Per the runbook, this proof stays outstanding and must not be assumed from the keyless probe. Both paths match the reviewed `mcp_remote` patterns (DCR and bearer header) used by existing providers. - Browser e2e suites not run: opt-in per `AGENTS.md`, and this change adds catalog data only, with no UI code. ## Risks - Low risk. The change is catalog data plus generated output. It adds no runtime code and touches no existing provider. - The free-profile method is a fixed keyless endpoint. If You.com changes or removes `?profile=free`, that method breaks and the entry needs a ledger update. The OAuth and API-key methods do not depend on it. - The authenticated tool catalog is discovered live at connect time, so provider-side tool changes appear through the normal catalog refresh and quarantine flow, not through this definition. - Rollback is a single revert; no migration and no state are involved. ## Model Used - Provider: Zhipu AI, via OpenRouter - Model: GLM-5.3 (`z-ai/glm-5.3`) - Context window: 200K tokens - Capabilities used: tool use (shell, file edits, live HTTP probes), long-context repository reading - The change was produced with AI assistance and reviewed by a human before submission. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
e26d787928 |
Shorten continuation prompts and verify question tool guidance (#13574)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents must continue tasks using user answers without losing earlier requirements or approval gates. > - The wake prompt mixed human decisions with prior tool evidence and repeated detailed question instructions. > - Those instructions belong with the question tool, with a short routing hint in the wake. > - The Runner evals need to prove that answers, approvals, and completed work survive later turns. > - This PR shortens the prompts, separates authenticated answers, and adds continuation tests with useful screenshots. ## Linked Issues or Issue Description Refs #13517. This is a follow-up to the merged onboarding skill and Runner E2E work. Related #13539 covers responses received while a run is active; this PR preserves its cases and adds continuation coverage. Existing continuation/recovery and question PRs were searched; none covers this prompt/documentation and eval change. **What existing behavior does this improve?** The instructions sent when an agent continues a task, the native human-input tool documentation, and the evidence captured by Runner full-stack E2E. **Current behavior** The wake repeats a long question-tool guide. Human answers appear alongside untrusted prior results. Screenshot capture can finish at DOM load while the task still shows a spinner, even when backend behavior checks pass. **Proposed behavior** Keep earlier requirements unless the user changes them. Treat clarification as distinct from approval. Give authenticated human responses a scoped field. Keep tool and agent results as evidence. Put detailed question behavior in the tool descriptor and retain one routing sentence in the native wake. Wait for the correct task and loaded conversation before taking screenshots. **Reason and benefit** Reduce repeated prompt text and make authority boundaries clear. Test that real question cards, later answers, approval gates, and completed child tasks still work. Make screenshots useful for human review. ## What Changed - Shorten shared continuation instructions for legacy and native runners. Separate authenticated user responses from tool results and agent summaries. - Remove the detailed question guide from native wake prompts. Keep its behavior in the canonical `request_human_input` descriptor and existing payload schema. Regenerate semantic contracts and fixture hashes. - Add five continuation cases across four local profiles. Add a dedicated choice-then-text case for native Codex and native Claude. All 22 cells join the shared full E2E campaign. - Cover revised scope, clarification without approval, hostile instructions in a handoff file, and reuse of a completed child after restart. Keep production instructions and fixed user facts. - Capture continuation screenshots only when the intended task and conversation have rendered. Add provider-free browser regressions for loaders and wrong-task capture. - Preserve current master’s extra tool and onboarding cases. The default campaign now contains 166 cells; 35 manual everyday cells remain separate. ## Verification - `pnpm -r typecheck`: passed after replay on current master. - `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed. - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm test:e2e:runner:browser-support`: 4 passed. These tests failed against immediate screenshot capture and passed after the fix. - Focused continuation and native-input tests: 36 passed locally. The tool-authority suite could not initialize embedded PostgreSQL locally, including one isolated retry; its 17 assertions did not run locally. The full remote server shards passed on this PR commit. - `pnpm build`: passed after replay on current master. `pnpm test:run` was attempted locally but hit the same embedded PostgreSQL initialization failure; the remaining local run was stopped after complete remote CI passed. This is not claimed as a full local test pass. - [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685): passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All server/chat/workspace/serialized shards, browser shards, Runner checks, typecheck, build, canary and policy checks passed. The isolated native Runner build and security checks also passed: 57 successful checks, with two expected Storybook skips. - Greptile reviewed the exact PR head at 5/5, with no findings or unresolved review threads. The PR has no merge conflicts. - [Live question-docs report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/): 3/3 passed at source `83dd132f2` before replay on master. Native Codex and Claude each asked a choice, waited, asked a text question, and saved both answers. Claude also passed a completed-child restart case. All three native turns are checked for absence of the old question block. - [Earlier continuation report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/): all five continuation cases passed on native Claude. The report retains campaign and revision provenance and separately shows two unresolved onboarding behavior failures. - [Before/after prompt report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/): full text, current recorded Claude inputs, and reproducible reference-token counts. The controlled wake comparison removes 401 reference tokens; the net counted input reduction is 339 after charging the larger tool description. These are text-size estimates, not measured billing savings. ## Risks - Prompt wording affects model behavior. Live results cover the stated cases, not every provider or conversation. Legacy profiles are registered but were not rerun for this change. - The optional continuation field changes prompt data only; there is no database migration or new production API. - Authenticated answer projection excludes generated summaries and agent-resolved interactions. It preserves the answer’s question or approval scope. - The screenshot guard can expose UI loading failures that earlier runs hid. Backend grading alone no longer makes those captures valid. - The two prior onboarding failures remain separate product issues: work before acceptance and a missing saved plan. This PR does not claim the entire onboarding suite passes. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, code execution, and browser verification. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted tests above; the full local database-startup limit is documented - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5d9b20ccf0 |
fix(ui): keep task composer available while pause state loads (#13562)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task messages enter through the shared task composer. > - The task page waits for a separate tree-control query to find active pause holds. > - The old loading guard disables the composer until that request settles. A slow or stalled request blocks messages on both desktop and mobile. > - This PR allows sends while that query is pending. The server still rejects paused board messages before saving a comment or waking an agent. > - Query failures and known pause holds still block the composer. > - Regression tests cover pending sends, successful responses, query failures, and late root or inherited pauses. ## Linked Issues or Issue Description Fixes: #13561 Related: #13569 adds pending/error coverage for the same fix. This PR now covers those cases and the late-pause transitions. The report overstates two details: `isPending` clears after a successful response, and the shared loading guard also affects mobile. The reproduction holds the request pending. It does not prove why the reporter's request remained unresolved. ## What Changed - Preserve the original two-line fix that removes the pending-state composer block. - Add eight page regression cases using a real QueryClient and a controlled API promise. Cover desktop and mobile, submission before the query completes, successful resolution, rejected resolution, and late root/inherited pause responses. - Repair an existing server CI failure in a separate commit. Reuse the shared unique-violation helper so Drizzle-wrapped duplicate inserts become retryable document conflicts. Add a deterministic regression and preserve unrelated database errors. - Repair an existing Inbox test race in a separate commit. Wait for workspace metadata, which resolves independently of the task list. ## Verification - Red: restore the pre-fix `IssueDetail.tsx` and run the new `composer tree control` cases. Both desktop and mobile pending-send cases fail with `Checking task status…`. The other six cases pass. - Green: restore the original PR fix. All 343 tests in IssueDetail, TaskChatThread, and TaskChatComposer pass. - The server comment/reopen and artifact-review suites pass (179 tests). They include POST and PATCH pause checks that return 409 before any comment, task mutation, or wakeup. - The document error regression fails before the shared-helper fix and passes after it. The document, artifact-review, and database-error suites pass (31 tests). The handler matches only the issue-document key constraint; revision and other constraint errors retain their original identity. - The Inbox suite passes (27 tests). - `pnpm check:token-gates` passes. - `pnpm -r typecheck` and `pnpm build` pass locally. The full CI matrix passes on `5af1ed0b43269247aaba406cf4fd4d7fe1a22e75`: 54 successful checks and two opt-in Storybook skips, including all general/serialized tests, runner checks, release checks, and eight browser shards. - One server shard initially hit an unrelated `EADDRINUSE` on test port 52000. Its single rerun passed without code changes. - Greptile completed successfully on that exact commit with 5/5 and no outstanding findings or review threads. - The serial local `pnpm test:run` was started, then stopped after the complete parallel CI matrix passed. It is not claimed as a completed local full-suite run. The focused local suites above did finish successfully. For a manual reproduction, delay the task's `/tree-control-state` response, open the task, and enter a message. Send should remain available during the delay. Resolve the response with an active pause hold and confirm that the pause takeover replaces the composer. Reject the request and confirm that the error blocks sends. ## Risks A user can attempt a send before the pause response arrives. The server remains authoritative and returns 409 for a paused task before saving or waking work. The known pause takeover and query-error block remain. There are no schema, API, or styling changes. The document change restores the existing conflict/retry behavior for wrapped database errors. It does not retry unrelated database failures. The Inbox change affects test synchronization only. ## Model Used - Original fix: Anthropic Claude Sonnet 4.6 (`claude-sonnet-4-6`), 200k context, tool use and code editing, as reported by the author. - Review, regression tests, and CI repairs: OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, tool use, and code execution. The session does not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: austinpilz <austinpilz@users.noreply.github.com> Co-authored-by: Dotta <bippadotta@protonmail.com> |
||
|
|
165b10bd98 |
fix: enable GitHub Actions MCP toolset (#13553)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The GitHub connector lets agents use repository tools through MCP. > - GitHub excludes Actions from its default MCP toolsets. > - Approval of Actions permissions therefore does not make workflow tools appear in Paperclip. > - This pull request adds Actions to the requested toolsets for discovery and execution. > - Users can refresh existing connections and use workflow tools under the existing access rules. ## Linked Issues or Issue Description **What happened?** GitHub Actions tools remain absent after the GitHub App receives Actions read/write access and the user refreshes actions in Paperclip. Paperclip does not request the Actions MCP toolset. **Expected behavior** Authorized GitHub connections expose workflow tools, including `actions_run_trigger` with `method: "run_workflow"`, so agents can dispatch an existing release workflow. **Steps to reproduce** 1. Connect GitHub to Paperclip with access to a repository that has a dispatchable workflow. 2. Grant the GitHub App Actions read/write permission and approve the installation update. 3. Refresh the connection's actions in Paperclip. 4. Observe that the workflow tools are absent. **Paperclip version or commit** Base commit: `fae698031`. **Deployment mode** Hosted instance with a managed GitHub connection. The same missing header affects PAT connections. No matching public issue or pull request was found in the duplicate search. ## What Changed - Send `X-MCP-Toolsets: default,actions` through the shared GitHub MCP header helper. This covers discovery, refresh, and execution for existing and new managed or PAT connections, including legacy rows identified through `transportConfig`. - Test catalog refresh, tool risk classification, and workflow dispatch through a mock MCP server. - Document tool names, workflow arguments, required GitHub permissions, and the refresh step. ## Verification - Passed both affected test suites: `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts server/src/__tests__/tool-gateway.test.ts` (391 tests). - Passed `pnpm check:token-gates` and `git diff --check`. - Live provider check: the default catalog returned 45 tools. `default,actions` returned 49 tools, with no tools removed. The four added tools were `actions_get`, `actions_list`, `actions_run_trigger`, and `get_job_logs`. - Live `actions_get` / `get_workflow` call succeeded. No workflow was dispatched during live verification. - Passed `pnpm -r typecheck` and `pnpm build` with the existing Rust toolchain added to PATH. - Rechecked server typecheck and build after the legacy-connection fix; both passed. - The full local test run has reported three skills-cache failures in `company-skills-service.test.ts`. All three reproduce on the untouched base commit (`fae698031`) on this macOS host: runtime-cache directory renames fail with `EACCES`. The full run remains in progress. - Greptile: 5/5 on `7d391e3c7`, with no unresolved review threads. - After deployment, use **Refresh actions** on an existing GitHub connection and verify the workflow tools appear. ## Risks - Refreshed GitHub catalogs expose more tools. Existing access, approval, and quarantine rules still apply. `actions_run_trigger` keeps GitHub's destructive classification because it also supports cancellation and log deletion. - GitHub still enforces token and installation permissions. Dispatch requires Actions write permission and a workflow with `workflow_dispatch`. - No database migration or saved connection edit is required. ## Model Used OpenAI GPT-6 through Codex, with code execution and tool use. The exact serving model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fae6980310 |
revert(apps): restore Google connector visibility (#13552)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - The Connectors catalog lists services that agents can use. > - PR #13551 temporarily hid Google connectors. > - We now want to restore their catalog visibility. > - This PR reverts that change and restores the previous catalog behavior. ## Linked Issues or Issue Description Refs: #13551 Revert the temporary removal of Google connectors from the UI. ## What Changed - Restore Gmail and eight Google Workspace entries to the catalog. - Restore the matching branding flags and original catalog and service tests. - Remove the temporary-hiding documentation note. This is an exact revert of commit `cf1e873ab24277d55ffd3ab06074f77014dc4015`. ## Verification - Passed: 507 catalog, UI, and connection service tests. - Passed: `pnpm check:token-gates` and `node scripts/check-app-brand-assets.mjs`. - Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter @paperclipai/ui... typecheck`. - Full local build and typecheck stop at the Rust runner because `cargo` is not installed. - Full local Vitest was not repeated because the unchanged base has confirmed macOS skill-cache permission failures. The full CI suites passed. - Passed: all GitHub CI gates; Greptile 5/5 on commit `4e3dddef0ebfef1f99001e7735822ed4cba852ab`, with no review threads. - Reviewer check: open Connectors and confirm that Gmail and Google Workspace entries appear again. ## Risks Low risk. This restores the previous catalog visibility and setup entry points. Connector implementations and saved connection data are retained. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact deployment ID and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cf1e873ab2 |
fix(apps): temporarily hide Google connectors (#13551)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - The Connectors catalog lists services that agents can use. > - We need to temporarily remove Google connectors from the UI. > - The catalog already separates visibility from retained definitions. > - This PR uses that setting so Google can return with a small change. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Connectors catalog and its setup entry points. **Current behavior** The catalog shows Gmail and eight Google Workspace connectors. **Proposed behavior** Temporarily hide those nine entries. Keep their definitions and existing connections. **Reason and benefit** Make the temporary UI removal easy to reverse. **Breaking changes** Fresh catalog setup no longer offers Google. Saved connections keep the existing management and reconnect paths. ## What Changed - Add the nine Google connector slugs to the existing hidden list. - Match the branding manifest visibility flags. - Update existing catalog and service tests. Keep backend Google connection coverage and document how to restore visibility. ## Verification - Passed: 507 targeted tests covering catalog definitions, URL matching, setup routing, connector UI, branding, and the connection service. - Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter @paperclipai/ui... typecheck`. - Passed: `pnpm check:token-gates` and `node scripts/check-app-brand-assets.mjs`. - Full local build and typecheck stop at the Rust runner because `cargo` is not installed. - Stopped the full local Vitest run after skill-cache permission failures. Three failures in `company-skills-service.test.ts` also reproduce on the unchanged base branch. The final connector service suite passes all 319 tests. - Greptile: 5/5 on the current commit, with no open review threads. CI is retrying one unrelated preview-server readiness timeout. That test file passes all seven tests locally. - Reviewer check: open Connectors in a company with no Google connections. Gmail and Google Workspace entries should be absent. Existing saved connections remain manageable. ## Risks Low risk. This uses the existing catalog visibility mechanism. No connector implementation, credential, or database schema is removed. Restoring visibility requires updating both the hidden list and branding manifest. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact deployment ID and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6fe8e30625 |
feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps gives agents governed access to external resources. > - Operators need to inspect Railway services, read logs, deploy code, and run container commands. > - Railway offers hosted MCP with OAuth, but broad remote actions hide their internal operations. > - This PR adds a branded connection and fixed direct operations through the existing gateway. > - Separate SSH keys enable container commands under the same grants and policies. > - Operators can require approval for an action and inspect the resulting audit record. ## Linked Issues or Issue Description **Subsystem affected** Apps catalog, connection setup, gateway execution, and connection documentation. **Problem or motivation** Agents need Railway access through Paperclip. Operators need to grant and revoke that access, inspect available actions, and govern deployment and container operations without giving agents provider credentials. **Proposed solution** Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and the gateway. Probe the actual credential before enabling fixed GraphQL operations. Use a dedicated grant-owned SSH key for bounded container commands. **Alternatives considered** A catalog entry alone cannot execute the missing operations. The hosted general agent has opaque internal effects. An unrestricted CLI runtime can bypass action policy and inherit ambient credentials. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps path and the Connected Apps direction in ROADMAP.md. It does not add a plugin or parallel connection service. Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway. They do not add this outbound Apps connection. The separate shared agent-picker fix is #13414 and is not included here. ## What Changed - Add the generated Railway catalog entry, official marks, provenance, and OAuth setup guidance. - Add fixed service/deployment status, bounded logs, and redeploy/restart/rollback tools. Block source deployment until the provider can atomically bind the approved repository and commit. - Verify API access with an explicit workspace before exposing direct tools. - Add grant-owned SSH key setup and a bounded runner with host verification, target checks, isolated state, and cleanup. - Block the opaque hosted railway-agent and accept-deploy actions. Preserve normal Allowed defaults and Ask-first policies for other actions. - Quarantine new or changed Railway schemas after initial discovery, including reconnect. - Add provider, lifecycle, gateway, SSH, UI, and browser fixtures. Document setup, limitations, and the release checklist. ## Verification - Security follow-up: removed the unsafe source-deployment mutation. Direct calls and old active catalog entries are denied before any upstream request, including normalized aliases. Refresh marks retired entries disabled. All 386 focused Railway, catalog and gateway tests passed, and server TypeScript checking passed. Full [GitHub CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144) passed on |
||
|
|
d0b67bfe71 |
feat: queue approvals and answers during active runs (#13539)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users guide running agents through messages, questions, and approval cards. > - Messages already wait in a queue when an agent is running. > - Card responses did not appear in that queue. Some question answers also steered a later run without a user click. > - A fast approval could invalidate the agent's review handoff and cause it to stop its own run. > - This pull request gives card responses the same queue controls and preserves the exact response during delivery. > - Users can wait for completion or explicitly send the response with Interrupt or Steer. ## Linked Issues or Issue Description Refs #13517, which is merged. This PR targets master and adds queued interaction responses on top of the onboarding changes. Related continuation work: #10519 and #12866. **What happened?** Accepting a proposal while its source run was active left a saved response outside the message queue. The agent could then lose its review path, reassign the task, and cancel itself. Answers to older questions could also steer another active turn without a click. **Expected behavior** Save the response immediately. Queue its continuation behind the active run. Deliver it after completion, or when the user explicitly chooses Interrupt or Steer. Preserve approval revisions and answer choices. **Steps to reproduce** 1. Let an agent publish a confirmation card while its run is still active. 2. Accept the card before the agent finishes its review handoff. 3. Inspect the message queue and the task's next run. **Paperclip version or commit** Reproduced on da8a3876c with the onboarding changes from #13517. **Deployment mode** Local development from source. The fix covers legacy adapters and native Runner turns. ## What Changed - Project resolved cards into the existing queue as immutable responses. Keep answers and exact approval revisions. - Require an explicit click to steer a response into a compatible native turn. Use Interrupt when a fresh session is required. - Preserve typed response context through interruption, cleanup waits, and normal queue promotion. Keep the direct answer channel for a provider blocked on its original question request. - Accept the source run's review handoff after its card resolves. Reject stale agent reassignment that would orphan a queued response. - Add deterministic regression tests and an `accept-while-running` case to the first-task suite. Require recorded timestamp overlap before that case can pass. - Keep the first-task skill name out of user-facing messages. ## Verification - Red-green: the original route failed the queue regression; the changed route passes it. - Focused server/UI tests: 139 passed, including 64 queue-route tests. - Runner harness unit tests: 314 passed. - Server, UI, and Runner E2E typechecks passed. UI token gates passed. - Full repository typecheck and build passed. Server typecheck passed again after review fixes. - Review regressions: 165 queue/reopen route tests, 53 wake admission tests, and 18 run identity tests passed. Approval acknowledgement recovery and both message/approval arrival orders are covered. - Full local test run: 12,401 passed; three new admission regressions ran against a cached pre-fix module. A fresh run of that entire suite passed (53 tests). The complete CI suite passed on the final commit. - Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional Storybook checks skipped. Every server/workspace/browser shard, Runner verification, build, typecheck/release registry, canary, policy, and security check passed. Greptile: 5/5, no unresolved threads. Earlier interrupted CI workers were replaced by this fresh complete run. - After integrating the updated parent: 314 harness tests, 119 queue/admission tests, 44 onboarding/question-delivery tests, and 13 native recovery tests passed locally. Full repository typecheck and build passed. - Clarified the skill wording preference: routine replies describe the action without announcing the internal skill; direct questions and permission/security/execution disclosures remain truthful. - The paid `accept-while-running` scenario is registered for all four local first-task profiles. It has not been run against a model in this change. - Rebased onto the merged parent at `11921075a`; the resulting tree exactly matches the locally verified integration tree. Final-head CI on `b53054807` passed: 54 successful checks, 2 optional Storybook checks skipped, no failed checks. Every new server/browser shard, aggregate verify/e2e gate, Runner, typecheck, build, canary, and security check passed on the first attempt. Greptile reviewed this exact head at 5/5 with no unresolved threads. ## Risks - Responses now wait instead of implicitly steering another active turn. A provider blocked on the original question still receives its answer directly. - Approval receipts cannot be edited, discarded, or reordered as comments. This preserves the recorded decision. - Interruption must still prove that the prior execution stopped. The tests cover cleanup waits and duplicate delivery. - The new paid overlap case can be unexercised if the model finishes before the click lands. It cannot pass without evidence of overlap. - No database migration is required. This repairs the existing approvals and execution controls; it does not implement the roadmap's work-stream queues. ## Model Used OpenAI GPT-6 through Codex. The exact deployed model ID and context-window size were not exposed in this session. Capabilities used: agentic reasoning, repository inspection, code editing, terminal commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
11921075a4 |
Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first task helps a new user define and approve useful work. > - That workflow needs reusable instructions and tests against the production experience. > - Native Codex and Claude must load the assigned skill, including after resume. > - Maintainers need recorded conversations and precise failed checks to judge regressions. > - This pull request adds the first-task skill and a suite in the shared Runner E2E harness. > - It keeps behavior results separate from informational quality scores and incomplete recordings. ## Linked Issues or Issue Description **What existing behavior does this improve?** The first onboarding task and the Runner E2E report used to review it. **Current behavior** Onboarding embeds its policy in a hidden brief. Native Codex drops the skill-instructions setting at the Rust boundary. The shared E2E harness has no onboarding suite or full conversation view. **Proposed behavior** Assign and invoke `/first-task` for the onboarding task. Send selected Codex skills as structured protocol inputs. Run twelve scenarios across legacy Codex, legacy Claude, native Codex, and native ACPX Claude. Include all 48 cells in full campaigns. Show recorded chat, question and approval cards, exact checks, instructions, and billing in the shared dashboard. **Reason and benefit** Measure the real onboarding experience before changing prompts. Distinguish infrastructure failures, behavior failures, and unexercised journey steps. **Breaking changes** No database migration or production API change. First-task instructions now live in an assigned skill. The user-edited persona is preserved; the skill includes the maintainer-approved proposal-mode mapping and saved-plan requirement. Related: #11043 is earlier onboarding work. #13422 already fixes native Claude model pinning, context delivery, and read permissions on master; this branch includes those fixes through its base. The new Claude recovery test supplements them. ## What Changed - Extract and assign the first-task skill while retaining the production greeting and opening question. - Carry the Codex skill-instructions flag through thread start and resume. Resolve explicit task skill references only against assigned skills and send native skill inputs. - Invoke an unambiguously selected assigned skill through Claude ACPX’s native slash-command parser on initial and resumed turns, retaining the entire task/wake envelope as its argument. Do not carry that invocation into ordinary tasks. - Restore the saved single-task proposal modes: confirmation card, or saved plan with revision-targeted checkbox approval. Explicit plan requests also require a saved plan. - Add first-response and complete-journey cases with fixed user facts, acceptance checkpoints, durable outcome checks, and accounting for child runs. - Fail the eval when choice questions have fewer than two real options. Recognize planning documents without treating them as completed work. - Add optional, bounded quality judging as explicit post-processing. - Render full conversations and static interaction cards in the shared report. Conversations start folded. Show original and regraded results and incomplete journeys distinctly. - Keep credential-persistence scanning outside the first-task behavioral suite; retain public evidence redaction. - Refresh generated capability references after the API-reference edits. - Correct shared native question guidance and tool schemas: choices need at least two meaningful options; open-ended questions use canonical text fields with the required compatibility payload. Verify both formats through real tool-authority persistence. - Disable announcements automatically for every isolated Runner E2E process and label the gallery environment/provider/target explicitly. - Remove CI races in the GitHub connection browser test and native session recovery test by waiting for the actual async work before asserting its results. ## Verification - `pnpm exec vitest run server/src/services/onboarding-first-task-assets.test.ts server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19 passed. - `pnpm --dir packages/paperclip-runner exec vitest run src/drivers/acpx/runtime-host.test.ts src/drivers/acpx/native-skill-prompt.test.ts src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command forwarding and the 1 MiB input boundary both failed before their fixes and passed afterward. Coverage includes changed skills on reopen, approval context, and an ordinary subsequent task. - Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64 first-task fixture and grader tests also pass. - Full repository typecheck and build passed locally. Server typecheck and Runner build passed again after the native-command change. - Full GitHub Actions CI passed on `23e56447b`: all server/workspace/browser shards, Runner verification, typecheck/release registry, build, canary, policy, and Docker checks. Greptile reviewed this exact head at 5/5 with no unresolved threads. The earlier broad local run had database startup/timing failures that passed isolated retries; the complete remote suite is green. - Merge verification against current master: 312 harness tests and 13 native recovery tests passed. Regenerated semantic contracts and fixture hashes pass their consistency check. Full local typecheck and build also passed on the stacked queue branch. After merging the latest master and preserving the GitHub setup timing regression in the split browser suite, both focused GitHub browser tests passed. Three CI timing/startup flakes passed local verification and one remote retry; all latest-head checks are green. - Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local mock API confirmed that `/skill-name` expands the assigned skill body before the model request and retains the task arguments. A prose mention does not. The probes made no paid model calls. The ACP probe used the current first-task skill body and retained the wake arguments. - [Full 48-case campaign and report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/): 44 passed after three interrupted Codex cases completed in targeted reruns. Original results, regrades, and all 51 executions remain in the report provenance. - [Claude campaign after the shared-question fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/): 10/12 passed with zero single-option failures. All 12 recorded the current assigned skill and corrected guidance. The failures exposed skipped skill invocation and a missing saved plan. This PR adds native command invocation and explicit saved-plan instructions; the subsequent report below still shows behavior failures. - [Fresh 12-case Claude report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/) at `78452129e`: 10/12 pass after correcting two false proposal-matcher failures. The recordings said “Here is the task I will create and run/complete” in approval cards; the old matcher missed that word order. Regression tests failed before the fix and pass after it. Original results and offline regrade provenance remain linked. No agent rerun was needed. Zero single-option-question failures; two behavior failures remain: direct work before acceptance on a plain first message, and an explicit plan request without a saved plan. Neither check was relaxed. The follow-up `82087ac7e` fixes command-prefix size accounting; `94aefb1f3` fixes only that proposal matcher. - Report browser checks confirm folded conversations, rendered cards, explicit Local/Daytona labels, and no page errors. The published-object audit scanned 1,306 text files across 2,154 objects with no credential-format findings or prohibited files. Image pixels and unknown token formats are outside that scan. ## Risks - Model behavior is nondeterministic. One campaign is evidence, not a guarantee. The two remaining Claude behavior failures are visible in the report and require further product work; this PR does not claim all onboarding scenarios pass. - The suite checks persisted Paperclip effects. It cannot prove the absence of arbitrary external effects. - Historical recordings can miss later journey steps. These remain incomplete, never passes. - Native profiles switch runtime after the production onboarding wizard because it does not yet expose a native option. - Quality scores are informational and cannot override behavioral failures. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, and code execution. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dcb04a8062 |
fix(claude-local): read a macOS isolated login from its suffixed Keychain item (#13519)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Connecting a Claude subscription during onboarding uses an isolated
login: the wizard points `claude` at a per-connection
`CLAUDE_CONFIG_DIR` and then verifies the credential before saving the
connection
> - The verifier reads `.credentials.json` from that directory — but on
macOS, Claude Code does not write a credentials file at all: it stores
the OAuth credential for a custom config dir in a per-directory Keychain
item named `Claude Code-credentials-<first 8 hex chars of sha256(dir)>`
> - So on macOS the connect step can never verify a successful sign-in,
and onboarding dead-ends at "Could not verify the local subscription"
(Linux works because the CLI falls back to writing the file there, which
is why the Docker-based smokes pass)
> - This pull request teaches the credential readers to consult the
login home's own suffixed Keychain item when the file is missing
> - The benefit is that macOS self-hosted users can connect a Claude
subscription during onboarding, while the standing isolation invariant —
an isolated login must never fall through to the machine-level operator
login — is preserved, because only the per-directory suffixed item is
ever read
## Linked Issues or Issue Description
No existing issue. Description follows the bug-report template:
**What happened?**
On macOS, connecting a Claude subscription during onboarding (or from
Connections) always fails with "Could not verify the local subscription.
Run the sign-in command shown for this connection, finish signing in,
then try Connect again" — even after `claude auth login` completes
successfully in the isolated `CLAUDE_CONFIG_DIR`.
**Expected behavior**
After finishing the browser sign-in for the printed command, clicking
Connect verifies the subscription and saves the connection.
**Steps to reproduce**
1. On macOS, run onboarding on a fresh instance and reach "Connect a
model" → Claude → Subscription.
2. Run the printed `export CLAUDE_CONFIG_DIR=… && claude auth login`
command in a terminal on the same machine and complete the browser
sign-in.
3. Return and click Connect. Verification fails every time. Inspecting
the isolated directory shows `.claude.json` with a fully populated
`oauthAccount` but no `.credentials.json`; `security
find-generic-password -s "Claude Code-credentials-<suffix>"` shows the
credential landed in the Keychain, where the verifier never looks.
**Paperclip version or commit**
Reproduced on `2026.915.0-canary.11` (`dffc2b3ca`) with Claude Code
2.1.231.
**Deployment mode**
Self-hosted, authenticated instance on macOS.
**Installation method**
`npx paperclipai onboard` (also affects any macOS install; Linux is
unaffected).
## What Changed
- `packages/adapters/claude-local/src/server/quota.ts`:
- New exported helper `readIsolatedClaudeKeychainToken(loginHome)` —
computes the suffixed service name (`Claude Code-credentials-` + first 8
hex chars of `sha256(loginHome)`) and reads only that item via
`/usr/bin/security`; returns null off macOS
- `readClaudeToken` with a custom `CLAUDE_CONFIG_DIR` now consults that
directory's suffixed item after the file reads miss (previously it
refused the Keychain entirely for custom homes). The unsuffixed operator
item is still gated behind the explicit `allowKeychain` opt-in with no
custom home, unchanged
- `server/src/services/local-ai-credentials.ts`: for anthropic isolated
logins, fall back to the suffixed Keychain item after the hardened
credentials-file reads miss. The file path is untouched and still
preferred; the hardened file reader (`readLocalAiCredentialFile` with
its uid/mode/symlink checks) is not bypassed
- Tests: adapter keychain suite extended (suffixed lookup for custom
homes, no unsuffixed fallback when the suffixed item is absent,
off-macOS null); server verifier suite extended (keychain fallback when
the file is missing, file preferred over keychain, absent-login failure
still never touches the ambient reader)
Security note: the suffix binds each Keychain item to exactly one auth
home, so reading it can only surface the login performed inside that
home. The account-isolation invariant the old code enforced by refusing
the Keychain outright ("never substitute the server operator's login for
a user's isolated login") is preserved — the unsuffixed item is never
consulted for an isolated login, and a new test pins that.
The suffix derivation was confirmed against a live login on macOS: a
real `claude auth login` into an isolated home left no credentials file,
wrote the full `oauthAccount` to `.claude.json`, and created a Keychain
item whose suffix equals the first 8 sha256 hex chars of the exact
`CLAUDE_CONFIG_DIR` string; reading it back with the same `security`
invocation returned the live token, which the new code path then
verifies via the existing quota probe.
## Verification
- `pnpm exec vitest run src/server/quota-keychain.test.ts`
(claude-local): 10 tests pass; full claude-local suite: 287 passed, 1
skipped
- `pnpm exec vitest run src/__tests__/local-ai-credentials.test.ts`
(server): 11 tests pass
- Reverting only the verifier change makes the two new server tests fail
— the suite reproduces the live bug
- End-to-end on macOS: a dev server built from this branch, fresh data
dir, full onboarding walk with a real `claude auth login` into the
printed isolated dir — the connect step verifies and saves the
connection
## Risks
- Low. The change is additive and fail-closed: when the suffixed item is
absent (Linux, older Claude Code versions, no login performed), behavior
is byte-identical to today — the file reads run first and the failure
message is unchanged
- The `security` call runs with the existing 10s timeout and swallowed
errors, matching the established unsuffixed-item code path
- No migrations, no API surface changes
## Model Used
Claude Fable 5 (`claude-fable-5`), extended thinking with tool use
(Claude Code).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
9fd2e50310 |
feat: create company skills from runner tasks (#13538)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner gives agents tools to change company resources. > - Users need agents to save reusable skills during a task. > - A saved skill needs a visible result that users can inspect and edit. > - This pull request adds `create_skill` and a task feed card linked to Skill Studio. > - Users can open the saved skill from the task and edit the same resource. ## Linked Issues or Issue Description **Subsystem affected** Runner tools, company skill storage, task feed, and Skill Studio. **Problem or motivation** The Runner has no dedicated tool to create a company skill. A user cannot follow a creation result from the task feed to the saved skill. **Proposed solution** Add a company-scoped `create_skill` tool. Save the skill with the existing company policy. Add one creation card to the task. Open a named sidebar tab from that card. Let the user open the same skill in Skill Studio. **Alternatives considered** An agent can write a local file, but that file is not a company skill. A second document copy in the task would become stale after a Studio edit. The sidebar therefore reads the saved skill directly. **Roadmap alignment** This extends the shipped Skills Manager, Skill Studio, and Skills Store milestone. The maintainer requested and approved this scope. Search found no duplicate `create_skill` PR or issue. Related UI validation work: #8715. This PR does not change that validation display. ## What Changed - Add the real Runner tool, its contract, and its mock implementation. - Validate the complete SKILL.md and derive company, task, agent, and run identity from authentication. - Apply the existing company skill policy. Do not assign the skill to an agent. - Make keyed retries return one skill and one creation event. Reject conflicting retries. - Make concurrent file creation safe. Never replace an existing published skill during creation. - Add a creation card, a named sidebar tab, and an Open in Skill Studio action. - Show saved Studio edits when the user returns to the task. - Add storage, policy, mode, retry, UI, and Product E2E tests. Document the tool. - Fix deleted-name reuse, onboarding panel persistence, immediate feed refresh, and mock validation parity from review. - Serialize Studio file edits and renames with skill deletion and recreation. Reject stale editor requests before they can change a replacement skill. - Generate the standalone mock parser and validator from the production contract. Use portable UUIDs so the browser scenario bundle builds. ## Verification - All latest-head PR checks pass on `145dd76a5`, including all server shards, browser E2E, Runner verification, build, typecheck, and release dry run. Greptile: 5/5 with no open findings. An interrupted CI runner was retried successfully. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Review regressions: 73 storage tests, 6 real API tests, 63 UI tests, and 61 semantic runtime tests passed. Parser synchronization passed. - CI exposed existing fire-and-forget Sentry test races. Reproduced the resumption race locally, then synchronized the related sweep and finalizer assertions on the actual report; all 27 tests across the three affected files pass. - Runner scenario browser build and strict content-security-policy check: passed. - Runner suite: 2,012 tests passed; 10 skipped. - `pnpm test:run`: the general-server batch had 12,416 passes and two failures. The old tool-count assertion was fixed; all 16 authority tests then passed. The chat webhook test had a socket error; it passed four isolated reruns. - Both workspace test groups passed. The isolated route suites completed. Two socket failures in the initial route batches passed on individual reruns; all remaining 61 files passed. - Product E2E `create-skill-studio`: passed with local Codex and local ACPX Claude. - Manual browser test: submit a task, observe the real tool call and creation card, open the sidebar, edit in Studio, save, and return. The task reached Done. The saved second revision and sidebar tab survived a server restart. - The new companion headless Runner Eval passed. Companion coverage PR: https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not run because no immutable runner image was configured. ## Risks - Database writes and local file writes cannot share one transaction. Recovery accepts only an exact file-for-file retry after a database rollback. Conflicting files remain untouched. - The sidebar displays the current skill. The feed card remains the historical creation receipt. - No database migration, dependency, or workflow change is included. - Remote Daytona behavior still needs a run with a configured immutable image. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and browser verification. OpenAI `gpt-5.6-luna` assisted with bounded implementation and eval work. Both used code execution and tool access. The host did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ae06329971 | test(server): hoist the route module graph in the issue ownership authz suite (#13524) | ||
|
|
e1f245a660 |
fix(server): recover sandbox leases stranded active after a restart (#13515)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server starts and stops provider sandboxes through environment leases > - A restart can leave a terminal run with an `active` lease > - The normal orphan recovery path cannot find that lease after the run ends > - This pull request adds a bounded sweep that changes the stranded lease to `pending_cleanup` > - The existing cleanup sweep then stops the provider sandbox on the same heartbeat tick > - The benefit is that stranded sandboxes stop and do not continue to create provider cost ## Linked Issues or Issue Description **What happened?** A restart can occur after the server writes a terminal run status but before it releases the related environment lease. The lease then stays `active`, and later recovery does not select it. A second path skips the lease when its environment row does not exist. **Expected behavior** The heartbeat recovery path must find an `active` lease that no live run can release. It must move that lease to `pending_cleanup`, and the cleanup sweep must stop the provider sandbox. **Steps to reproduce** 1. Start a run that owns a provider sandbox lease. 2. End the run and stop the server between the run-status write and the lease-release write. 3. Restart the server and allow the heartbeat recovery sweep to run. 4. Confirm that the lease reaches `pending_cleanup` and the provider sandbox receives a stop request. **Paperclip version or commit** This pull request targets the current `master` branch at the base commit used for review. **Deployment mode** The change applies to local development and server deployments. **Installation method** Built from source with the repository test commands. **Agent adapter(s) involved** Not adapter-specific. The change applies to core heartbeat recovery. **Database mode** The change uses the existing database tables. It adds no migration. ## What Changed - Add `sweepOrphanedActiveLeases()` to heartbeat recovery. - Select only stale `active` leases that have no live run owner. - Skip leases with a different live lease for the same provider resource. - Preserve retained leases and write a failure reason for recovered leases. - Limit each sweep to 20 rows. - Run the recovery sweep before the pending-cleanup sweep. - Add focused tests for the recovery guards and same-tick cleanup. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-orphaned-active-lease-sweep.test.ts` — 10 tests pass. - `pnpm vitest run server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts` — 22 tests pass. - `pnpm --filter @paperclipai/server typecheck` — exits 0. - The complete CI suite remains the final check for the repository. ## Risks The sweep changes only stale `active` leases that no live run can release. The stale threshold, live-resource guard, retained-lease guard, and page limit reduce false recovery. The change adds no endpoint, schema change, or migration. ## Model Used OpenAI Codex, GPT-5, tool-enabled coding agent with repository inspection, GitHub CLI, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a8d32e5e61 |
feat(sandbox-providers): add CreateOS sandbox provider (#13434)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent work runs in sandboxes that provider plugins supply > - Operators can choose a provider to run agent work > - CreateOS adds another provider with workspace-preserving pause and resume > - This pull request adds a CreateOS provider plugin > - The benefit is that operators can preserve a workspace between runs without keeping its compute active ## Linked Issues or Issue Description Refs #13203 and the earlier closed #13096. This continues the CreateOS contribution from @bhautikchudasama and @ashwaq06. The branch preserves the original implementation commit. Thank you to both contributors. When squash-merging, preserve the original author's credit in the squash commit body: ```text Co-Authored-By: bhautikchudasama <BhautikChudasama@users.noreply.github.com> ``` The original fork rejects maintainer pushes. This branch includes the merge-conflict resolution and review fixes. The request is described below using `adapter_request.yml`. **Agent or provider** CreateOS sandbox API (https://api.sb.createos.sh). **Why this adapter is useful** CreateOS can pause a sandbox and resume it by ID. The workspace survives the pause. This adds a reusable-lease option to the existing sandbox provider system. **How the agent is invoked** Build and install the local plugin as described in its README. Open Instance Settings, then Environments. Select the `createos` driver. Supply an API key and shape. The driver then supplies sandbox leases for agent runs. **Are you willing to implement it?** Yes. This pull request is the implementation. ## What Changed - Adds the `createos` sandbox provider under `packages/plugins/sandbox-providers/createos`. - Calls the CreateOS HTTP API directly. The package adds no vendor SDK. - Implements the environment lifecycle hooks, incremental process output, and binary workspace sync. - Registers the optional bundled provider and its trusted host credential fallback. The fallback is limited to the official API origin; custom endpoints require an explicit key. - Lists the package in the release manifest with `publishFromCi: false` until its first npm publish is bootstrapped. - Waits through delayed pause/resume state updates without duplicate action requests. - Cancels queued API requests promptly while preserving request spacing. - Uses direct CLI invocation in the setup guide so paths and IDs are passed without an extra shell expansion. - Includes current master and retains its existing Git-subfolder containment fix. ## Demo Fresh setup and a run against a CreateOS sandbox. https://github.com/user-attachments/assets/e71b9e06-c006-4fb9-b847-52dfd68f6110 https://github.com/user-attachments/assets/43b5ac75-66bd-4f76-8563-67e4c7759084 ## Verification All 25 jobs in [CI run 34884260542](https://github.com/paperclipai/paperclip/actions/runs/34884260542) passed at commit `f8d0997677024b784fdadf9d44a84c01cb4e813c`, including typecheck, build, native runner verification, server and workspace tests, browser tests, and the canary release dry run. Greptile reviewed the same commit at 5/5 with no unresolved review threads. GitHub reports no merge conflicts. The remaining merge gate is code-owner approval for the new `package.json`, as required by `.github/CODEOWNERS` and the `master` ruleset. Reviewers have been requested automatically. Local checks passed: - Provider: `pnpm typecheck`, `pnpm test` (52 passed, one live smoke skipped), and `pnpm build`. - Host: focused credential and bundled-plugin tests (17 passed), plus CLI invocation safety (39 passed). - Release: package manifest check and release policy tests (18 passed). The full local `pnpm test:run` attempt caught the README command issue; its focused rerun now passes. The full local run stopped after its general-server group: 7,804 tests passed, with unrelated embedded PostgreSQL startup failures and 10 failures in unchanged runtime-skill-cache tests (`EACCES` on directory rename on macOS). It did not reach the later test groups. Local `pnpm -r typecheck` and `pnpm build` reach the runner package and stop because this machine has no Rust/Cargo installation. The corresponding CI checks passed on provisioned runners, as linked above. The live CreateOS smoke requires explicit provider credentials and was not run during this review. It is available with `CREATEOS_LIVE_TEST=1 pnpm test` in the provider directory. The author supplied the demo links above. ## Risks The provider is opt-in and is not installed by default. It is available through a local-path install or explicit image inclusion. npm publication remains disabled until a maintainer bootstraps the package and enables publishing. Sandbox creation has no idempotency key. An ambiguous create response can leave a resource that requires provider-account inspection. Process tracking is in memory; durable lease recovery belongs to the host. The provider does not advertise guaranteed expiry, interactive login, snapshots, duplex channels, or ingress. Live native-runner qualification remains outside this PR's tested claims. ## Model Used Original provider implementation: human-authored by @bhautikchudasama, as reported in #13203. The original description reports Claude Opus 5 assistance. Review and follow-up fixes: OpenAI GPT-6 via Codex, with code review, editing, and tool execution. The precise runtime model variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks; full-suite environment limits documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: bhautikchudasama <bhautikrchudasama@gmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9adeabb590 |
fix(connections): unblock personal MCP auth discovery (#13497)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections let people give agents access to external tools. > - A personal connection needs the current user's authorization. > - A new MCP URL must be probed before Paperclip can discover its sign-in method. > - Requiring a personal grant before that probe prevents sign-in from starting. > - This pull request permits the initial probe for a creator-owned draft with no credentials. > - People can complete personal setup while later requests retain authorization checks. ## Linked Issues or Issue Description **What happened?** Connecting an unknown MCP URL with "Just me" failed with HTTP 502 and "This connection needs the current user's authorization". Paperclip checked for a personal grant before contacting the provider. The health wrapper also changed the expected authorization error into a server error. **Expected behavior** Discover OAuth and start browser sign-in. Create a personal grant after consent. For a public endpoint, discover its tools and create the empty personal grant after a successful probe. Keep missing authorization on later health checks as HTTP 422. **Steps to reproduce** 1. Add an unknown remote MCP URL with no saved credentials. 2. Select "Just me". 3. Check the link. Before this fix, the request fails before sign-in or tool discovery. **Paperclip version or commit** The three original regressions fail against `6cfe4acff` with the service fix removed and pass with it restored. **Deployment mode** The defect was reported in production and reproduced in local server tests with isolated PostgreSQL. Related work: Refs #11831. Refs #11144. Searches found no duplicate fix. ## What Changed - Allow an initial credential-free probe only for the creating user's personal draft with unknown authentication and no supplied credentials. - Leave OAuth grant creation to the callback. Create an empty personal grant only after a public probe succeeds. - Preserve the personal grant for URLs that already contain a credential. - Make empty personal grant creation conflict-safe without overwriting a concurrent grant or duplicating its creation audit. - Commit the empty grant and audit atomically. Retain the established public draft identity after a later catalog failure, so a failed retry cannot remove a successful retry's grant. - Run the following catalog/default-profile step in its own transaction, so failures discard partial catalog, profile, binding, and audit changes without deleting the established identity. - Verify archived personal connections retain their owner: another user is rejected before probing, while the original owner can resume setup. - Preserve `user_authorization_required` and HTTP 422 in health failures. - Cover the connect and OAuth callback routes, real loopback HTTP, credential-bearing URLs, and later health checks. - Document personal setup and the test fixtures. ## Verification - Red/green: the original three tests failed with the exact reported error before the fix and passed after it. - The credential-bearing personal URL regression also failed before its guard was added. - Focused suite: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/generic-mcp-connection.test.ts src/__tests__/tool-access-service.test.ts` passed all 385 tests across the two suites. An earlier run had a socket hang-up in an existing agent-permissions test; the unchanged suite passed on rerun. - The concurrent rollback regression failed before its fix because the successful retry's grant was deleted. It now verifies the grant and draft survive and a later normal health check succeeds. - Database fault injection during profile-entry insertion reproduced partial catalog writes before the transaction fix. The regression now verifies unchanged catalog rows, no partial profile/bindings, a retained grant, and successful retry. - CI's first serialized-server shard 3 attempt failed an existing peer-agent mutation test (the real run-context guard ran despite the test's mock). The test passed in isolation and all 108 tests in that suite passed unchanged locally. The single failed-shard rerun passed without code changes. - Final-commit CI: all 32 applicable checks passed on `639f037987352cab6084c4ebfa5dbf7b0aed6046`, including all 385 affected tests, the full test matrix, browser suite, build, typecheck, release checks, and security checks. The two Storybook-only checks were not applicable and skipped. Greptile is 5/5 with all review threads resolved. [Successful CI run](https://github.com/paperclipai/paperclip/actions/runs/35027478353). - `pnpm -r typecheck` passed. - `pnpm smoke:mcp-fixtures -- --require-paperclip` passed. - `pnpm build` passed. - Full local `pnpm test:run` was attempted: its general-server group finished with 12,372 passed, 4 failed, and 70 skipped tests. The run started before review edits; its two MCP failures used the old cached service (including an insert without the new conflict clause). All 385 focused tests pass on the final code. The other failures were existing workspace-cleanup and runtime-port tests; their unchanged suites passed on rerun (66 passed, and 25 passed/3 skipped). The local command stopped before later groups. The final-commit CI matrix is the full-suite merge gate; this local run is not claimed as green. ## Risks - The initial probe must not become a general authorization bypass. It is restricted to the creating user's draft. Normal health checks retain authorization enforcement. - Public endpoints get a personal grant with no secrets only after they answer successfully. Credential-bearing URLs keep their existing grant. - The concurrent-probe regression seeds the catalog and default profile to isolate grant creation. Existing first-time catalog/profile creation races are outside this change; this does not claim to make the entire setup flow concurrency-safe. - No database migration or UI change is required. OAuth tests use a simulated provider; the public endpoint test uses real loopback HTTP. ## Model Used OpenAI Codex, GPT-6-based assistant for regression tests and PR preparation; a GPT-5-based Codex assistant assisted with the initial implementation. Exact runtime model IDs and context-window sizes are not exposed in this session. Both used reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |