mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
codex/slack-managed-setup
258
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
34ffc905ea |
fix(slack): recover rejected grants and consume creation confirmations
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
483c466008 |
fix(slack): make manager refresh and installation recovery resumable
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
785ae16699 |
feat(slack): add gated managed setup alongside customer-owned apps
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
57e977be72 |
feat: integrate Pi 1.0 into the experimental Runner (#14921)
## Thinking Path > - Paperclip manages AI agents and their work. > - The experimental Runner owns provider processes and durable sessions. > - Pi needs working task execution and human controls. > - The five-PR stack must preserve changes already on master. > - Each layer now carries the complete integrated source for a safe sequential fallback. > - This PR belongs to native GitHub stack #15602, ending at #14956. ## Linked Issues or Issue Description Refs #14436, #14631, #14743 and #14956. Ship Pi 1.0 through the experimental Paperclip Runner. The five PRs are #14921, #14922, #14923, #14924 and #14956. The user authorized the complete merge after checks pass. Existing `pi_local` execution is unchanged. Accounting and wider provider/platform qualification remain deferred. ## What Changed - Recover missing final replies after workspace finalization changes owners, using accepted-turn evidence without rerunning work or granting external-chat publication. - Preserve the admitted Pi instruction root across warm runs, while retaining changed-root rejection. - Give Pi a bounded 15-second default shutdown grace so stop, drain acknowledgement and durable suspension can complete. Explicit deadlines and other providers retain their existing behavior. - Integrate the Pi 1.0 runtime and master contracts. - Use Pi profile 22. Preserve explicit caller-selected models and exact native thinking levels. Keep Pi's wrapper, helper, extension and question/control behavior unchanged from the qualified profile-19 runtime. - Preserve master's Dot lifecycle and consent fields, configured task environment, status guards and current Codex/Claude dependency versions. Cursor stays qualified. Copilot stays pending; profile 17 binds the changed shared protocol validation sources. - Exclude general AWS IAM credentials from Pi static/custom provider bindings and selected task projections; preserve the provider-scoped Bedrock bearer key. Profile 21 is retained as historical provenance. Rust and cloud install probes use the current declaration. - Patch bundled brace-expansion 5.0.9 to the exact official 5.0.12 payload. Pin the patch and complete runtime closures. Include the patch in normal installed setup tooling. Keep the upstream Pi shrinkwrap as provenance and permit only this exact security correction. - Include current attestation files in the Docker build context. Keep the repository lockfile unchanged from master. CI and private image builds resolve manifest changes before their frozen installation. ## Verification - Full local `pnpm -r typecheck` passes, including Runner Rust, server and UI. Focused integration checks pass: 194 Runner admission/environment tests, 63 profile/credential tests with one expected skip, 152 Dot/UI configuration tests, and Pi transcript/notice tests. - Full local `pnpm build` passes on the final source. - Fresh final-source checks pass: all 698 Rust workspace tests (32 binaries), 156 credential/profile/controller tests with one expected skip, Runner TypeScript typecheck, and 20 package/setup/sandbox tests. - The profile-21 Pi materializer passes on the native host with the official pinned Node 24.21.0 and its npm. It verifies all 150 locked packages, the patched dependency and the exact closure. Setup/package bundle tests and UI token gates pass. - The old hashes were reproduced for all three supported targets before calculating the patched graph. New closure hashes are darwin-arm64 `282022db10150c6632b3444df421342e7d534bdf5d5fb1097a2e79d0625a2bcf`, darwin-x64 `64e251e19009f755c0b04f73ce2138246faab71a961b0f13d75ebfcc34bef12e`, and linux-x64 `713b1fdff42fb56a1518bdc084f181d70bee8ebadc3e4b1d76321ed9108c8410`. Independent native platform execution is separate from graph identity reproduction. - Historical cloud qualification remains unchanged: all seven core cases pass on shipping source `10dc43c9ec65d88c2f782d62afb296d09494f215`, harness `1a4408a48cfb5a1f094a311141c257c92cd7a893`, image `sha256:5b3a775b383591bda1b0c1889e509acc70ce7f37c53f09733c81d59037f02280`, and accepted Sonnet 4.6/low fixture. All 215 canonical files and all seven cleanup checks pass independent verification. These are profile-19 results and are not relabeled as fresh profile-22 runs. - Current Pi digest: `sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`. [The readiness plan](https://github.com/paperclipai/paperclip/blob/codex/pi-production-readiness/doc/plans/2026-10-02-pi-production-readiness.md) preserves campaign and failed-attempt provenance. - Merge only after every PR's current-head CI and fresh review pass. Linux CI covers the full suites, build and browser tests. The local embedded Postgres API-authority suite cannot start on this macOS/Node 26 host, so Linux CI must confirm that suite. ### Fresh profile-22 core qualification — 2026-10-08 All seven accepted core cases pass canonically on Pi profile 22, with `openrouter/anthropic/claude-sonnet-4.6` and native-confirmed low thinking. This model is a fixture; production accepts the caller's explicit Pi provider/model. Runtime/install source: `3241a992f2a7703e59e97ed0fd3e5d6405de4401`. Frozen accepted harness: `1a4408a48cfb5a1f094a311141c257c92cd7a893`. Immutable cloud image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:506f22db7edd78f37c0c40bec1cc084af1850455026dbf467194bfbb8fcef141`. Pi digest: `sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`. [Hosted Linux image and clean-install verification](https://github.com/paperclipai/paperclip/actions/runs/37868328023) passes, including all 20 source-bound archives, normal CLI/Pi setup, companion import and the production pack reader. This exact installation source includes the latest master integration and the corrected Pi warm instruction-root fence. Full local typecheck/build and current-head hosted CI verify the final stack. All 13 focused real-root regressions pass. The full local executor suite passed 662 tests; 15 database tests could not start the Mac embedded PostgreSQL service. Hosted Linux CI passes the full required verification and E2E checks. These fresh results keep their own source identity; profile-19 results remain historical. | Core path | Canonical campaign | Retained archive SHA-256 | | --- | --- | --- | | File edit, validation, download and Done | `pi-core22-replyfix-0-1791511228` | 23 files; `a473e8603a3dd4737863291f8d3d1e392391f0b16d433c3e0e0e9d8baf7a97b0` | | Pending question and controller restart | `pi-core22-replyfix-1-1791511376` | 33 files; `6b829c4eb74e1f32a89c692a4ae7130dbfc1c6d3cf13915effe2103d9e242c8e` | | Three-turn session/process/workspace continuity | `pi-core22-replyfix-2-1791511587` | 23 files; `7a87021f8f9a3fdd3c58bb4467f8d82c635e3ea4795d6e75f144d9aa14818df8` | | Four typed questions and browser reconnects | `pi-core22-replyfix-3-1791511881` | 42 files; `9e31755252be1f4f9cb0626c984c142d4d1ae5f5bee3a7af08444db8d12c280a` | | Plan approval and completion | `pi-core22-replyfix-4-1791512031` | 22 files; `a0383ce1aab38e7b5a25ce0e9dd3bebea5c037ebd96ae6b29dae19015da2ae2c` | | Same-turn steering and permission denial | `pi-core22-replyfix-5-1791512261` | 39 files; `c929b8c7070f0b66aedc17e65ca46e6beab1e363926ac9f7e2a75fb250f05949` | | Stop during pending permission | `pi-core22-replyfix-6-1791512390` | 33 files; `7f0a58ae0f4d5bfc76149435f4e322537089c5bd16e7ffe9b5ad71f10a621a07` | All 215 canonical files (28714587 bytes) are independently hash-verified. All seven cleanup grades pass, with no owned runtime process or temporary root after each case. Automatic retries are zero. The owned cloud host stopped normally after retention. The prior profile-22 warm attempt remains failed and separately retained: archive SHA-256 `1e54eba5ec72b50cee1534b23d1d1d4f21a090006b8a64501ba70db972abfde5`. Its original canonical classification is preserved. Diagnosis reproduced a product bug comparing an agent-files root against an unset checkpoint-only field. The fix stores the admitted physical root separately from the adopted per-run collection capability. The real-root regression fails before the fix and passes afterward, including rejection of a changed physical root. Fixture, grader, model and all seven accepted case IDs are unchanged; this fresh campaign tests final-reply publication after file registration first. The intermediate restart attempt also remains failed and retained: archive SHA-256 `5dcaefdf1d17cf4cd54fd4cf810f45e736667392339b8ce7caf08bb4e225277f`. Its original canonical classification is preserved. Pi resumed, wrote the verified answer and completed its task; exact runner suspension was proven, but idle stop consumed about 5.2s and left under 3s for the drain acknowledgement. The Pi-only default shutdown grace is now 15s, preserving a full 5s drain round trip and a finite suspension reserve. Explicit caller deadlines, other provider defaults, literal drain receipts and exact suspension identity checks remain unchanged. The timing regression fails before this correction and passes afterward; all 18 focused settlement tests and Runner typecheck pass. The final-source file attempt is also preserved as failed (`candidate_failure`), archive SHA-256 `db6767b6773ea618997927ac77bdb005a5ac81492c7b9c0ffbc900449f829bc9`. Native edit, validation, exact downloadable artifact and Done/succeeded all passed, and the exact final reply was durably recorded. A workspace recovery owner completed before the live heartbeat reached presentation, leaving that reply absent from task chat. Recovery now materializes only a completed final reply from the accepted turn of an ordinary internal Done task, preserving issue/run/contract binding, suppression, external-chat authorization and same-run deduplication. The database regression covers the generated file-preparation receipt, suppression, unapproved external continuation and replay. Server typecheck and all 49 response-selection tests pass; hosted Linux verifies the database regression because embedded PostgreSQL cannot start on this Mac. The delayed-final-answer database regression passes on [the final root-source Linux server shard](https://github.com/paperclipai/paperclip/actions/runs/37868262553/job/113628594152), alongside 1,108 passing tests. The first root Runner shard had one unchanged durable-resume test exceed its 5-second timeout; the identical top-source shard and the isolated exact test passed. One rerun of that failed job and its required aggregate passed without source or test changes. The original failed job log and the single-rerun receipt remain retained. ### October 9 merge verification Current merge head: `5a8fe63512a7166aaef5cf50065a25008aa8b44b`. All current-head checks pass, including `ci / verify` and `ci / e2e`; exact-head Greptile review is 5/5 with no unresolved threads. Current master conflicts are resolved. The user authorized the maintainer override of the code-owner review gate after these checks. The seven retained live core cases remain bound to source `3241a992f2a7703e59e97ed0fd3e5d6405de4401` and its recorded cloud image. ## Risks - The security correction changes the dependency closure and profile identity. Old sessions must reopen on the new profile. Exact identities and credential bindings fail closed. - The runner remains experimental and requires explicit selection. Legacy Pi Local is unchanged. Caller model IDs pass through; the E2E model is a fixture. - Accounting and the broad platform/provider matrix remain deferred. This merge does not publish a release or deploy a service. ## Model Used OpenAI GPT-6 through Codex assisted with reasoning, repository inspection, editing and tool use. The exact serving ID and context window are not exposed in this session. Final live qualification uses Pi 1.0.0 with `openrouter/anthropic/claude-sonnet-4.6` and native-confirmed low thinking. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2b688a8fe3 |
feat(connections): add verified MCP providers and setup fixes (#15621)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents governed access to external tools. > - Several provider MCP servers lacked a supported catalog entry or failed during setup. > - Real browser tests identified specific registration, session, and form defects. > - This pull request adds seven catalog entries and fixes the shared paths used by eleven verified providers. > - Users can connect these providers through the existing Apps flow and control each action. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Apps catalog and remote MCP setup, discovery, and action tester. **Subsystem affected** Shared app definitions, the server tool services, and the Apps UI. **Current behavior** Seven providers lack a catalog entry. Airtable can select an advertised client-metadata flow that fails. Calendly rejects the registration name. Tavily needs an initialized session before a call. Firecrawl exposes invalid defaults for optional nested form objects. HTTPS setup can display a callback that differs from the server callback. **Proposed behavior** Add Calendly, Exa, Firecrawl, GSC Wizard, Parallel Search, Tavily, and Windsor.ai. Preserve Airtable, Linear, Make, and PostHog. Use reviewed provider options through the shared MCP and OAuth paths. Display the callback supplied by the server. Omit untouched optional object inputs. **Reason and benefit** Eleven providers passed bounded reads through the real Paperclip browser action tester. This PR includes that verified set and its required shared fixes. Unfinished providers remain outside this change. AgentMail keeps its existing integration. **Breaking changes** No database migration. Existing OAuth credentials and action policies keep their ownership and access rules. Airtable's reviewed method re-registers a retained client-metadata binding through DCR. Tavily's reviewed method initializes sessions before dispatch. The existing customer, managed, and Vercel OAuth gateway paths now refresh on upstream 401 and return `oauth_refreshed_retry_required` (409), requiring an explicit caller retry. They do not replay the rejected call automatically; the next invocation initializes a fresh credential-scoped session when required. Related catalog work: #15545. This branch preserves current master catalog entries and does not add another Google Workspace integration. Open and closed provider PRs and public MCP issues were searched. No duplicate for this verified set was found. The change extends the shipped Connected Apps roadmap item. ## What Changed - Add seven catalog entries with official provider branding and reviewed permissions. - Update Airtable, Linear, and Make metadata and show Make in Apps. - Add authoritative provider source overrides that use the existing generation and review checks. - Add the reviewed Airtable DCR option and compatible Calendly registration names. - Initialize Tavily sessions before discovery and calls, including retained connections. Keep credential-scoped session caching and single call dispatch. - Use the server's callback URL in the OAuth setup UI. - Omit untouched optional object defaults; keep supplied empty strings, false, zero, and object values intact in the action tester, with strict required-child validation. - Require an explicit retry after OAuth refresh instead of replaying a tools/call with missing session headers. - Record all eleven successful reads, write policies, auth-method limits, and qualification caveats. Omit credentials and private account data. - Find chat connector cards through catalog search in browser tests, so pagination does not hide Slack or other later entries. ## Verification - Real browser qualification: eleven bounded reads passed. Catalog refresh and reload proof are recorded in `doc/connections/verified-mcp-qualification-2026-10-08.md`. - Provider writes and actual agent-adapter sessions were not run. Alternative documented auth methods and expiry refresh remain unqualified. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - Focused shared/catalog tests: 41 passed. Source override tests: 3 passed. Catalog regeneration fixture suite: 9 passed. Latest form regression suites: 31 passed. Gateway suite: 38 passed. Other focused UI suites passed before these last fixes. - UI token gates and both token sync checks: passed. - The catalog regeneration fixture includes the authoritative provider overrides; its full suite passes. The focused OAuth socket case passed on isolated rerun. - Complete local UI suite: 7,866 passed. CLI suite: 511 passed, 6 skipped. Complete shared suite: 892 passed. - The unchanged AgentMail/ClickUp discovery fallback suite passes all 45 cases. An OpenAI login test passed on isolated rerun. - Local broad database coverage is limited by macOS PostgreSQL shared-memory exhaustion (`shmget: No space left on device`, not disk space). The broad run was interrupted after diagnosis; no host settings or other running services were changed. - CI found that the Slack browser test assumed its catalog card was on the first page. The test now uses catalog search; the Slack case passes locally. Adjacent GitHub and iMessage cases could not start locally because embedded PostgreSQL initialization failed before browser assertions. - Final commit [`febe4520e`](https://github.com/paperclipai/paperclip/commit/febe4520e13d4b3a4121eafc3d22540c2b11379d): all 54 checks passed, with none pending or failed. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37847182065) passed typecheck, build, all general and serialized test suites, all eight browser shards, runner checks, canary dry run, and the aggregate verification gate. - Greptile reviewed that exact final commit at 2026-10-08 21:33 UTC and returned 5/5 with no outstanding findings or unresolved threads. - PR UI preview against the existing local test server: Calendly catalog/setup observed; Firecrawl read passed after expanding More options with nested optional inputs untouched. This retest proves frontend behavior, not the revised OAuth backend against a live provider. <details> <summary>Browser evidence</summary>    </details> ## Risks - Provider registration and consent behavior can change. Airtable's DCR option is explicit and keeps the existing issuer, resource, redirect, and PKCE checks. - Tavily uses initialized sessions without automatic call retries. After a successful OAuth refresh following 401, callers now receive a retry-required result. A later explicit retry can still require reauthorization if the provider rejects the refreshed token. Provider writes are not live-qualified. - Optional object cleanup affects the shared action tester. Regression tests cover absent objects, required children, defaults, and populated values. - GSC Wizard's underlying Google data scopes and Google flow completion were not independently verified. Its account reported paid/trial metadata of unknown origin. No purchase was performed. - No database migration, new AgentMail integration, or change to existing action grants is included. ## Model Used OpenAI GPT-6 through Codex assisted with implementation, research, tool use, and code execution. The root backend model ID and context window are not exposed in this session. The cheaper subagents used OpenAI `gpt-6-luna`. No model context size is inferred. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted checks and complete UI/CLI/shared suites passed; the broad local database run was blocked by the macOS PostgreSQL startup limitation documented above. The full database/workspace suite passed in CI. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
428b617fc9 |
test(evals): qualify native question resume paths (#15616)
## Thinking Path > - Paperclip lets agents pause tasks for human answers and continue the work. > - Native questions can yield a run or pause inside the provider's running turn. > - Those paths have different run identities and different forms. > - The documentation-placement test requires a semantic three-turn journey. > - A valid provider question therefore needs separate behavioral qualification. > - This pull request adds an opt-in two-answer journey with exact source and resume checks. > - Qualification exposed a Vite startup crash in idle-handler traversal, so the branch also distinguishes Connect mount paths from Express Route objects. ## Linked Issues or Issue Description **What existing behavior does this improve?** Product E2E evaluation of native task questions and answer delivery. **Current behavior** The semantic documentation case rejects the provider adapter's optional Other field before answering. Its three-run expectation also cannot prove a provider question that resumes within the same run. **Proposed behavior** Keep that semantic test intact. Add a separate opt-in case that classifies each recorded question by authoritative identities. Accept the optional Other field only for a verified provider question. Require both user answers, the correct continuation, and one saved final document. Related: #15564 added strict lifecycle checks. #15522 introduced idle-request tracking; this PR corrects its Connect/Vite compatibility after reproducing the startup crash. #14591 changes ACP question drafts and provider support; this PR has no overlapping files with that change. ## What Changed - Add the explicit-only `question-resume` suite for native Codex and ACPX Claude. Reuse the existing neutral choice-then-text task request. - Bind each question to its company, task, agent and run. Require an applied semantic tool receipt or the exact provider request identity. - Require a provider answer to continue the same run. Require a semantic answer to start one new run with the matching interaction and source-run wake fields. - Preserve question forms and both real board answers through the final checkpoint. Reject missing inputs, duplicate answers, extra runs and substituted identities. - Keep the existing lifecycle, output, task ownership and budget checks. Limit the new case to one attempt and one to three runs. - Add mutation calibrations and document the separate qualification claims. Keep production guidance, scheduling, provider code and the historical semantic case unchanged. - Fix startup with Vite middleware: treat Connect string mount paths separately from Express Route objects when traversing idle-request handlers. A real-Vite regression verifies startup and retained async work through an idle hold. ## Verification - 81 focused continuation, question and calibration tests passed before review. After the reference-preservation correction, all 50 question/resume tests pass, including two new negative cases that failed against the previous grader. - Final eval support suite passes: 1,921 Vitest tests and 128 Node tests, with one intentional skip. - Eval and workspace typechecks pass. Workspace build passes. - Catalog discovery finds only the two declared opt-in cells. Existing continuation remains 23 cells; default paid scope is unchanged. - Original [campaign 37831726391](https://github.com/paperclipai/paperclip/actions/runs/37831726391), source `1ac4b2873466896ff2892f49be82904927b6c57d`, failed during server startup before any test or provider run. Its original `transient_infrastructure` result and `cleanup: not_started` are preserved. Local reproduction identifies the Connect string route traversal crash; the question grader was never reached. - Setup-corrected [campaign 37833872898](https://github.com/paperclipai/paperclip/actions/runs/37833872898) measures source `9781a772389d41421df55759004676e1b905e512` with trusted workflow `8e59efc50b162126f33816841e06509634eba65c`. The workflow bytes are unchanged from the first dispatch. Same single selected native Claude cell, one attempt, at most three runs, ten-minute cell deadline, and 1,000-cent company/agent hard stops. Original result **PASS**, 24/24 checks and cleanup pass. Exactly three successful native Claude runs: two applied Paperclip `request_human_input` questions and one finishing run. Both real board answers persist, exact interaction/source-run response wakes match, one revision-1 task document contains Afternoon and the supplied reference, and final task state is done with no active lock, retry, recovery, monitor, pending interaction or child task. No model retry. - Current review head `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e` differs from the measured source only in the question/resume grader and its tests (11 insertions, 3 deletions). The saved document must contain the complete submitted reference, including its prefix and punctuation. Separate provider-free replay of the retained successful campaign passes all 24 continuation/lifecycle checks under this stricter grader. The original result and its measured source remain unchanged; no new provider run was made. - Observed paths were **semantic → semantic**. The provider built-in question/optional Other path was not exercised by this new live attempt. Its form acceptance has retained-original calibration; same-run answer/resume has deterministic positive/negative calibration. This neutral-task pass alone does not qualify the provider path; the subsequent explicit bridge trial below is separate evidence. Neither trial regrades the old Claude failure or establishes a causal behavior/performance comparison. - Subsequent explicit provider-path [campaign 37841107684](https://github.com/paperclipai/paperclip/actions/runs/37841107684) measures the final PR source `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`, using trusted workflow `2a5f65c9ae501b49b2a38ce7edd04209b806c9d0`. Exact cell: `continuation.runner-acpx-claude.local.provider-question-bridge`, Claude `claude-sonnet-5`, suite fingerprint `fc77203bf85fd99b06ebc53cd52982186c0ae0dc369fcc38a9ec3c737e9a31e5`. **Original PASS, 15/15 checks, cleanup passed, one actual provider run and zero eval rerolls.** The built-in question produced one choice plus its optional Other companion. The browser selected the requested reference, left Other blank, and submitted. Exact runtime request/response IDs and the selected option match the saved interaction; the same paused native run created one revision-1 document and finished Done with no pending interaction, lock, retry, recovery, monitor or child task. Independent evidence and runtime-receipt audits pass; marked initial/final screenshots were visually checked. [Original public report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37841107684-1/index.html). - This separate provider-directed trial qualifies one choice-answer bridge and traversal of an empty optional Other field. It does not establish natural tool selection, a typed Other answer, two consecutive built-in questions, free-text-only input, restart recovery or general reliability. No production, prompt or grader change was made for this dispatch. The old failure and the separate neutral semantic-path result retain their original grades. - Preserve within-run friction: one `write_task_document` input-schema rejection and two HTTP 409 “Document does not exist yet” responses preceded a successful `write_document` call. These are three unsuccessful tool calls inside the same run, not three extra provider runs; no flawless document-tool behavior is claimed. Billing reports token usage but remains unpriced/incomplete, so actual charges are unknown and local/hosted runtime is unmetered. Both company and agent hard stops were 1,000 cents, with one attempt, one selected cell, concurrency one and a ten-minute cell deadline. - Billing records all three model runs with token usage, but no priced dollar receipt (`unpriced`, `complete: false`). Actual charges are unknown; local/hosted runtime is unmetered. Numeric zero reported cost does not mean free. - Startup repair: real-Vite reproduction passes; 35 targeted idle-tracking tests pass, 34 database-dependent checks skip because embedded PostgreSQL is unavailable locally. Workspace typecheck and build pass again after the repair. - Full local repository database tests were not repeated because embedded PostgreSQL was unavailable in the preceding workspace verification. Full Linux CI now passes on the final review head. - Final-head [CI run 37837804284](https://github.com/paperclipai/paperclip/actions/runs/37837804284) passes. Complete current-head audit: 53 successful check-runs, two intentional Storybook skips, and separate Snyk success. The ready transition’s contributor and security checks also pass (security scan: no findings). [Fresh review](https://github.com/paperclipai/paperclip/pull/15616#issuecomment-6068092601) is 5/5 on `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`; zero unresolved threads and no merge conflicts. ## Risks This is a new behavioral definition. The neutral two-question trial records semantic-tool use; the separate provider-directed trial records one built-in question. They do not establish consistent tool selection, documentation placement, default hiring, remote environments, crash recovery or general reliability. The separate explicit bridge trial covers one provider choice, an empty optional Other companion and same-run completion. Typed Other, two consecutive built-in questions and free-text-only provider input remain unqualified. The observed document-tool schema rejection and two 409 responses are retained as a separate follow-up, without attributing their cause. The historical case and its original grades remain intact. The reference question must still be text-only; the oracle does not turn an arbitrary extra question into an optional companion. Missing or unfamiliar identity evidence fails closed. The only production change distinguishes Connect string mount paths from Express route objects during idle-handler traversal. Unknown route objects still fail closed. No API, schema, migration or workflow change is included. ## Model Used OpenAI Codex, GPT-6-based, with tool use and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0ac194450a |
fix: make Copilot provider-pack wrappers portable (#15586)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Container images include a movable provider pack for agent execution. > - pnpm executable wrappers can contain the temporary build directory. > - New optional Copilot packages add wrappers that the build does not replace. > - This pull request gives those installed wrappers relative executable paths. > - Image publication can finish while the existing path check remains enforced. ## Linked Issues or Issue Description Refs #15572 and #15560. The small portability loop comes from cryppadotta's larger Copilot runtime PR #15560. This separate fix repairs image publication without waiting for that feature's qualification and runtime changes. That PR can remove its duplicate loop after this lands. **What happened?** The standard Docker build stops with `Provider pack shim copilot-linux-x64 retains its temporary build path`. The failure occurred before and after #15522. See [the failed master build](https://github.com/paperclipai/paperclip/actions/runs/37802065316). **Expected behavior** Installed native Copilot wrappers resolve their pinned executable after the provider pack moves. Optional packages that are absent do not gain a command. The builder still rejects wrappers with temporary paths. **Steps to reproduce** 1. Run the provider-pack build from the affected master revision on Linux x64. 2. Let `pnpm deploy --prod` install the optional Copilot package. 3. The wrapper scan rejects its temporary `NODE_PATH`. **Paperclip version or commit** Master `3367b75ccce34d02f355cda1f1ed3fe0b34cf93d`. **Deployment mode** Docker and provider-pack builds. ## What Changed - Replace installed Copilot platform wrappers with relative native executable launchers. - Move the existing executable wrapper writer into an importable helper. Retain the same launch behavior for Node, Claude and OpenCode. - Test relocation, argument handling, exit status, absent optional packages and missing executable packages. Register the tests in the existing Runner test preparation command. - Hash the helper in Daytona image identity and prove that helper changes invalidate the image cache. - Document the packaging rule. Keep dependencies, provider qualification and the final temporary-path check unchanged. ## Verification - 17 Node packaging tests passed across the new wrapper tests, provider-pack release tests and candidate selection tests. - Six existing bundled remote-provider-pack tests passed using the Runner Vitest configuration. - Nine Daytona image identity tests and Runner E2E typecheck passed after the cache-input correction. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - Real `pnpm deploy --prod` reproduction on macOS ARM64: the original Copilot wrapper contained the temporary path. After the rewrite and directory relocation, the actual executable returned Copilot CLI 1.0.88 with exit code zero using only `/usr/bin:/bin` in `PATH`. - Syntax checks, `git diff --check` and the pre-push secret scan passed. - The full local `pnpm test:run` reproduced the same five skill/connector fixture failures observed earlier in this workspace. It was stopped after current-head clean-checkout CI passed; later local phases were not run. This local run is not claimed as passing. Focused packaging tests, workspace typecheck/build and all hosted CI passed. - [Hosted Docker verification passed](https://github.com/paperclipai/paperclip/actions/runs/37804902678): Linux AMD64 and ARM64 image builds, multi-architecture publication and the process-reaping smoke check. This run tested `779d94989c54d9abbeba0838194c951186af66a6`; the only later changes are Daytona cache identity and its regression test. Provider-pack build code is identical. Local Docker did not respond within the bounded probe. - Current head `db2c19e270b8d5a7bab5db39d3de0e4761cf57c3`: 54 successful checks, two conditional skips, no failures and no merge conflicts. Apex is 5/5 with no unresolved comments. ## Risks The helper uses each installed package's exported executable. An installed wrapper with a missing package still fails the build. This change does not execute Copilot during image construction, alter dependency pins, change runtime admission or weaken the temporary-path check. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use and code execution. The exact serving model identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; the full local-suite limitation is documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
941a3fa991 |
Pin Copilot native dependencies with the maintained lockfile refresh (#15572)
Pin the three optional GitHub Copilot 1.0.88 native packages for Runner and server using the maintained lockfile workflow. Synchronize the package contract and bound initial render readiness in the deliberately throttled browser fixture. Current-head CI and focused checks pass. Co-Authored-By: Dotta <cryppadotta@users.noreply.github.com> Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
24a15c5afa |
test(evals): verify durable waiting across continuation checkpoints (#15564)
## Thinking Path > - Paperclip manages agents and durable tasks across provider runs. > - Human questions must preserve task ownership and stop work until a real answer arrives. > - The continuation eval checks the final work, but its lifecycle oracle misses several broken waiting states. > - A correct final answer can hide a stale execution lock or a lost intermediate answer receipt. > - This pull request checks each wait and retains every question and run identity. > - A source audit records which waiting operations belong to native runtimes and which still require legacy API calls. ## Linked Issues or Issue Description **What existing behavior does this improve?** The existing Product E2E continuation oracle for human question and approval journeys. **Current behavior** A pending interaction can pass the waiting check even when its task has the wrong status, a stale lock, or a retry. Only the first pending question and first run receipts are checked at the final checkpoint. **Proposed behavior** Require a healthy wait on the same task and assignee. Preserve all intermediate run receipts and every pending question's answered identity. Allow a paused native provider question only when its pending runtime request identifies the running native run and execution lock. Related: #15548 and #15554 cover earlier bookkeeping slices. #15544 changes production continuation summaries; this PR changes the eval oracle and does not overlap that fix. ## What Changed - Check waiting state, execution ownership, and all answer/run identities in the existing lifecycle oracle. - Add negative calibrations for broken records and positive coverage for both semantic waits and paused provider questions. - Retain task activity at each checkpoint for inspection of successful persisted mutations. - Verify 1,000-cent company and agent budget limits before continuation work. Disable automatic cell rerolls. - Document runtime ownership, existing coverage, and the remaining instruction decision. Keep production instructions unchanged. ## Verification - `pnpm test:e2e:runner:unit`: 1,867 Vitest tests pass, one skips; 128 Node tests pass. - `pnpm test:e2e:runner:typecheck`: passes. - `pnpm -r typecheck`: passes. - `pnpm build`: passes. - Negative calibration: 27 added cases fail against the prior oracle and pass with these checks. - Four existing local continuation cells are selected for a separate bounded live canary. Live results are pending; no behavioral pass is claimed here. - The full local `pnpm test:run` suite was not repeated because embedded PostgreSQL was unavailable in the preceding workspace verification. Required Linux PR CI must pass before readiness. ## Risks The stronger oracle can expose existing product or fixture defects. A paused provider question and a terminal semantic wait have distinct valid states. The four-cell canary does not qualify approval/review, dependency unblock, crash races, remote execution, or general task quality. Task activity records successful persisted writes, not failed API attempts; repeated progress comments are not automatically defects. Historical eval grades remain unchanged. No production scheduling, prompt, tool, schema or migration changes. ## Model Used OpenAI Codex, GPT-6-based, with tool use and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2f0c485dec |
fix(skills): ship the completion helper with the installed skill (#15554)
## Thinking Path > - Paperclip manages work for AI agents. > - Legacy agents use the Paperclip skill to save task status and comments. > - The skill names a script relative to the task workspace. > - That script exists only in the Paperclip source repository. > - Agents in other workspaces can hit a missing command or search for it. > - This PR ships the helper inside the skill and uses the installed skill path. > - The repository command remains available through a forwarding wrapper. ## Linked Issues or Issue Description Fixes #9527. Refs #15548 for the preceding runtime checkout guidance. Related: #6052 addresses LF line endings for the repository helper; this change addresses helper delivery and path resolution. ## What Changed - Bundle the existing issue update helper with the Paperclip skill. Preserve its HTTP checks, echoed-status check and two-attempt limit. - Resolve the command from the installed skill directory. Use a verified PATCH when that path is unavailable, without searching the filesystem. - Keep the repository command as a wrapper that works from any directory. - Test shell execution and exact status/comment payloads through both provider skill-home layouts, including paths with spaces. - Add helper sources and existing verification tests to stock-harness admission. Record an absent historical helper explicitly. Add the missing declaration for the admission fingerprint export. ## Verification - Complete directly affected source suites: 30 tests pass. They cover skill delivery, preserved multiline comments and links, authentication headers, empty responses, mismatched status, transient retries and definitive rejections. - Product E2E typecheck passes. Support suites: 1,835 Vitest tests pass, one is skipped; 128 Node tests pass. - Full local build and workspace typecheck pass. - Full local repository tests are not claimed as passed. Embedded PostgreSQL was unavailable in this worktree during the preceding task; Linux CI will run the repository gates. - The authorized matched Codex/Claude comparison is pending. It uses the existing assigned-skill case and original oracle, one initial attempt per profile and variant. - CI and a fresh Greptile review are pending. Keep this PR in draft until readiness gates complete. ## Risks - Correct path resolution depends on the harness supplying the installed skill path. The instructions use verified PATCH when that path is unavailable. - The helper still requires Bash, curl and jq. Its existing retry and response-verification behavior is unchanged. - Tests use the shared skill-directory symlink mechanism and an HTTP fixture. Real provider completion behavior still requires the bounded live comparison. - This fix does not redesign native completion, legacy recovery or ambiguous transport handling. ## Model Used OpenAI Codex, GPT-6 family. The exact model build and context window are not exposed in this session. Used code editing, shell tools and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
71cd0a2621 |
fix(skills): honor the current run harness checkout (#15548)
## Thinking Path > - Paperclip manages work for AI agents. > - The runtime claims eligible assigned tasks before it starts an agent. > - The wake tells the agent when the runtime already holds that claim. > - The legacy skill still requires another checkout in every case. > - This PR makes the skill honor the current task and run claim. > - Manual checkout and server ownership checks remain in place for other cases. ## Linked Issues or Issue Description **Where is the issue?** `skills/paperclip/SKILL.md`, in the scoped wake procedure and Step 5. **What's wrong?** The wake can say that the harness already checked out the issue. The skill still tells the agent that it must call checkout. These instructions conflict. **Suggested fix** Skip the second checkout only when the runtime wake explicitly confirms the claim for this issue and run. Retain manual checkout when that statement is absent or the agent selects another task. Refs #14948 for the existing shared prompt reduction. ## What Changed - Honor the explicit runtime claim in the scoped wake procedure and Step 5. - Keep context reads, status writes, deliverable handling and conflict rules. - Add checks for normal and resumed wake text and excluded automatic claims. - Retain successful checkout HTTP activity for legacy stock-task evals. Bind each receipt to the exact company, task, agent and run. Keep this observation separate from the original task grades. ## Verification - Checkout observation calibration: nine tests pass. - Focused skill, wake and database ownership tests: in progress. - Full repository build, typecheck and tests: in progress. - Planned live comparison: the existing assigned-skill document case on legacy Codex and Claude. One attempt per variant and profile. No automatic retries. The baseline and candidate share the observation code and task oracle. - Live results are pending. This draft does not claim behavioral qualification. ## Risks - Agents may misread prompt guidance. The API still enforces ownership; the text grants no new authority. - The exception is specific to the current issue and run. It does not remove ordinary legacy completion writes or authorize another task. - Activity measures successful checkout HTTP calls. Failed attempts require separate run-log inspection. Missing or mismatched observations cannot count as zero calls. - One trial per profile cannot establish general reliability, speed or cost trends. ## Model Used OpenAI Codex, GPT-6 family. The exact model build and context window are not exposed in this session. Used code editing, shell tools and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d66acb7ac1 |
feat: automate Slack bot app setup and installation (#15413)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors give each agent a customer-owned bot and task-backed conversations. > - Manual Slack setup requires app creation and copying durable credentials. > - Operators need a shorter setup that an assisting agent can use safely. > - This pull request creates the app through Slack's Manifest API and installs it through OAuth. > - Durable registration state supports recovery without creating another app. > - A four-screen wizard, automatic avatar upload, and OAuth account linking reduce setup work. > - Connector settings and per-turn tool guidance support daily use after installation. ## Linked Issues or Issue Description **Subsystem affected** Native Slack bot setup, company secret storage, chat connector management, and agent tool guidance. **Problem or motivation** New Slack bots require manual app creation and copying a signing secret and bot token. Interrupted setup can create duplicate apps. The setup and management screens contain unnecessary controls. Agents also need guidance for native questions, files, thread replies, and governed Slack actions. **Proposed solution** Use a temporary app-configuration access token to create a customer-owned app. Save durable secrets in the vault. Bind OAuth to the initiating actor, company, endpoint, registration revision, scopes, and configured origins. Preserve manual and existing-app recovery. Link the installing user's account, send a welcome DM, and advance from saved server evidence. Keep request-URL recovery instructions available if automatic connection detection waits. **Alternatives considered** The Slack CLI adds installation requirements. Socket Mode changes transport. A shared Paperclip-owned app changes app ownership. These alternatives are outside this change. **Roadmap alignment** This extends existing chat connectors and secrets capabilities. Related public work: #14037 and #13954 cover Slack MCP prerequisites and user OAuth. No duplicate bot-registration PR was found. ## What Changed - Share one reviewed manifest builder between automatic registration and manual setup. - Add replay-safe migration 0318 and company-bound registration state with vault references and uncertain-creation recovery. - Add registration, installation, callback, and resume APIs with short-lived, single-use OAuth state. - Save installation credentials before downstream checks and preserve bot identity constraints. - Reduce automatic setup to four screens. Keep advanced app details, manual recovery, and existing-app setup. - Upload the agent avatar with the Paperclip dark background. Link the OAuth installer's account and send setup DMs. - Show agent and connector-owner avatars. Simplify settings, access, and conversation screens. - Discover joined Slack channels and enable them by default. Start a task from a bare mention and admit same-thread follow-ups. - Refresh Slack tool guidance each turn. Add native-form, file, approval, and delivery regressions plus manual model probe definitions and sanitized acceptance records. - Update deployment/database docs, OpenAPI, redaction, removal cleanup, production Storybook stories, and provider browser tests. - Merge current master and move the registration migration after its latest migration without rewriting published commits. The completed Slack success view intentionally has a single centered **Done** action and no **Save & exit**, as explicitly requested by the product owner. `DESIGN.md` records this exception; unfinished setup steps retain the aligned wizard footer. ## Verification - Passed after the master merge: repository typecheck, full build, Storybook build, design-token gates, module-boundary gates, and migration generation. - Passed: all 352 focused Slack deterministic tests and all 14 affected provider browser tests. Browser tests use controlled provider fixtures and a separate throwaway instance. - Passed on current head `c5d01e0e2`: the complete GitHub test matrix (general server, chat, all workspaces, serialized server, and Runner), all eight browser shards, typecheck/release registry, build, canary dry run, security checks, and policy gates. There are 52 passing checks and no pending or failing checks. - Greptile completed on the exact current head with 5/5 and no actionable findings or open review threads. - Local repair verification passed 93 focused tests, including same-app reinstall after revocation and rejection of consent started before revocation, the AgentMail browser journey, and repository typecheck. Local build and Storybook build also passed. The redundant local full-suite rerun was stopped after the complete current-head CI matrix passed. - Real Slack setup and agent replies were exercised in the authorized isolated test drive during the setup iteration. - The ten additional model probes were attempted with legacy `codex_local`, `gpt-5.6-sol`: five passed, two failed, and three were partly verified. Native runtime is not qualified. See `server/src/services/connectors/slack/evals/2026-10-08-acceptance.md` for evidence and limits. - Passing model probes cover native forms, downloaded file bytes, bare mentions with thread replies, explicit posts/reactions, and saved approval denial. - The controlled uncertain-write probe found wrong delivery-check IDs. The canvas fallback attempt used an invented tool name. Search pagination/native search, a private-source denied-tool receipt, and distinct board/webhook origins remain unqualified. Reviewer path: enable Chat connectors, start Slack chat setup, select an agent, enter an app-configuration access token, and approve Slack installation. Send a message to the bot and confirm that setup advances to success. Inspect settings and allowed channels. See `doc/connections/SLACK-AUTOMATIC-SETUP.md` for deployment and recovery. ## Risks - Slack app creation has no provider idempotency guarantee. A timeout after dispatch stays uncertain until the operator checks Slack. - OAuth needs a stable public HTTPS board origin. Webhook ingress may use a separate configured HTTPS origin. Workspace policy can delay installation. - Migration 0318 can replay safely on instances that applied the earlier development migration. - OAuth installation now links the installer to the initiating Paperclip user. Identity checks and company access rules still apply. - Joined channels now enable bot responses by default. Linked-user authorization and per-action approval rules still apply. - Model behavior has the documented delivery-check and canvas fallback failures. A passing CI run does not establish that every model probe passed. - Removing the connection does not delete the customer's Slack app. No new first-party telemetry is added. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser verification. The runtime does not expose a more specific authoring model ID or context-window size. The live bot probes used OpenAI `gpt-5.6-sol` through `codex_local` in legacy mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b5342febe5 |
fix(runner): require current-turn completion after connection continuations (#15514)
## Thinking Path > - Paperclip manages agents, tasks, permissions, and execution budgets. > - Native tasks can resume after a connection decision in the same provider conversation. > - Each turn still needs an accepted completion report. > - Compact continuation messages did not explain that reports from earlier turns cannot finish the new turn. > - The connection evaluator could also grade before the final reply was stored or reject valid unavailable-access wording. > - This PR clarifies the current-turn report requirement and fixes those observation boundaries. > - The original connection instructions and strict native completion gate stay in place. ## Linked Issues or Issue Description Refs #15489. The reduction remains draft while this separate repair is qualified. Refs #15471 for the earlier connection continuation work. ## What Changed - Add a current-turn completion reminder to compact continuation inputs. - Keep final prose insufficient for completion. Preserve permissions and retry policy. - Wait for the final successful task run's saved, attributed decline reply within the existing deadline. - Use one bounded explanation matcher for both decline checks. - Wait for a recorded tool-action rejection to dispatch its bound continuation, with strict company, task, agent and source-run checks. - Retain the exact grading input before later API refreshes. - Add failure and delay regressions and update the Runner and evaluator docs. ## Verification The fresh comparison has **15/15 original passes on each variant**: 15 unchanged pass pairs, zero new failures, and no pending pair. There are 30 case attempts and **65 actual agent runs** (baseline 33; candidate 32). All runs succeeded. All 30 cleanups passed. No model attempt was retried. | Profile | Baseline | Candidate | | --- | --- | --- | | Native Codex `gpt-5.6-sol` | 5/5 | 5/5 | | ACPX Claude `claude-sonnet-5` | 5/5 | 5/5 | | OpenCode `openrouter/deepseek/deepseek-v4-flash-0731` | 5/5 | 5/5 | Each profile covers service approval, service decline, connection decline, provider decline, and selection of the second provider. The saved replies, approved briefings, decisions, fixture observations, final task states, and native completion records were inspected. All 18 saved decline-grade snapshots match their original captured inputs and checks. Result, API snapshot, and final ledger run sets agree. - [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37711658378) · [public candidate report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711658378-1/index.html) - [Baseline campaign](https://github.com/paperclipai/paperclip/actions/runs/37711675579) · [public baseline report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711675579-1/index.html) - Candidate source: `79905343bba280d462765faad19a26e7f179259e`. Baseline source: `7c5e120158f1385a1fc5f20be41f66c58a605534`. Both use master context `fc6304dfe5f2e446e09bd052a7b45f51e930f250`. The trusted workflow source is separately frozen at `dd777f4b7343305c4e6f44c422f44a1d78e12e4f`. - Both variants use the same evaluator, fixtures, models, permissions, 720-second cell deadline, and 1,000-cent company and agent hard stops. The only production difference is the compact continuation reminder. - Suite hash: `aee30b74b4d38ada08777798db0932fbc64e368bb427d29caa8cfd86c7f59747`. Definition hash: `ea9e17f54fe0af1acbb2337ddfab8a7e92db8488bd0f59510a63095ca0229b60`. - Provider-free transport capture: startup/resume input stays at 55,726 bytes. Compact continuation input grows from 53,334 to 53,535 bytes. The 42-tool catalog stays unchanged. These are Paperclip input bytes, not complete vendor prompt tokens. - Focused evaluator tests: 31 pass. Native contract, transport delivery, and session tests: 194 pass. Evaluator support: 1,819 TypeScript tests and 128 Node tests pass, with one intentional skip. - Full build, workspace typecheck, and evaluator typecheck pass. Current-head CI passes all required gates. The current rollup has 51 successful check runs, two intentional Storybook skips, and a successful Snyk status. Review is 5/5 with zero unresolved threads. - CI attempt 1 had one initial runtime-fixture health timeout. The exact test and its full 164-test file pass locally. One CI shard retry passed. The original CI failure, its dependent verify failure, and the retry remain visible in [CI history](https://github.com/paperclipai/paperclip/actions/runs/37711235060). - The broad local `pnpm test:run` attempt was interrupted after about 49 minutes (exit 130). It recorded one failure in the unchanged Zep memory-connector disabled-setting test. That test and the full 388-test tool-access file pass in separate local checks; current-head CI also passes. The local cause is not established, and this broad local attempt is **not** claimed as passing. Three earlier local failures also pass in their isolated checks; their original logs remain retained. - The first two setup admissions were cancelled before provider jobs to include the review correction. They made no provider calls. The completed campaigns above are the first and only model attempts for these corrected variants. Cost evidence stays separate from behavior. Original result summaries report only OpenCode amounts: baseline $0.039192096 and candidate $0.063221620. Final run ledgers also retain estimates for Claude (baseline $1.232439000; candidate $1.333611200) and Codex (baseline $2.340324400; candidate $2.129335600). These estimates do not replace the original summaries. Local and GitHub runtime are unmetered here. Invoices are unknown. This is not a cheaper or faster claim. ## Risks - A single matched trial cannot prove general equivalence or causation. The reminder is an instruction change, not a new completion enforcement rule. - Candidate OpenCode service-decline finished within one continuous run; its baseline used two. That pair passed the task outcome, but it does not qualify the reminder on a resumed decline turn. No extra paid run was used to replace it. - The text matcher is bounded evidence of an explanation. It does not prove reasoning or consumption of feedback. Bounded stdout excerpts do not prove that every extra attempted tool call is absent. - Missing saved replies or continuations still fail at the original deadline. Failed native completion remains a failure even when final prose is correct. - The unresolved local-suite discrepancy above remains a validation limit. Full remote CI and both focused local reproductions pass. - The original connection instructions stay in place. These results do not qualify the reduction in #15489. Its original 11/15 versus 12/15 grades and two new failing pairs remain unchanged. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, test inspection and eval analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the focused local tests listed above and they pass; the interrupted broad local run is disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ae6f95ed7a |
feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults. Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ceabc3bc88 |
fix(e2e): wait on the server's own deferral signal in the signoff wakeup helper (#15345)
## Thinking Path > - Paperclip coordinates work for AI agents through server-managed tasks and heartbeats > - Signoff policy tests exercise stage transitions that wake the next participant > - A stage transition can defer a wake while the previous run still holds the issue execution lock > - The test helper used a fixed wait and could fail before the server released that lock > - This pull request waits on the server deferral signal and verifies the returned run identity > - The benefit is a stable test with clear failure details and no repeated wake request ## Linked Issues or Issue Description **What happened?** The signoff policy end-to-end spec failed intermittently with `No issue-bound heartbeat run became available for agent <id>`. The server had deferred the wake while another run held the issue execution lock. **Expected behavior** The test helper waits for the named blocking run to finish, then reads and verifies the new run for the requested agent and issue. **Steps to reproduce** 1. Run `npx playwright test --config tests/e2e/playwright.config.ts tests/e2e/signoff-policy.spec.ts`. 2. Exercise a signoff stage transition while the previous stage run still holds the issue execution lock. 3. Confirm that the helper waits on the server deferral signal and returns only a matching run. This change relates to [PR #11299](https://github.com/paperclipai/paperclip/pull/11299), which also touches the signoff policy end-to-end spec. ## What Changed - Call the wakeup endpoint that reports the server deferral reason. - Wait for the named blocking run instead of using a fixed delay. - Verify that each returned run belongs to the requested agent and issue. - Bound read and wait loops and report the last server signal and lock fields on failure. - Keep one wakeup request per helper call. ## Verification - `npx playwright test --config tests/e2e/playwright.config.ts tests/e2e/signoff-policy.spec.ts` passed locally. - All five signoff policy cases passed locally. - The spec keeps all 60 assertions. - The diff contains only `tests/e2e/signoff-policy.spec.ts`. - The full CI end-to-end shards must pass before merge. ## Risks - The helper now depends on the wakeup endpoint's deferral fields. - A server change to those fields can fail the test with a named diagnostic. - The change affects end-to-end test behavior only. ## Model Used OpenAI GPT-5 Codex, exact runtime model `gpt-5-codex`, with tool use and code review assistance. The runtime does not expose a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1640c5b6ab |
feat(ui): render HTML artifacts in a secure sandbox (#15447)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents deliver reports as task attachments and workspace files. > - The board already renders Markdown, but it shows HTML reports as source or rejects their preview. > - Reports can need inline scripts to build charts and tables. > - Artifact scripts must not read board cookies, storage, or the parent page. > - This pull request adds an opaque-origin HTML preview and view controls beside Download. > - Users can explore a report, inspect its source, and download the original file. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task attachment panel and workspace file viewers. **Subsystem affected** ui/ and the server workspace file preview service. **Current behavior** HTML attachments show source text. Workspace HTML files cannot be previewed. **Proposed behavior** Render self-contained HTML reports in a sandboxed iframe. Place eye and code controls beside Download. Preserve the original file for raw view and download. **Reason and benefit** Users can read and filter agent reports in Paperclip without giving artifact scripts access to the board session. **Breaking changes** HTML previews now open rendered. Scripts and styles must be embedded in the report. Remote resources, API requests, forms, popups, and host navigation are blocked. Workspace HTML content uses the existing bounded UTF-8 JSON response. **Additional context** Related closed proposals: #3293 and #4857. This change uses the current attachment and workspace viewers and adds no report-serving endpoint. The Artifacts and Work Products roadmap item is complete; this improves its existing preview behavior. ## What Changed - Add a shared HTML iframe renderer with `sandbox="allow-scripts"` and no `allow-same-origin`. - Install a restrictive CSP before artifact markup. Keep inline report scripts and styles. - Add shared rendered/raw icon controls beside Download in the attachment panel, workspace panel, and file sheet. - Return workspace HTML as bounded text inside JSON. Keep download and path-access protections. - Add Storybooks for reports, security probes, workspace viewers, a task journey, and mobile layouts. - Document the security boundary and browser test command. ## Verification - `pnpm -r typecheck` and `pnpm build` pass. - Token gates and the static Storybook build pass. - Four browser tests pass. They cover report filters, raw mode, task entry, workspace viewers, cookies, storage, host DOM, resource requests, forms, popups, and host navigation. - The focused UI suite passes 26 tests. The file-resource server suite passes 36 tests. - Local full-stack acceptance passes with an actual 52 KB HTML report. Its chart and filter render. Raw mode preserves the source. Download bytes match the original file. - The complete local UI suite passes: 698 files and 7,754 tests. - `pnpm test:run` was attempted. Its server group recorded two unrelated timeouts and eight connector failures. All failed cases pass in isolated reruns, including the complete 388-test connector suite. The remaining local run was stopped after the full CI test coverage passed, to release test database resources. The full local command did not complete cleanly. - The separate local CLI group passes 511 tests. Three worktree database cases fail to start embedded PostgreSQL because this Mac has exhausted its shared-memory allocation. One isolated rerun fails at the same database startup step. No CLI code was changed. The corresponding CI test jobs pass. - All 56 PR checks are clean on `c771255f1a2e86bd8825f22eafb5d14a5e342cce`. Greptile gives 5/5 with no actionable findings. The security scan passes. The branch has no merge conflict. - Review the **HTML artifacts** section in Storybook. Run `pnpm exec playwright test --config tests/html-preview/playwright.config.ts` for the browser checks. ## Risks - Never add `allow-same-origin` to this iframe. Its opaque origin is the cookie, storage, and host-DOM boundary. - Reports that depend on CDN assets need to embed their dependencies. - The frame can navigate itself. Sandbox restrictions remain after navigation. This is not full network isolation. - Inline scripts can consume browser resources. This change does not isolate CPU or memory use. - No database migrations or API shape changes are required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository editing, code execution, and browser tool use. The runtime does not expose a more specific model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — feature and UI checks pass; the full local run has the infrastructure limits recorded above. All CI test jobs pass. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d0db8820db |
fix(connections): repair native baseline and approval continuations (#15420)
## Thinking Path > - Paperclip manages AI agents and the tools they may use. > - Connection setup separates provider preference from permission to use a tool. > - The first native connection baseline could not exercise its intended decisions. > - The browser used mutable task titles, and the provider fixture already granted access. > - Native provider-choice instructions also disagreed with the preferred question format. Schema rejection gave no field guidance. > - This pull request repairs those test preconditions and native guidance, then fixes restart/approval defects exposed by the corrected baseline. It also restores missing OpenCode tool-error evidence. > - The benefit is an inspectable baseline before any further instruction reduction. ## Linked Issues or Issue Description Refs #15407. The original 15-cell baseline remains 0 PASS / 15 FAIL. Ten cells stopped on stale titles, two Codex cells had schema denials, two OpenCode cells used already-granted tools, and one Claude cell returned no native result. No intended user decisions were submitted. The exact invalid Codex field and underlying Claude failure cause remain unknown. [Original campaign](https://github.com/paperclipai/paperclip/actions/runs/37562577199) · [Original report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37562577199-1/index.html) ## What Changed - Match the browser's task route and visible identifier instead of a title the agent can change. - Start native provider-choice fixtures with no agent tool access. Verify the public effective-access records. - In the positive case, select Arcade, then grant its exact HubSpot tool through the real access card. Require both saved decisions and exactly one observed call. - Return a canonical `providerQuestionSet` for native input and retain the equivalent legacy `providerQuestion`. - Keep invalid input rejected. Return bounded schema locations and required field names without submitted values. - Preserve Claude's exact session/content identity while allowing authenticated registered instruction-copy paths to rotate on a new run. - Reject duplicate approval reports for an existing exact tool-action card before they create another human review. - Wait for a recorded service-approval continuation within the existing deadline; retain missing or failed continuation grades. - Forward OpenCode tool activity through the runner facade, preserving bounded errors and execution-part identity without inventing host-call joins or exposing arguments. - Preserve original grades, costs, scope limits and diagnoses in the dated repair report. ## Verification - Eval typecheck passes. Support suite: 1,801 PASS, one intentional skip; Node checks: 128 PASS. - Connection/schema tests: 51 PASS. Real-server public fixture setup: one PASS with zero providers. - Browser support regression: five PASS, including renamed and wrong tasks. - Focused Rust safe-feedback test: one PASS. - Repository typecheck and build pass before the latest master replay. Post-replay connection/shared/real-server fixture checks: 52 PASS; eval typecheck passes. The browser review fix additionally passes all five browser checks and seven suite checks. - The full local repository run was interrupted incomplete after about 45 minutes, with five integration failures retained. All five pass in a separate targeted invocation (1,250 unrelated tests skipped). No full local-suite pass or root cause for the initial local failures is claimed. - Corrected frozen source `162cc90fdabe7f505b88ae095044531b82784c92`: **10 PASS / 5 FAIL** across the [passing Codex canary](https://github.com/paperclipai/paperclip/actions/runs/37575158761) and [remaining 14 cells](https://github.com/paperclipai/paperclip/actions/runs/37576261807). The canary passes all 17 checks. Claude's two provider-choice continuations fail on restart, Claude service approval exposes an early evaluator rejection, Codex service approval creates a duplicate approval, and OpenCode provider-second times out after both decisions with no HubSpot call. No original result is regraded. - Final ledgers count 31 actual runs: 27 succeeded, two failed, two cancelled during cleanup. All 15 cleanup/budget checks pass. The late Claude continuation is absent from its earlier workflow snapshot; it remains in the result/API/final ledger. Original evidence retains 279 hashes. Recorded LLM subtotal $0.04553787 is incomplete billing, not actual total cost; local runtime is unmetered. - New repair regressions reproduce the Claude attach failure, duplicate approval acceptance and dropped OpenCode tool events before their respective fixes. Nine Rust attachment checks, 127 ACPX host/adapter tests, 33 completion/control-plane checks, nine eval deadline tests, 59 OpenCode proxy/driver tests, one Rust tool-error/redaction check, and TypeScript/Rust composer parity pass. Eval typecheck, repository typecheck and build pass. Existing support coverage is 1,802 PASS plus 128 Node PASS, one intentional support skip; two additional deadline tests also pass. - New-source full CI/review and live canaries are pending. The next bounded selection is Claude provider-decline, Codex service-approve and one OpenCode provider-second diagnostic with repaired event evidence. No broader campaign or instruction-reduction qualification is claimed. - Initial corrected campaign [37574251834](https://github.com/paperclipai/paperclip/actions/runs/37574251834) was cancelled during shared build after review found the breadcrumb whitespace assumption. Its matrix job has zero steps and no provider execution. The real adjacent-span browser regression now reproduces the old failure and passes after the fix. ## Risks - The corrected baseline remains 10/15. The new restart/approval fixes require live qualification; OpenCode evidence forwarding does not itself establish or fix its prior behavioral failure. - The positive provider case now expects three runs, including separate access approval. Its new results are distinct from the original invalid fixture. - The old Claude missing-result cause and rejected Codex field are unknown. These repairs do not retroactively explain or erase either failure. - Path rotation must preserve prompt, custom instruction, skill/content identity and protected provider settings; regression checks reject stale or changed content. No connection authorization, JSON schema, budget, cleanup, or final-result requirement is relaxed. Historical Everyday prompts and gateway setup remain unchanged. ## Model Used OpenAI Codex, GPT-6. The exact deployment variant and context window are not exposed in this session. Used repository inspection, code editing, test execution and retained-evidence analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
caf120105c |
test: prepare neutral native connection guidance evals (#15407)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need to discover connections, obtain consent, and continue from saved decisions. > - We want to reduce repeated instructions only when measured behavior supports the change. > - The existing decline tasks tell the model not to retry. One provider-decline check can pass without an explanation or an observed service counter. > - This PR adds neutral tasks and stricter saved-evidence checks before any connection instruction reduction. > - Production instructions remain unchanged. The new cells are configured, not live-qualified. ## Linked Issues or Issue Description Refs #15218. Refs #15389. **What existing behavior does this improve?** The Product E2E connection workflow evaluation and its instruction measurement provenance. **Current behavior** Some decline prompts supply the policy they intend to test. The provider-decline workflow does not require a saved post-decision explanation. Its old no-call check can use a missing fixture counter as zero. OpenCode has no connection cases in the original Everyday matrix. **Proposed behavior** Add an explicit-only suite with five connection stories on native Codex, ACPX Claude, and OpenCode. Require an explanation attributed by exact run ID after a saved decline. Observe the provider fixture counter. Preserve the original cases and grades. ## What Changed - Add fifteen configured cells with one attempt, twelve-minute deadlines, and verified 1,000-cent company and agent budget stops. - Remove procedure hints from the three new decline prompts. Keep a user-permitted explanation fallback and the existing positive controls. - Require saved decline state, one decision, unchanged connections, observed zero service calls where applicable, and a post-decision explanation from a successful run on the same task. - Add negative grader calibration and test the actual fixture budget payloads. Exclude the suite from default and generic selection. - Extend the existing full-catalog measurement source manifest with connection descriptions and schemas. Add an audit of fixed text, tool descriptions, returned instructions, and unqualified behavior. - Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and preserve its new Cursor suites. No production, credential, workflow, or lockfile change. ## Verification - Before rebase: Product E2E support passed 1,424 TypeScript tests and 128 Node checks. Six catalog measurement tests, repository typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery passed. - The full pre-rebase repository test run was stopped when master advanced. Its partial result is not a pass. - After rebase and the review correction: repository build/typecheck, Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128 Node checks, six measurement tests, and exact fifteen-cell discovery pass. The duplicate local full-suite run was stopped incomplete after about 20 minutes once complete CI passed; no local full-suite pass is claimed. - Review found that the initial grader read `runId` instead of public `createdByRunId`. A regression calibration reproduced both rejection of valid public comments and acceptance of the wrong alias. The fix uses the actual field and binds the evidence type to the shared `IssueComment` contract. A subsequent type-only import path correction passes Product E2E typecheck. - Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes [complete CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545): 51 successful checks and two intentional Storybook skips, plus separate Snyk success. Fresh Greptile review is 5/5 with the single review thread resolved and no new findings. The PR is clean and mergeable. - Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm test:e2e:runner -- --list --suite native-connection-guidance`. The measurement uses `PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json pnpm exec vitest run --project @paperclipai/server server/src/__tests__/native-procedure-measurement.test.ts`. Validation used pinned pnpm 9.15.4. - No paid provider campaign was started. There is no baseline/candidate behavior result for these new cells. - The audit records 654 UTF-8 bytes of fixed connection guidance. A clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41 supplied tools, 53,341 normalized bytes at start/resume, 50,949 at compact continuation, and a 48,195-byte authenticated OpenCode MCP catalog. These are byte counts, not tokens, bills, vendor-private prompt sizes, or savings from this PR. ## Risks - This is eval preparation. Passing support tests do not establish live model behavior or qualify an instruction reduction. - The explanation oracle checks attributed saved output. It does not prove cognition or arbitrary prose truthfulness. One saved interaction also does not prove the absence of repeated idempotent tool calls. - Successful new authentication and tool refresh, existing-connection agent grants, independent work while waiting, explicit retry after decline, and blocking when mandatory work remains still need separate coverage. - Notion setup decline does not execute a real Notion service. Positive service approval uses an already installed deterministic service; it does not qualify new connection creation. - The original historical failures remain unchanged. Future comparisons must freeze source, fixture, model, input, and grading controls and retain every actual attempt. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code execution. The exact serving model ID and context-window size were not exposed in this session; they are not inferred. No model provider was invoked by the eval suite in this PR preparation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a6306ba606 |
feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2d0c138122 |
Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also use assistants in Codex, Claude, and other MCP clients. > - The existing assistant connection can read work and create tasks or comments. > - It cannot edit tasks, exchange files, or manage normal agent and project settings. > - These operations must retain the person's permissions and Paperclip's execution rules. > - This pull request adds an explicit operation registry and separately consented configuration access. > - Assistants can manage work without receiving credentials or runner authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: assistant MCP, domain routes, consent UI, storage, and Product E2E. **Problem or motivation** A connected assistant cannot update tasks, maintain documents, attach files, or configure existing agents, projects, and skills. Users must leave the assistant for these routine actions. **Proposed solution** Add named tools and a restricted API registry to direct connections. Require separate configuration consent. Reuse domain routes and retry receipts. File uploads save the attachment when the byte transfer succeeds. **Alternatives considered** Arbitrary REST forwarding would expose administration and credential operations. Runner impersonation would bypass execution ownership. Separate upload completion calls add unnecessary client state. **Roadmap alignment** Checked ROADMAP.md and related MCP pull requests. This extends the human-authorized connection from #14933. It does not replace the runner or introduce agent impersonation. Companion Cloud routing and directory isolation: https://github.com/paperclipai/paperclip-cloud/pull/678. ## What Changed - Add task editing, finish/block, documents/revisions, deliverables, agent settings/instructions, projects/repositories, and skills/files. - Add an allowlisted API search/call registry with identical field restrictions, scopes, and retry identities. - Add unchecked configuration consent. Existing write grants retain their current authority. - Add hashed, expiring file transfer tickets and atomic upload receipts. No completion call is required. - Preserve company boundaries, human attribution, active native execution ownership, and execution review gates. - Add protocol/domain tests, consent stories, and eight paid Product E2E workflows. - Repair two CI fixture races: await cold route setup before assertions, and wait for asynchronously loaded connection copy. Both fixture suites pass (24 + 48 tests). ## Verification - Consent revision: one write-access checkbox controls requested work and configuration permissions in browser and device flows. All 16 consent tests, UI typecheck/build and token gates pass. Updated interactive stories cover default approval, opt-out and viewer restrictions. The paid browser helper uses the new exact label. Real GPT-5.4 Mini Product E2E passes 2/2 at `64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic retries, cleanup passed; $0.04149375 estimated assistant cost plus unpriced worker usage. Raw results, usage and source fingerprints are retained in the worktree. UI and Product E2E typechecks pass. - Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks pass; two optional Storybook checks skip. Greptile 5/5 on that head, no unresolved review threads. Final consent head `64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing and Greptile 5/5 with no unresolved threads. The unchanged Cursor sandbox test had one 10-second timeout, passed in local isolation, and passed its single CI rerun; the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37554106934). The existing chat retry-denial browser test had one visibility failure; its single rerun passes, and the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37542735691). - Full workspace `pnpm -r typecheck` and `pnpm build` pass at final runtime source `b2196fae1`. UI token gates pass. - 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests pass, including one-connection PostgreSQL OAuth and concurrent upload retries. - Paid Product E2E: all eight expanded cases qualified across Mini, Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell timed out before application startup. Final affected-case qualification passes 9/9 on all three models with grader v16, including the failed cell. Automatic retries disabled; failures, costs, source hashes and independent durable-state/file assertions are retained in [the verification record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md). - Actual Codex CLI, Claude Code and OpenCode clients completed local reads/mutations. Codex wrote a report, Claude updated it in a later conversation, and OpenCode uploaded/downloaded a file with matching SHA-256 and registered the attachment. Revoking the CLI grant rejects subsequent bridge initialization. - Butter staging is verified on final runtime `b2196fae1` ([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)). A fresh OpenCode workspace fetched the copied invitation, configured remote MCP, started OAuth and reached real consent with configuration unchecked. Invalid transfer tickets return 403 through Cloud. Human approval for the new persistent staging grant is pending; hosted task/file success is not yet claimed. The final transaction fix is deployed. - Full local `pnpm test:run` passed 15,614 general-server tests but stopped on two macOS timeouts. The heartbeat test passed in isolation; the existing 40,000-file Git stress fixture timed out again. Its Linux CI lane passes. Later local full-suite phases did not run after the timeout; this is not an all-green local full-suite claim. - Instructions and security limits are in `doc/public-mcp.md`; the saved plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`. ## Risks - This expands the experimental direct MCP surface. Explicit schemas and domain permissions must stay synchronized. - Migration 0311 adds transfer tickets and upload receipts. Expired orphan cleanup must not remove committed attachments. - Configuration requires a new consent request containing that scope; the single write-access choice controls it alongside work mutations. Refreshing an old grant does not add it. - The public directory keeps its original ten tools through the companion Cloud change. - Hosted consent/work proof remains the final delivery gate. The PR stays draft while approval of the new staging grant is pending; code checks and review are green. Merging is a separate action. ## Model Used OpenAI Codex (GPT-6, tool use and code execution). The exact serving model ID and context window are not exposed in this session. Paid evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001; claude-sonnet-4-6. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite macOS limitation disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ebd7bc0c9 |
test: isolate accounting fixtures and include all adapter suites (#14989)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
fe9d865250 |
fix(native): wake parents after revised child completions (#15372)
## Thinking Path > - Paperclip manages agents and the tasks they perform. > - A parent needs a new notification when a child finishes requested revisions. > - Native completion currently identifies that wake by the parent and child only. > - An already consumed notification can suppress the child's next completion. > - This PR binds ordinary child notifications to the committed status decision. > - A replay keeps the same identity, while a new completion can wake the parent. > - Live evaluation then exposed new child results being swallowed by an active parent. > - The repair preserves that input for a fresh turn and prevents premature Done from discarding it. ## Linked Issues or Issue Description Refs: #15218. The procedure experiments remain unshipped. This fixes native completion wake identity and durable delivery to a busy parent. Related open work: #10559 and #4507 address legacy wake deduplication, #11179 addresses delivery during an active parent run, and #13044 addresses watchdog signals. This PR changes native committed-decision keys and their delivery to an active parent, while preserving exact watchdog behavior. **What happened?** A native child completed, its parent consumed the notification, and feedback reopened the child. The second completion found the old completed wake and created no new notification. **Expected behavior** Each new ordinary child completion can notify the parent. Replaying the same committed completion must not add a notification. If the parent is already running, the new result must remain available for a fresh turn; an unread result must survive a parent Done claim. **Steps to reproduce** Complete a native child, consume its parent wake, reopen and complete that same child, then inspect the parent wakes. The regression fails on unchanged master because only one wake exists after two completions. **Paperclip version or commit** Baseline: `0fe47882cfcb12082035113c59ca96091c46ebfc`. **Deployment mode** Native runner with the standard server and PostgreSQL control plane. ## What Changed - Include the durable status-decision ID in ordinary child completion wake keys. - Defer new ordinary native child completions behind an active parent, carrying the committed decision identity and revised summary. - Keep a parent in progress while that result is still queued, claimed or deferred; recheck before the status commit. Use the existing continuation without reviving cancelled tasks. - Lock the parent before ordinary child status writes, making the notification/Done ordering explicit without upgrading an implicit foreign-key lock. - Cover exact sequential delivery, dispatcher replay, and queued/claimed/deferred versus consumed/current-run completion identities in database tests. - Apply it when the child is also a dependency and when it is only a child. - Preserve the stable key for exact `task_watchdog` origins to avoid repeated watchdog loops. - Extend real database conformance coverage for both relationships, ordinary and near-match origins, watchdogs, revised summaries, and replay. - Reconcile the working checklist with the merged guidance PRs and record the bounded next step. - Reuse the current composer helper for Everyday task creation: capture the returned task ID, preserve the exact prompt and chosen assignee/project, and cover the setup with paused-agent browser tests. The same setup correction is present in both comparison variants. ## Verification **Ready for review and merge at `e2fc0c8e3ddb84dd9bc045704c3d1ecb23ee3447`: both original live cases pass, zero new failures against the frozen baseline, all checks green, CLEAN/MERGEABLE and out of draft. Not merged.** | Profile | Original baseline | Initial candidate | Fixed candidate | | --- | --- | --- | --- | | native Codex / `gpt-5.6-sol` | PASS | FAIL | PASS | | ACPX Claude / `claude-sonnet-5` | FAIL | PASS | PASS | One new pass, one unchanged pass, zero new failures and no pending pairs against the original baseline. The earlier failed candidate is preserved; it was fixed and measured at a new source, not regraded or rerolled unchanged. ### Current source and live evidence - [Completed campaign 37523025407](https://github.com/paperclipai/paperclip/actions/runs/37523025407) and [public report with original screenshots](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37523025407-1/index.html). Report and screenshots return HTTP 200; screenshot bytes match original retained evidence. - Measured candidate `e2fc0c8e3ddb84dd9bc045704c3d1ecb23ee3447`; reused frozen baseline `7d700f43e93e89e4c196f8799b0aa3cef41d20da`. Trusted workflow `88ff98b83d15eaa6640ee1848df4f0ad9bc82af3` is distinct from measured sources; workflow blob `0600886144d3e22ea2e4a38329a79177882f3948`. Prompts, fixture, oracle, model/profile and input controls are unchanged. Both variants include the same composer setup repair. - Both original candidate grades PASS: all 31 checks per cell, parent and child Done, independently tested revised ZIPs, successful cleanup. Exactly one first attempt per profile on this source: 10 actual runs, no retries. Reusing nine baseline runs gives 19 matched runs; the baseline was not rerun. - Claude directly exercised the repair: its revised child completed while the parent was active, the parent's Done claim became InProgress with `native_child_completion_pending`, and a fresh parent turn then completed with the revised artifact. - Both persisted final provider comments refer to the same attachment whose parent-registration hash matches the revised ZIP tested by the original oracle; native final evidence also references that attachment. Codex provides a clickable download link. Claude describes the new `--max-length` behavior but uses a backticked attachment ID, without a clickable URL in the final prose. The artifact is registered on the parent and independently downloaded/tested; prose-link usability remains a presentation limit outside the original oracle. Artifact identity and observed execution do not prove cognitive review. - 68 focused scheduling/conformance/arbiter tests pass. Four regressions fail against the original production files while four controls pass. The database scheduler test proves one sequential continuation with the revised summary and replay deduplication. It uses a mock adapter and is separate from the live proof. - Full local `pnpm -r typecheck` and `pnpm build` pass. [Exact-head full CI 37522126627](https://github.com/paperclipai/paperclip/actions/runs/37522126627): 51 successful checks, two intentional skips, separate Snyk success. Fresh exact-head Greptile [5/5](https://github.com/paperclipai/paperclip/pull/15372#issuecomment-6022179172), no new actionable findings; all review threads resolved. - Ready-transition Contributor trust and Superagent Security Scan both pass. The security scan completed at 2026-10-06T20:26:53Z with zero annotations. Final total: 53 successful check runs, two intentional skips and separate Snyk success; aggregate SUCCESS, source unchanged, zero unresolved threads. - The explicit parent lock precedes child writes. Its controlled PostgreSQL ordering also passes on previous production through an implicit foreign-key lock; this is hardening, not a reproduced additional live failure. The four original failing regressions remain the before/after proof of the active-parent repair. ### Preserved failures and accounting - Initial candidate `a2ae2324ce7692e704fc43e99904b076d6246fee`: Codex PASS → FAIL, Claude FAIL → PASS. Equal totals concealed a new failure and did not qualify that source. Its Codex parent ended Blocked after the revised child result coalesced into the active parent; there was no retained later parent execution and the independent ZIP oracle was never reached. A later server backstop log did not prove recovery. - In the initial Claude cells, both parent finals referenced an earlier parent ZIP while the oracle tested the revised child's ZIP. The original candidate PASS did not establish latest-artifact delivery. Earlier parent ZIP bytes are absent, so different hashes alone do not prove missing functionality. Those original grades and content findings are unchanged. - Original baseline [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37509893406-1/index.html); original candidate Claude [artifact](https://github.com/paperclipai/paperclip/actions/runs/37509916344/artifacts/11434353106); failed candidate Codex [campaign](https://github.com/paperclipai/paperclip/actions/runs/37513070706). The original cohort contains 17 actual runs, no retries, successful cleanup. - Campaign 37509916344 was cancelled after its Codex job waited over 16 minutes without a runner or steps. Completed Claude evidence was preserved; only the missing Codex first trial was dispatched in 37513070706. The trusted workflow revision differs but workflow bytes are identical. The cancelled campaign has no published HTML. Earlier campaigns 37506546098 / 37506550087 retain four composer setup failures, zero agent runs; the identical fixture-only repair restored setup without changing the prompt or outcome oracle. - Intermediate `3c1cf6830ce9db4123b834ffb99c594c383a8986` [campaign 37518652522](https://github.com/paperclipai/paperclip/actions/runs/37518652522) was cancelled after a concurrency review finding. Both paid-cell steps started and both tasks were created. Only invocation policies survived; no grade, run inventory or cleanup receipt. Provider activity and charges are unknown: two incomplete attempts, not passes or zero-provider setup failures. - Cumulative accounting: **27 known actual runs** (17 original + 10 fixed-candidate), plus unknown activity in those two cancelled intermediate attempts. The 19-run matched comparison reuses nine baseline runs and is not additional execution. Reported LLM amounts are zero with original billing `complete=true`; actual charges are unknown and local/hosted runtime is unmetered. No free-run, speed or cost claim. - Original and new results pass the canonical result validator. Retained source, input, result/API/story/final-ledger/usage run identities and artifact hashes were audited. Each downloaded package omits the pre-upload-declared `playwright-output/.last-run.json`; primary result, API, final ledger, story, screenshots and reached ZIP oracles are retained. No full-package completeness claim. - The initial redundant local full test invocation was stopped after 2,492.5 seconds once that head's CI passed; completed groups recorded 23,045 passes and 87 skips. That local invocation remains incomplete. Current full CI is the repository-wide test evidence. ## Risks A revised completion can schedule another sequential parent run and its normal budget use. A parent Done claim remains non-terminal while a newer native child result awaits delivery. Company scope, governance, workspace-finalization, terminal cancellation and exact watchdog behavior remain enforced. This changes native completion authority and scheduling, so the durable identity, replay and concurrent-commit controls matter. The two live trials qualify the observed revised-child handoff, not broad task quality or causal/general equivalence. Claude's final prose still gives an attachment ID without a clickable URL, although the registered revised artifact passes the original download oracle. Completion before any new child result exists, removal of unfinished dependencies and broader instruction reduction remain separate questions. There is no schema or prompt change. ## Model Used OpenAI Codex, GPT-6 family, with repository inspection, shell execution and code editing. The exact serving model ID, reasoning setting and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ca0d26a99 |
fix(connections): recover missing personal AI credentials in chat (#15376)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e38d6d16b6 |
feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect accounts and choose an agent harness and model. > - The runtime change in #14970 supports custom providers on those connections. > - Normal setup must stay simple while advanced users can choose a compatible gateway. > - Shared connector rows and access controls keep these choices consistent. > - This pull request refines the agent setup UI and adds review stories and repeatable browser qualification. > - The qualification checks real tools and downloaded outputs, not only a successful run status. ## Linked Issues or Issue Description Refs #14970, #37, #13083, #14104, #14565, #12692. The core implementation in #14970 is merged. This branch incorporates its squash commit and targets `master`. Both PRs contain our implementation. #14016 is a reference only and is not a dependency. This PR has 96 changed files. ## What Changed - Complete model-provider connector presentation beside other connectors. Each row uses the existing Connect action and connection list. Tags are stored without category UI. The base PR includes the provider forms and routes. - Show persistent Subscription, API Key, and Advanced choices. Label Advanced as Custom Gateway. Reuse provider logos, connection lists, and permissions controls. Default access to the organization and all agents when permitted; keep narrowing controls under Advanced. - Keep Configure reachable before subscription sign-in, so users can select a supported environment when the default cannot sign in. Testing and saving still require a connection. Show the execution environment in Configure. Preserve the confirmed Connect choice. Editing a method, credential, saved account, or advanced choice requires that current choice to connect before testing or saving. Use matching model and thinking-effort dropdowns and retain connection icons in selected values. - Preserve the new harness model default when switching an existing OpenCode agent to Codex or Claude, and resolve user-selected model names with the effective harness. - Load popular OpenRouter models through the shared connection-model discovery path. Keep explicit model lists and manual model entry available. - Group onboarding, connection setup, agent runtime, management, recovery, and production-component stories under AI Connections / Provider routing. - Add an explicit-only provider-connections browser suite for managed local or existing local/staging targets. Use private browser profiles and credential handoffs. Support human-assisted subscription sign-in without sharing passwords or tokens in reports. - Verify persisted connection identity, runtime probes, tool execution, exact artifact bytes, completion, and context-dependent follow-up. Retain source/model provenance, cost bounds, closed error diagnostics, original failures, and cleanup evidence. - Add Gemini startup-model and skill-root fixes, Grok private-history detection, ACP filesystem regression fixtures, selected-workspace handling for local Hermes, and artifact-helper workspace fallback. - Keep managed Grok runtime homes disposable. Remove host-side transcript retention/restoration because private file modes do not isolate same-user agent processes. Ignore earlier development archives and use a fresh task handoff when history is unavailable. Verify the absence of restored transcripts with a separate same-user process. - Capture stopped-run diagnostics before deleting an attached-company fixture agent. Track creation and owned sign-in receipts; revoke only this attempt's accounts and never adopt a concurrent campaign's newly created account. Preserve failure signals and final status through cleanup. - Require the requested environment in the saved agent and every run, including follow-ups. Reject a forced incompatible target. Keep one cancellation state through startup, every cell, reporting, and teardown for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after interruption. Document qualification limits. ## Verification - Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes master `d9f600043`. The security fix in `a758fde31` passes full workspace typecheck, production build, and 119 connection/Grok regressions. The unchanged UI passes all 126 configuration/model-discovery tests and token gates. The final published-guide correction passes Grok adapter typecheck. Earlier head `eebd8225c` passed the complete deterministic runner suite (1,404 Vitest tests and 128 Node tests) and all CI jobs. Current-head CI run `37520147514` passed all 47 jobs, including the full sharded Vitest and browser matrix, production build, and canary dry run. All 55 checks completed: 53 successes and two expected skips. The current-head security scan passed, Greptile is 5/5, and no review threads remain open. - A separate same-user process reproduced reading a restored Grok transcript before the security fix. The regression now finds no transcript. Existing fresh-session fallback and ordinary session metadata behavior pass. - The final account-choice and cleanup fixes pass 85 setup tests and 26 qualification-harness tests. Regressions verify that editing a connection invalidates confirmation, Configure remains reachable before sign-in, diagnostics are captured before fixture deletion, and concurrent campaigns cannot adopt or revoke each other's accounts. UI and E2E typechecks pass. - The Storybook build and actual Chromium production-component stories passed during this change. Review the neighboring AI Connections / Provider routing stories, regular connector rows, three connection modes, model discovery, and the single execution-environment control in Configure. - Cancellation smoke verified authenticated cleanup before browser close for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during startup and reporting, missing-file ACP resource errors, and preserved permission denials. Both ACP runtime versions and 54 ACPX/Grok regressions passed. The deterministic connection-intent browser suite passed two tests. - Historical local qualification retained 43 passing API/gateway cells out of 46, with downloaded outputs and follow-up receipts. These attempts span earlier builds; they do not qualify this exact commit or staging. Subscription combinations, Gemini overloads, and the unresolved follow-up failure remain recorded rather than counted as passing. - Use `pnpm test:e2e:runner -- --list --suite provider-connections` to inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md` for credentials, target URL, sign-in assistance, budget, evidence, and cleanup. Paid live tests remain opt-in. ## Risks - The core implementation in #14970 is merged. This PR adds no database migration of its own. - Subscription login needs an interactive provider session. Dedicated accounts and staging qualification remain follow-up work; this PR does not certify every login combination for production. - Managed Grok transcript resume is deferred until provider history has an OS isolation or authorized broker solution. Follow-ups start fresh with Paperclip task context; earlier live Grok results do not qualify this behavior. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Live overloads and one unresolved follow-up timeout remain recorded. The stock CLI is unchanged, and those cases are not marked as passing. - Real-provider tests spend credits and use private credential/evidence directories. The launcher requires explicit selection and checks target ownership. It must not attach to a developer's database by accident. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local remain outside custom provider setup. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a7a2ed63a0 |
fix(ui): restore wheel and touch scrolling in task selectors (#15374)
Restore wheel and touch scrolling in nested task selectors and keep mobile picker controls reachable above the keyboard. Stabilize the affected server route test harnesses for Vitest 5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
22a3ea3414 |
Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Public MCP lets people use their organization from an external assistant. > - Operators need a visible control for this experimental access. > - Hosted users should select an organization once and then approve its permissions. > - This pull request adds the setting, invitation-first setup, and browser or device consent. > - Connections provides a copyable invitation with public instructions that grant no access. > - Users reach browser consent from their assistant and return to inspect or revoke access. ## Linked Issues or Issue Description Builds on merged foundation #14846. This PR now targets master. Related settings convention: #13905. **Current behavior** The preview uses an environment variable to enable MCP. Hosted consent repeats organization selection. Assistant access has no entry in Connections, so users must already know the endpoint and how to reach consent. **Proposed behavior** An administrator enables Settings → Experimental → Assistant connections (MCP). A hosted connection shows the selected organization and its icon, then asks for permissions. Requested write access starts checked when the user’s role permits it; the user can opt out before connecting. Direct instance connections show an organization picker with the first available organization selected. The selection stays fixed across refetches and still requires an explicit Connect action. Connections includes Assistant Connection (MCP). Its setup page explains the canonical endpoint, client configuration, browser authentication, and connected access. It connects as the current person and does not select or impersonate an agent. **Reason and benefit** Operators manage access with the other experiments. Users select one organization, and both the UI and server enforce that choice. **Breaking changes** The old enable variable has no effect. Preview operators must enable the setting once. Apply the additive consent-request migration before deploying the tenant, then deploy the compatible Cloud broker. Existing direct requests and grants keep their behavior. ## What Changed - Simplify OAuth and device consent: show the Paperclip logo beside “Connect {client} to Paperclip”, fall back to “your assistant”, and show the identifying origin plus its favicon below, with the callback URL also visible when different. Remove the hosted-organization creation action. Default to the first available organization without silently changing it on refetch; preserve company restrictions and write opt-outs. Keep the button row contained on narrow screens. - Make Copy invitation the primary action, using the shared animated AgentSetupPrompt and a collapsed manual setup section with icon-labeled line tabs. Remove redundant link actions, copy-status text, the extra first-prompt well and revocation explanation from the setup page. Serve shared version-aware HTML and Markdown instructions without private organization data. - Support guarded Client ID Metadata Documents alongside dynamic registration, and include authorization response issuer identification. - Add RFC 8628 device authorization with separately hashed codes, expiry, shared request quotas, persistent polling backoff and atomic redemption. Reuse human consent, role checks, scoped grants, audit and revocation. - Add CLI device login and a local stdio bridge. Store credentials separately with private permissions and serialize rotating refreshes. - Add device consent stories and five cold-start paid Product E2E cases with independent grant, configuration and durable-work assertions. - Add `enablePublicMcp` to the settings validator, normalizer, feature catalog, and toggle UI. Check it live for OAuth, tools, subscriptions, and event delivery. Keep connection management and revocation available while disabled. - Default the MCP origin to the existing auth public URL, with strict validation and an explicit override. - Persist the optional OAuth `company_id` restriction. Describe only that company and reject approval for any other company, even if the person belongs to both. Keep active-membership and role checks. - Show the Paperclip icon and a large organization icon during consent. Return the saved company logo through the company-scoped request response and reuse the standard fallback icon. Use the requested concise permission labels: “Read all of your Paperclip data” and “Allow write access and creating tasks as me”. Use concise permission copy, retain a compact client and callback-origin disclosure, and remove the footer link. - Default requested write access on for eligible roles. Preserve opt-out across organization changes and refetch, reset defaults for a new request, and submit read-only access when the request or role does not allow writes. Align the shared checkbox with its label. - Use organization wording in consent, management, settings, and walkthroughs. Keep the organization fixed for hosted requests and retain direct-instance choice. - Add an Assistant Connection (MCP) card to the Connectors catalog, a setup page in the app shell, and a return link from Experimental settings. Include Codex, Claude Code, OpenCode, and generic remote MCP instructions. - Read the live gate and canonical server URL through authenticated setup metadata. Show only the current person’s grants for the selected organization, refresh after consent, and support revocation. Surface catalog status failures with an explicit retry action; do not present them as an empty connection list. Opening setup grants no authority. - Start the eight guided chapters in Connections. Keep presenter notes and chapter controls around real product pages in the app shell. Explain the terminal, consent, delegation, retrieval, and revocation handoffs. Mark conversation examples as illustrative. Cover first use, client setup, connected, loading, and error states. Keep the existing consent and management stories. - Keep the paid-eval setup and browser helper aligned with the setting and consent button. ## Verification - Warm-standby integration fix `ec64ea05e`: public MCP ingress now follows the Cloud claim guard; MCP and discovery paths return 503 instead of SPA HTML while unclaimed. Event polling checks the in-memory claim before reading the persisted experimental setting. All 97 focused OAuth/Cloud tests and server typecheck pass, including new request and timer regressions for idle-before-claim and resume-after-claim behavior. Fresh review is 5/5 with no unresolved threads, and all security scans pass on this final head. All browser shards, typecheck, build, canary installation and other test groups passed on the first attempt. The unchanged Cursor sandbox default-command test timed out at 10 seconds; the exact test passed locally without edits in 587 ms. The single failed-job retry passed, with the original failure retained in workflow 37500711895. All 54 final-head checks pass on `ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs are intentionally skipped). - Final master integration `8457828fc`: merged foundation #14846 and current master, preserving the invitation changes and all 33 files from the two newer upstream changes. No migration renumbering was required. All 95 focused OAuth/Cloud integration tests, full recursive typecheck and token gates pass. All CI gates passed on that integration head; review identified the warm-standby issue fixed above. - Security-review fix `8c1d0b696`: commit shared global/per-source admission before outbound CIMD work, preserve failed-attempt receipts, and validate resource/scope before fetching. Added migration `0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive typecheck and production build pass. The security scanner passed that commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18 authorization requests sharing one proxy across two service instances use just two fetches. All 70 OAuth/metadata tests and server typecheck pass after that refinement. Final follow-up `8ebeae84c` reports admission-storage failures as retryable HTTP 503 instead of invalid client metadata. Its regression proves no outbound request before admission and successful retry after storage recovers. All 71 OAuth/metadata tests and server typecheck pass. Final-head security scanning passes; Greptile is 5/5 with no unresolved findings. CI passed all browser shards, typecheck, build, token gates and canary installation. One unchanged adapter-utils bridge test raced a response-file write (expected a JSON error, received the safe file-changed error). The exact test passed locally without edits. The single failed-job retry passed; the original failure is retained in workflow 37490609192. All 54 checks now pass on final head `8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final integration above now targets master. - Integration with current master: preserved the new Connections source filters and pagination, kept all eval suites, and regenerated the consent/device snapshots as migrations 0303/0304. All four MCP migration SQL hashes are unchanged from the staging versions. Full recursive typecheck and production build, 132 focused UI tests (including catalog filtering), 89 server authorization/settings tests, 26 migration tests, 120 eval calibration tests and token gates pass. Review follow-up `4820ce74c` also keeps active assistant grants in Installed, with pending/error recovery and revocation/company-isolation coverage. All 76 setup/catalog tests, UI typecheck and token gates pass after that fix. The unchanged signoff browser test timed out waiting for a heartbeat in CI at `4820ce74c`; the exact test passed locally without code changes, and the preceding CI head passed that shard. That same unchanged test failed at the reviewer stage in the next CI run. All five signoff tests passed three times locally (15/15), without test changes. All eight browser shards pass at final head `8ebeae84c`; no browser-test edits or failed-browser-job retries were needed. - Setup-page refinement at `9ab009178`: all 17 focused setup/consent tests pass, along with UI typecheck, production build, Storybook build and token gates. Browser exercised the shared prompt preview and client tab switching, and the updated InvitationCopied Storybook interaction checks its clipboard fixture. All final-head CI checks pass at `9ab009178`, with no unresolved review findings. Deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524. Verified the actual page, tab switching and line styling, removed actions/copy, and successful native copy/paste of the complete Butter invitation into a local-only test field. The existing Claude grant was left intact. - Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates pass. UI typecheck and production build passed again at `4e4d5e4d9`; Storybook build and eval-helper typecheck passed for `28101cf91`. Follow-ups let the primary button wrap on narrow screens, preserve a distinct callback URL, and use only bundled icons to avoid pre-consent requests to client-selected sites. Browser-verified the real consent component in desktop and 320px mobile stories, including default selection, write access and preserved opt-out. Updated E2E heading/default-selection helpers. All CI checks passed at `dc8e9fd11`, with review 5/5 and no unresolved threads. The Butter preview publication needed a retry because npm initially accepted the DB package before making it visible; the retry succeeded and `dc8e9fd11` deployed. Verified a fresh, unapproved native Codex CIMD request on Butter: default organization/write selection, known-client heading and icon, distinct callback origin, and removed creation action. No grant was approved for this UI check. Prior paid runs below retain their exact source provenance; this UI-only follow-up did not rerun paid qualification. - Source-pinned paid matrix at `2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign `local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config, unavailable host, denied consent and reconnect/later retrieval, with independent configuration/grant/task/run/document assertions. Original failures, transcripts, source fingerprints and billing remain retained. - Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`. Latest `0637b9f1c` shares that same guidance across HTML, Markdown and manual UI after review; generated Markdown is verified byte-identical to the paid-evaluated version. Shared build, server/UI typechecks, token gates and 63 auth/metadata tests passed again. Every CI gate passed at prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads. - Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval calibration tests, server/UI/eval typechecks, token gates and Storybook build passed. Full recursive typecheck and production build passed during implementation; CI also passed them at `2992ef271`. - Local full-suite limitations: a large-file Git streaming test times out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh MCP reruns passed, and the corresponding CI groups passed. No claim that the local full suite is green. - Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD consent; device CLI reaches verification/consent. New grants await human approval. Existing local OpenCode retrieved a saved result in a fresh conversation through its previously approved grant. - Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received the exact copied invitation, read public setup, configured its server and started PKCE consent. Its shell command timed out; background retry reached the client's own callback deadline while approval remained pending. Latest instructions cover that handoff. **No completed Butter read/delegation/result retrieval is claimed.** - Cloud companion https://github.com/paperclipai/paperclip-cloud/pull/672 passes checks/review and deployed. Anonymous setup and device-protocol routing verified. Core `2992ef271` deployed successfully and the actual Claude web flow now reaches consent. Its extra JWT-bearer metadata is filtered to implemented grants; unsupported token grants remain rejected. Final `0637b9f1c` deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300; live HTML and Markdown both contain the final guidance. The superseded instruction-only build was canceled before deployment. This is a core-only staging preview; private Cloud plugins are omitted. ChatGPT web is signed out, so browser connector use is unverified. - Screenshot gallery begins at Butter's dashboard and distinguishes real setup/pending consent from local reuse and fixtures. It records the timeout finding. New persistent access needs human confirmation before the remaining actual-client acceptance work. - Manual path: Connectors → Assistant Connection (MCP) → Copy invitation → paste into assistant → configure and start authorization → sign in and approve → verify `paperclip_connection` → delegate → retrieve the saved report later. - Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md` and `doc/public-mcp.md`. ## Risks - Apply additive, replay-safe migration `0304_curvy_shadow_king.sql` before using device authorization. The public setup link carries no credential. Device codes and tokens stay private; neither sharing instructions nor installing a plugin authorizes access. - Apply additive migration `0305_chubby_vin_gonzales.sql` before deploying the shared metadata admission gate. It retains at most 60 short-lived, hashed-source receipts per instance and rejects excess attempts with 429. - CIMD metadata fetching is a new external-input boundary. It requires HTTPS, exact client ID and redirect validation, bounded responses and guarded DNS/network access. Client names remain self-reported. - Device support is per-instance. The central Cloud broker retains its existing grant support. Host installation and tool reload capabilities vary by client; instructions describe manual settings and restart requirements. - Consent names the registered client in its heading and displays its identifying origin below. Known-origin icons are bundled; all other origins show a neutral site icon without contacting client-selected sites. Client names are self-reported; the callback origin is the recipient check. The Cloud chooser also displays the original client and receiving origin before tenant handoff. - A user who accepts the preselected write permission can create tasks and comments. Task creation and comments can start or wake agents and use execution budget; the consent label uses the concise wording explicitly requested by the maintainer. Scope requests, role checks, and the final Connect action still apply. - Migration `0303_supreme_garia.sql` adds one nullable UUID column with `IF NOT EXISTS`. Requests without a company restriction keep the direct-instance picker. The binding stays recorded if its company is deleted; consent then fails closed. - Deploy tenant support before the Cloud broker sends `company_id`. Unknown or inaccessible organizations must never fall back to a different company. - The setting defaults off. Disabling access does not cancel work already delegated. Existing tokens and unexpired subscriptions can resume when enabled again; revocation remains separate. - The catalog entry is visible for discovery while the feature is off. Setup instructions, OAuth, and tool execution remain gated. No access is granted by viewing the entry. - Assistant sign-in starts in the external client so it owns PKCE and callback state. Client command syntax can change and links to official setup documentation are included. - An authenticated instance and valid public URL are required. Hosting, paid execution, and store publication remain separate rollout steps. ## Model Used OpenAI GPT-6 in Codex, with tool use and code execution. The exact serving model version and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks pass; unrelated local full-suite timeouts are explicitly recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e34abee670 |
feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path > - Paperclip gives teams durable tasks, agent execution, budgets, and approvals. > - People also use assistants in Codex, Claude, and other MCP clients. > - Those assistants need a scoped connection that preserves the person’s permissions and attribution. > - Delegating a task must not turn the assistant into the assigned agent. > - This PR adds opt-in user OAuth, ten first-party tools, browser consent, and workflow packages. > - Paid product evals verify the resulting tasks, documents, attribution, retries, and access boundaries. > - The team keeps working after the assistant conversation ends. ## Linked Issues or Issue Description **Problem or motivation** A person cannot connect an external assistant to an existing team through browser consent and safely delegate durable work as themselves. **Proposed solution** Expose an opt-in `/mcp/paperclip` endpoint with individually described first-party operations. Bind every connection to a person, client, company, resource, and scopes. Reuse domain authorization and scheduling. Package shared team-review, delegation, and follow-up workflows for OpenAI/Codex and Claude. **Alternatives considered** Related PRs #9393 and #12549 cover earlier remote MCP and board-operator approaches. This change uses user OAuth and a bounded public catalog. It does not expose a generic executor, operator administration, static shared board credentials, or external agent execution. Registry listing work in #9851 is a separate distribution step. **Roadmap alignment** This maintainer-requested implementation extends the governed MCP gateway, activity attribution, durable work products, and hosted deployment direction in `ROADMAP.md`. It implements the first release of the saved design plan; external agent participation and granted third-party tools remain later releases. ## What Changed - Add MCP 2.0 discovery and task status/comment/document Events on the same authenticated endpoint. Persist subscriptions and delivery receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt callback material, recheck permissions/Cloud membership, and bound retries/expiry. Older MCP clients keep their existing tools. - Add discovery, dynamic client registration, S256 PKCE, resource validation, rotating refresh tokens, revocation, and company consent. Store credentials as hashes and recheck membership at execution. - Add tools for connection identity, agents/projects, task search/read/create, human comments, documents/deliverables, and pending-approval links. Preserve current domain permissions and scheduling. - Add durable mutation receipts across reconnects. Matching retries replay results; uncertain outcomes keep the same request ID and require inspection. - Add consent and connection-management pages, OAuth log redaction, shared plugin workflows, and separate OpenAI/Codex and Claude package outputs. - Add eight paid Product E2E cases across three models, independent durable-state grading, usage evidence, cleanup, and report integration. Add task-document guidance and regenerate the runner capability inventories. - Add migrations 0301 and 0302, the dated implementation plan, result notes, and direct-client setup instructions in `doc/public-mcp.md`. ## Verification - Merge integration `e180b1948`: resolved conflicts with current master, preserved both eval registries, regenerated capability catalogs, and regenerated migrations as 0301/0302 while keeping the original replay-safe SQL byte-identical. Local migration safety/snapshot tests (26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval catalog/grading tests (198) pass. Token and capability gates pass. Full recursive typecheck passed. Fresh Greptile review is 5/5 with no unresolved findings. CI is green on this exact head (55 successes, two intentional skips, one neutral result): one unchanged Cursor sandbox test timed out at 10 seconds, then passed locally in 856 ms. A single retry of that failed shard and the aggregate workflow passed. Merge remains blocked on the repository code-owner approval rule. Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55 successes, two intentional skips and one neutral result. [The earlier CI run](https://github.com/paperclipai/paperclip/actions/runs/36901592350) includes all test shards, browser tests, typecheck, build and canary dry run. Greptile was 5/5 on that commit with no unresolved review threads. GitHub still requires code-owner review under the repository merge rules; passing checks do not bypass that approval. Paid source fingerprints remain separate below and in the dated result note. - Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku 4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback, signature verification and report retrieval in a fresh conversation. A final Mini regression passes after the quota/status fixes. All evidence validates. Bounded tunnel startup retries occur before provider calls and remain visible; failed earlier attempts retain their original grades. - The earlier complete seven-case matrix passes **21/21**, with a separate **3/3** delegation regression. Two preceding matrices also passed 21/21 each. A complete 24-cell matrix including Events has not been run. [The dated results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain exact source fingerprints, failures, model IDs and partial costs. - Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass after merging master. Server typecheck passes after the final quota/status changes. Eval typecheck and all 892 eval-support tests pass. - All 33 real MCP/OAuth tests pass. The preceding combined MCP, redaction, private-address and DNS-rebinding run passed 129 tests; two later MCP regressions cover quota reuse and unchanged-status suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI then found a null checkout result in the existing concurrent-workspace path; logging now uses optional status access. All 12 closed-workspace tests and all 33 MCP tests pass after that correction. The exact-start event calibration exposed a timestamp gap; scanning now includes the subscription start, with all 33 MCP tests and server typecheck passing. These two narrow corrections follow the paid regression. - A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0 subscription/delivery, current membership loss, unsubscribe, legacy SDK tools, refresh and revocation. Its callback transport is a fixture with independent HMAC verification. The paid Events campaigns separately prove public HTTPS delivery. - Earlier component qualification passed UI 7,117 tests, CLI 502, shared 832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module boundaries, migration order and plugin regeneration passed. CI covers general/serialized suites, eight browser shards, runner checks, typecheck, build and canary dry run. - **Local full-suite limitation:** the earlier monolithic run was not clean. It encountered overlapping schema rebuilding, Mac database shared-memory limits and isolated CLI/fixture failures. Targeted reruns passed. The existing >32 MiB Git filename stress test still hit its 300-second Mac timeout. The additional serialized sweep stopped after 62 passing suites once CI passed. Original failures and partial logs remain; this PR does not claim a wholly green local monolithic run. - Local Codex CLI and Claude Code OAuth login and MCP SDK interoperability were verified. Public-store installation, actual ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted newcomer provisioning remain release gates. Enablement is moving to **Settings → Experimental → Assistant connections (MCP)** in the stacked follow-up [#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge both for the intended setup experience. This foundation branch alone still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set `PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and connect to `/mcp/paperclip`. Select a team and allow writes in browser consent. Configure an available agent and budget, then delegate and retrieve results later. For Events, rescan the deployed plugin catalog in ChatGPT Work Cloud; the host supplies its webhook credentials when the user asks to watch a task. See [the setup runbook](doc/public-mcp.md). ## Risks - Events are at-least-once and may arrive out of order. No replay cursor is advertised. Clients must refresh finite subscriptions, read current state and avoid comment feedback loops. Callback material uses the instance secrets master key; hosted subscriptions require the updated Cloud broker and are bounded to five minutes/the access proof expiry. - ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging subscription remain deployment gates. Local signed-webhook and paid model evidence does not claim those surfaces have been exercised. - Disabled by default. Merging adds schema and opt-in code; it does not deploy a public endpoint, publish a store listing, create a team, or start paid agents. - Migrations 0301 and 0302 are additive and idempotent. Their SQL is unchanged from the earlier preview numbers, so hash-aware upgrade reconciliation preserves prior staging applications. Normal instance upgrades must apply it before enabling MCP. - Task creation and comments can schedule paid agent work. Consent and tool descriptions disclose that effect. Revocation blocks future calls but does not undo delegated work. - Public deployments need edge rate limits and credential-safe logging. Internal dispatch is restricted to the closed catalog and carries a request-local verified actor. - Hosted onboarding requires the companion Cloud broker, encryption-key configuration, and tenant rollout. Self-hosted direct connections can use this PR alone. - Store acceptance and agent-mode participation are not claimed. Checked-in plugin endpoints are development defaults; rebuild packages for a real deployment before installation. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A more specific serving version and context-window size were not exposed by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`, `claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted/component checks; full local-run limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
16b7db35ff |
Shorten planning skills and measure task decomposition (#15296)
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3290d97417 |
Scope the release smoke opening question to its card (#15339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The release lane checks onboarding before it promotes a nightly candidate. > - The first task shows an unanswered-question summary and an open question card. > - Both display the same prompt, so the smoke test's page-wide text locator fails strict matching. > - This pull request scopes that assertion to the open interaction card. > - The smoke still requires the actual prompt and verifies that onboarding does not start an agent run. ## Linked Issues or Issue Description **What happened?** [Scheduled release run 37441718130](https://github.com/paperclipai/paperclip/actions/runs/37441718130) failed `smoke_nightly / smoke` on both attempts. The page-wide `getByText("What would you like to do?")` matched two elements: the timeline summary and the open question card. This skipped `publish_nightly`. **Expected behavior** The test confirms that the seeded task's open question card displays the exact prompt. A summary row alone must not satisfy that check. **Steps to reproduce** Run the Docker onboarding smoke against published candidate `2026.1006.0-canary.11`, then run the release-smoke Playwright spec. The old assertion fails after the greeting appears. **Paperclip version or commit** Release workflow commit `f858207161ba29c01c82f4674aef83d91b74480f`; published candidate `2026.1006.0-canary.11`. The assertion is unchanged on the current master base `9b3fe260bac576d622ffdfefb923db73cc5273f7`. **Deployment mode** Docker smoke harness, authenticated/private, with its existing mock provider. Related: #13166 added the first-task chat assertion. I also reviewed open #12316, which fixes a separate bootstrap race and does not change this selector. Searches found no duplicate selector fix or public issue. ## What Changed - Scope the exact opening-prompt text to `task-chat-interaction`. - Explain why the timeline summary cannot satisfy the assertion. - Preserve the rest of the smoke flow, including the 15-second check for no agent runs. ## Verification - Local Docker harness ran the exact failed published candidate, `2026.1006.0-canary.11`, with the existing mock provider. - The old spec reproduced the same two-element strict-mode error in Chrome. - The fixed spec passed the full authenticated onboarding flow and the 15-second no-run check: 1/1, 32.6 seconds. The executed spec copy was byte-identical to the changed repository file; only artifact output paths and the local port were overridden. - Independent Chrome checks: the scoped locator passes with both summary and card present. Summary-only, wrong prompt, hidden prompt, and duplicate active prompts each fail as intended. - `pnpm typecheck` and `pnpm build` passed on Node 24.21.0. Canonical Linux CI passed all 55 checks (53 success, 2 intentional Storybook skips), including the full aggregate tests, build, typecheck, all eight E2E shards, and post-ready security scan. - Greptile scored the exact head `5a30e667396600a052a4041c819652c19b8c68ff` at 5/5 with no actionable issues. Independent review found no issues; there are no review threads. - `git diff --check` passed. The diff and PR text were scanned for secrets and private identifiers. ## Risks Low risk: one assertion in a release smoke test changes. A missing or hidden prompt still fails, and multiple matching prompts in interaction cards still fail strict mode. This does not change the product, release selection, or publishing. The scheduled release lane still needs its next normal successful run to prove nightly recovery. ## Model Used OpenAI GPT-6 through Codex, with reasoning, shell tools, and an independent Codex review. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e0b63e5a5 |
feat(connections): add experimental task-pinned AI routing (#14967)
## Thinking Path > - Paperclip manages AI agents and their work. > - AI Connections separate account access from models and harnesses. > - A pool must act as one connection while retaining each task’s account. > - Core must enforce member access and preserve session and recovery rules. > - A plugin supplies rotation policy without receiving credentials. > - This change adds durable routing and native connector setup and management. ## Linked Issues or Issue Description **Subsystem affected** AI Connections, Connectors, plugins, run dispatch, and session compatibility. **Problem or motivation** Operators need to rotate new tasks across saved accounts while each task keeps its account and session. Pool setup must fit the existing connector catalog and account workflow. **Proposed solution** Add an experimental router binding, a capability-gated plugin hook, and transactional task pins. Plugins declare native pooled connectors through `aiConnectionRouter`. Core hosts the existing-account picker, ordering step, and account settings. Related usage contract: #14936. Companion private plugin: https://github.com/paperclipai/paperclip-cloud/pull/643. **Roadmap alignment** This extends Apps and AI Connections. Core supplies generic enforcement and native connector UI; the private plugin owns rotation and quota policy. The prior duplicate search found no matching router implementation. ## What Changed - Add a router binding without changing existing concrete bindings. Keep the instance flag and new pools disabled by default. Require manual operator configuration. Show no routing toggle in Experimental settings on either open-source or Cloud installs, even after routing is enabled. - Persist company-scoped pools, one shared cursor per pool, and pins keyed by company, pool, agent, and task. Commit pins and cursor advances together with revision checks and bounded retries. Persist run-ID affinity before allocation. - Pass only authorized metadata and normalized usage to plugins. Core retains credential handling, member access checks, runtime qualification, and recovery evidence. Probe outside locks with a shared 15-second budget and freshness cache. - Resolve routing before credential preparation and backend selection. Preserve pins through turns, session resets, removed members, and quota waits. Retain admitted recovery after disable or uninstall. - Separate credential session epochs from token generations. Verified refresh preserves the epoch; reconnect and manual replacement change it. Include the credential slot ID in session and usage-cache identity, so reconnecting an indexed legacy account invalidates its old session even when both epochs are zero. - Validate pool member installations before accepting saved-agent bindings and recheck compatibility when the harness changes. Install only authorized members in the new-agent transaction and record their IDs in local activity. Pool membership cannot install a restricted shared connection. - Preserve pool bindings when agents hire teammates through either creation API or native caller runtime inheritance. Block stale manager credential references; retain explicit child authentication precedence and reject incompatible inherited pools. - Add native connector registration through plugin metadata. Reuse the Connectors catalog, setup header, account header, sidebar, dialogs, and usage display. Setup selects and orders saved connections. Advanced settings hold usage rules and member runtime defaults. New-account setup opens in another tab. - Use revision-checked pool archival from the Connectors catalog and account page. Keep task pins, cursors, recovery evidence, and underlying connections. Reject ordinary connection updates or removals that bypass pool revisions. - Add pool selectors, composer models, override notes, quota status, run details, activity records, and local run-log records. Keep session-adoption copy minimal. - Show **Used by** below the pool connections. List current company agents with shared avatars and profile links. Include paused agents; exclude terminated agents and agents using another pool. - Add Core stories for the generic connector workflow and runtime surfaces. Cloud stories reuse these production routes and tokens through a preview-only alias. ## Verification - Final head `73cb953bca30ed83e4505dd820edd9b5edffd28b`: full workspace `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass locally. - All 422 focused connector/settings/shared-contract/migration tests and all 156 database-backed AI connection, hiring, reconnect, and durable-routing cases pass (69 hiring cases rerun after the final auth-precedence fix). The merged shared contract retains connection instructions and pool metadata. The pool migration is generated at sequence 0299 after the latest upstream migrations; this PR makes no lockfile changes. - All four full-app Playwright tests pass on the final head after a cold restart and migration, against the installed private plugin and isolated database, with no route or pool-API mocks. They cover hidden routing controls after manual opt-in, native pool creation, ordering, membership edits, rename, paused defaults, enabling/save/refresh persistence, stale edits, cancellation/removal, preserved underlying accounts, unavailable routers, and Used by avatars and profile links. Exact command: `PAPERCLIP_CONNECTION_POOL_E2E=1 AI_CONNECTIONS_TEST_COMPANY_ID=a37b9625-5ecf-4e29-8081-04df3d6e7d6f AI_CONNECTIONS_TEST_URL=http://127.0.0.1:3108 pnpm exec playwright test --config tests/ai-connections-app/playwright.config.ts connection-pools.spec.ts`. - [Native setup, ordering, and management screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-6006976278) address the review follow-up. [Earlier selector, quota, and run-detail screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-5971537316) show the runtime surfaces. Core previews: `pnpm --filter @paperclipai/ui storybook`, then **Connectors / Pool host** or **AI Connections / Connection pools**. Cloud owns its host-backed plugin stories; both repositories’ Operator Setup Required story assertions pass. - Live acceptance used OpenAI/Codex and Anthropic/Claude ACPX, resumed both exact sessions after restart, preserved pinned accounts through explicit reset and controlled quota deferral/recovery, and committed only two allocations across fourteen runs. A later UI-created task test again rotated OpenAI then Anthropic and resumed OpenAI through follow-up/restart/quota recovery. That later Anthropic execution was blocked by its saved OAuth token expiring (provider 401). No live usage probes ran. - The full local `pnpm test:run` was attempted earlier and did not complete because of macOS embedded PostgreSQL bootstrap/shared-memory failures and the 40,000-file Git fixture timeout. The focused database suites above now pass; full-suite verification is provided by the split CI lanes. The preceding CI run had one runtime readiness timeout; it passes locally both alone and inside the larger runtime suite. That larger local suite also encountered an embedded PostgreSQL setup failure and two macOS temporary-path alias assertions; those two assertions pass with canonical TMPDIR=/private/tmp. All final-head CI checks are terminal green, including full general/serialized server suites, Runner checks, browser E2E shards, canary verification, build, and typecheck. Greptile is 5/5 on that exact head with no unresolved threads. ## Risks - The migration adds routing tables and a credential epoch column. Install the private plugin only with the compatible Core contract. - Routing and each pool require opt-in. Production distribution and fleet defaults remain unchanged. - Unknown usage stays eligible. Known pinned exhaustion waits; revoked access requires operator repair. - Legacy adapters require compatible members. Runner model and effort overrides remain limited by qualified backend support. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and browser testing. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes:` / `Refs:` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted suites; full-suite limitations are reported above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
72ff3a9f27 |
Measure native tool context and expand bounded workflow evals (#15218)
## Thinking Path > - Paperclip manages persistent agents and their assigned work. > - Agents receive both fixed instructions and tool definitions. > - Moving a procedure into a tool description still adds model context. > - We need to measure the complete delivered catalog and test real outcomes. > - The tested reductions saved little space and introduced failing outcomes. > - This PR keeps measurement, bounded eval coverage and the original evidence. > - Production prompts, tools and runtime behavior stay unchanged. ## Linked Issues or Issue Description Refs: #15151, #14961, #14948, #14985. This adds the measurement and eval coverage needed to assess further native instruction changes. The attempted hiring and dependency reduction failed qualification and is excluded from the final diff. ## What Changed - Measure the actual standard-mode tool authority, including all 39 tools and their input schemas. Capture scripted native start, resume and continuation payloads and the OpenCode MCP declaration list. - Add OpenCode to the two explicit-only local hiring/reuse and delegation/feedback stories. Preserve the original task requests and independent oracles. - Apply one attempt per selected story and explicit company and lead-agent budget stops. - Select managed hiring credentials from the requested profile, including OpenRouter. - Run the existing Node test files under Node instead of collecting them as Vitest suites. - Retain sanitized comparison reports, original failed grades, source hashes and evidence gaps. ## Verification Final source: `2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8`. Local verification passes: - Full repository `pnpm -r typecheck` and `pnpm build`. - Eval-support unit tests: 1,252 Vitest checks and 128 Node checks. - Eval TypeScript check and six final-source measurement tests. All 36 normalized components across nine scripted deliveries and the OpenCode MCP catalog match the baseline exactly; the complete standing projection is 49,200 bytes. Fresh review of this exact head is [5/5 with no remaining findings](https://github.com/paperclipai/paperclip/pull/15218#issuecomment-5996723170); all three review threads are resolved. [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/37354539126) passes all 47 jobs, including repository typecheck, build, tests, runner checks, browser shards and the canary check. One initial annotation-test timeout is retained in attempt 1; its seven-test suite passed locally, and the affected CI lane passed on one targeted retry. No source or paid eval rerun was needed. Both ready-transition security scans passed with zero annotations. The PR is out of draft and conflict-free. GitHub still requires code-owner approval for the `package.json` test-script change; its requested reviewers are already set. A byte-for-byte comparison against master context `a65ca0950834a85bb93bcc4b4042ecacdebfef53` confirms no production changes under `packages/` or production `server/` paths. The only server addition is a measurement test. No new paid rerun is needed to compare unchanged production bytes. This does not claim that existing product defects have been fixed. The rejected corrected experiment had baseline **5 PASS / 1 FAIL** and candidate **3 PASS / 3 FAIL**, including **two newly failing pairs**. The later readiness experiment had candidate **3 PASS / 3 FAIL** and baseline **3 PASS / 2 FAIL / one setup cell without a behavioral grade**. Its five comparable pairs had two new failures, two new passes and one unchanged pass. The missing baseline Codex cell never reached its provider step because Docker setup timed out. The observed missing behaviors include parent continuation, waiting for the latest child revision, revised ZIP delivery and OpenCode credential persistence. Those failures remain failures. Source review and passing CI do not regrade them. The original reduction saved 460 bytes; its first repair saved only 125 bytes, and the larger unqualified runtime repair increased the full projection. None of those production changes is shipped here. Read the [report](https://github.com/paperclipai/paperclip/blob/2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8/doc/plans/2026-10-05-native-procedure-guidance.md) and its linked sanitized receipts for exact sources, original campaign links, pair-level results and evidence limitations. ## Risks - Full-catalog bytes are not model tokens, invoices, private vendor prompts, lazy-loading behavior or truncation proof. - The two stories are explicit-only and do not prove general coding quality or arbitrary resume behavior. Single trials do not establish causation or performance trends. - The retained OpenCode candidate credential guard failed. Cleanup removed the original provider database, so the precise persistence mechanism remains unknown. The guard is unchanged. - ACPX provider-execution IDs lack proven host-call mapping. No extra-work or feedback-consumption claim is inferred by matching names, order or counts. - PostgreSQL cannot start locally while the host's shared-memory slots are exhausted. Hosted CI must supply the full database checks; the full local database suite is not claimed green. ## Model Used OpenAI Codex, based on GPT-6, with repository tools and code execution. The exact deployment model ID and context-window limit are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2c43b39167 |
refactor(ui): share task composer and project/worktree controls (#15091)
Use the shared composer for new tasks, including project and worktree selection, rich mentions and slash commands, and remembered task settings. Save project preferences after successful task updates. Show known agent models by name and use Default for unknown runtime defaults. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b43073d11f |
feat(connections): sync and group accounts managed by aggregators (#15254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external tools. > - Aggregator gateways can expose accounts that users already connected upstream. > - The Apps catalog did not show those accounts or their current provider status. > - Separate cards and setup tasks also made account ownership unclear. > - This pull request discovers upstream accounts and groups them under one app card. > - Users can find connected apps while each provider keeps control of its accounts. ## Linked Issues or Issue Description **Subsystem affected** Connections across the database, shared contracts, server, and board UI. **Problem or motivation** Users cannot see which apps are connected through a saved aggregator gateway. Native and upstream accounts need one app card. Discovery must preserve company, user, gateway, and credential boundaries. **Proposed solution** Sync account metadata from Composio, Arcade, and supported Executor gateways. Keep upstream account management in each provider. Use source chips and search to browse the catalog. Preserve native setup and the gateway's existing access policy. **Alternatives considered** Creating a local executable connection for each upstream account would duplicate authorization state. Using an agent task for routine Composio setup would add an unnecessary step. The board now calls the saved gateway directly for that setup. **Roadmap alignment** This extends the shipped Connected Apps and MCP Tool Gateway features in ROADMAP.md. The duplicate search found no open PR for managed account discovery. Related work: Refs #13755, Refs #13941, Refs #14725, Refs #13855. Open PR #12906 covers adjacent toolkit routing work. ## What Changed - Add provider-neutral discovery, sync, and refresh APIs. Preserve the Composio API paths. - Cache observations by company, saved gateway, viewing user, and credential version. Retain stale observations after failed or incomplete scans. - Add optional Arcade account sync credentials in the vault. Discover Executor accounts through its supported inventory interface. - Group native and upstream accounts in one app card. Imported account menus open their provider. Gateway menus own refresh and sync setup. - Add Paperclip, Composio, Arcade, Installed, and All chips. Show 50 catalog entries per page. Keep connected accounts above discovery. Keep explicit provider searches scoped. - Simplify Composio app setup and refresh its connected app list on the gateway Permissions page. - Add a compact agent access card and task creation defaults for connection setup. Preserve explicit blocks and approval policies. - Add two replay-safe migrations, service and UI tests, Storybook journeys, and acceptance stories. ## Verification - Passed the repository typecheck, full build, token gates, and migration ordering check. - Passed the focused provider adapter, connection interaction, and catalog tests after rebasing onto master. - Passed all nine database sync and migration replay tests using a disposable database on the test-drive PostgreSQL cluster. Removed that database after the run. - Verified Arcade cursor pagination against its official Go SDK and passed all eight adapter tests, including short and incomplete pages. - Passed all 45 interaction tests after making the exact requested tools and their Allowed/Ask first permissions visible before granting access. Verified the compact card in Storybook. - Passed the complete UI suite on the final code: 683 files and 7,432 tests, including the corrected Composio destination assertions. Passed 130 focused tests for the UUID, management-link, and health-status corrections. - Passed 22 Composio setup/sync tests, 23 connection-intent service tests, and the connection migration test in separate disposable databases. Database startup alone was substituted; the suites exercised their real SQL and services. - Passed all 10 OpenAPI route checks and the full-stack connection-intent browser test, including scoped consent, agent continuation, and task completion. - The local full runner encountered embedded PostgreSQL startup failures on this loaded macOS host. The earlier in-flight run also held the pre-fix Arcade transform; a fresh run of the final provider suite passes. The final-head CI is queued during GitHub’s active Actions incident: https://www.githubstatus.com/. The previous run also lost several runners simultaneously; its real catalog assertion failures are fixed and the fresh complete UI suite passes. - Tested the real test-drive server in the embedded browser with a live Composio gateway. Detected Airtable and Circleback. Verified refresh progress, account rows, source chips, search scope, and 50-entry pagination. - Arcade and Executor coverage uses provider fixtures. Live credentials were unavailable. - Storybook builds successfully and includes grouped native/provider accounts, stale and unavailable discovery, optional Arcade setup, and mobile states. The acceptance document records the simulated and live coverage separately. - Greptile reviewed final commit `217b024c27b5933e773ce9419c4e92b1032042c6` at 5/5. All six review threads are resolved, security scans pass, and the PR has no merge conflicts. The outstanding remote checks are `ci / Select trusted runner` and `review`, queued by GitHub. They need to complete before merge. ## Risks - Provider response changes can break inventory discovery. Failed scans retain observations and show stale status. - Composio scans only the supported catalog and can take time. Large inventories run in the background with progress and a bounded lease. - Arcade requires a project API key and user ID when the gateway cannot supply them. This key is used only for discovery. - Executor discovery depends on the server's exposed inventory tools. Unsupported servers report unavailable discovery. - Cached account rows do not grant access or create executable connections. Gateway policies still govern tool use. Account deletion and per-app authorization remain upstream. - The migrations add tables and one nullable column. Replay preserves existing rows and company-scoped foreign keys. ## Model Used OpenAI Codex, based on GPT-6. The session does not expose a more specific serving model ID or context limit. Used reasoning, repository tools, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7498705642 |
fix(ui): stabilize mobile task reading and document navigation (#15228)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People read tasks and send instructions from phones as well as desktop browsers. > - The mobile footer let page text show through its labels, and small text fields made Safari zoom on focus. > - Small scroll changes made the footer switch direction, and its changing page padding moved the conversation. > - Task pages also showed comments before question cards and run history arrived, so the composer and reading position moved again. > - Document links also used native navigation, which reset the reading position or reloaded a task through its UUID URL. Desktop tabs were crowded in the mobile drawer. > - This pull request keeps navigation steady, opens documents in the mounted task, and gives mobile readers a full-height panel with a vertical tab selector. > - The benefit is a stable task view and smoother scrolling on mobile. ## Linked Issues or Issue Description **What happened?** The mobile footer was translucent. Safari zoomed when a person focused a small text field. The footer switched abruptly while scrolling. A large task could show saved comments, then move the page again when a question card or run history arrived. In a local test with delayed responses, a late question moved the mobile composer by about 374 pixels. Opening a plan from the feed could reset the view or reload the task through a UUID link. The mobile document drawer left part of the feed exposed above small desktop tab controls. **Expected behavior** The footer has an opaque surface and moves smoothly after deliberate scrolling. Text fields do not cause automatic focus zoom. A task shows its initial conversation and composer together at the final scroll position. Background refreshes keep the existing conversation visible. Document links open in the mounted task with its existing cache and reading position. Mobile documents fill the viewport, show a clear close button, and offer a vertical list of open tabs. **Steps to reproduce** 1. Open a task with many long comments and a pending question in iOS Safari. 2. Delay its interactions, activity, and runs responses by different amounts. 3. Reload the page and watch the conversation and composer move as each response arrives. 4. Scroll down and back up, including small direction changes and edge bounce. 5. Focus the task composer, search field, and new-task title and description. 6. Open a plan or another task document from the feed, including a link that uses the task UUID. Close the panel and check the reading position. 7. Open several documents on a phone. Switch tabs and close both active and inactive tabs. **Paperclip version or commit** Developed from `1c07b5903` and rebased onto `59015846a`. **Deployment mode** Built from source. Tested in an isolated local test drive with iOS 26.5 Simulator Safari and Chrome. A temporary local proxy delayed independent responses for the layout test. Related work found in the duplicate search: - Refs #14727. It made saved replies appear before supporting history. This PR keeps its parallel requests and narrows the tradeoff in favor of a stable first layout. - Refs #14667. This open PR takes a different approach with per-run placeholders and retries. This PR fixes the observed question/composer movement and mobile navigation behavior. - Refs #13095 and #13597. These earlier fixes added task scroll anchors and skipped transcript waits for scheduled retries. - Refs #6550. Earlier mobile board polish. - Refs #9467. This related open PR changes list and generic tab reflow. The task-pane selector uses a separate component. ## What Changed - Give the mobile footer an opaque semantic surface. - Set a base-size floor for editable text on touch devices to prevent Safari focus zoom. Preserve larger title text. - Share mobile scroll tracking between both layouts. Accumulate scroll distance, ignore edge bounce and changed document bounds, and update once per frame. - Use shared motion tokens for the footer and composer. Keep page padding stable and honor reduced motion. - Wait for the initial question cards, attachments, work products, activity, runtime selection, plan, and relevant transcript history before the first reveal. Skip scheduled retries and older runs outside the initial comment window. - Bound the first reveal to 15 seconds. A stalled supporting request leaves saved conversation and the composer accessible with an explicit loading notice. - Keep concealed mobile history from stretching the document. Keep the composer mounted but concealed until the same reveal. Keep both visible during later refreshes. - Route first and repeated same-task document clicks in place. Recognize UUID and identifier links. Preserve the thread history entry and feed position. Keep modifier clicks, downloads, external links, and classic document behavior. - Give the mobile task panel the full viewport and safe-area padding. Use a visible X and 44-pixel touch controls. Replace the horizontal tab strip with a vertical selector that wraps titles and supports keyboard focus. - Add four interactive Storybook states for a few tabs, long names, many tabs, and the last tab. Reuse the production selector and tab controller. - Add navigation and tab regressions, update first-reveal regressions, and document the behavior in `DESIGN.md`. ## Verification - 392 tests passed across the seven focused task-loading, scroll, mobile-navigation, layout, and composer suites. After review fixes, all 339 tests across the four affected suites passed, including stalled-loading fallback on mobile and desktop and the motion-token catalog. - All 442 focused document, tab, task-thread, and scroll tests pass. The final click-propagation cleanup also passes all 136 task-detail tests. UI typecheck, UI production build, and `pnpm check:token-gates` passed. - All four cases in `artifact-tab-arrival.spec.ts` and `text-attachment-tabs.spec.ts` pass locally, covering desktop and mobile selection, composer focus, document rendering, and downloads of the original bytes. - `pnpm --filter @paperclipai/ui build-storybook` passed. Open the mobile tab stories under `Prototypes/Task detail/Mobile tabs`. - A local diagnostic proxy measured cached plan content at about 250 ms after the first click. The HTML load count and task request count did not change. Feed scroll stayed at the same position. First, repeated, and UUID document links were tested at phone and desktop widths. Task-reference links close their preview before the document reader opens. - In Chrome at desktop and phone widths, the delayed-response task showed one complete reveal. The late question no longer moved an already visible composer. - In iOS Simulator Safari, verified the large-task reload, opaque footer, navigation hide/reveal, and search/new-task/composer focus without automatic zoom. - Full workspace typecheck and build passed. The updated UI also passes typecheck, production build, and token gates. - All 54 checks pass on the final commit `bf5c9914e61833e7cc8794a69d72cf8c7057b952` (two additional checks are intentionally skipped). Greptile reviewed that commit at 5/5, and all review threads are resolved. - Full local `pnpm test:run` was attempted but stopped after server-fixture failures. Embedded PostgreSQL startup failure reproduced in an isolated native-interaction fixture after five startup attempts. The broad run also reported a rapid Slack callback ordering test failure. These server paths are unchanged by this PR, and their CI shards pass on the latest head. The full local suite is not claimed as passing. - A localhost proxy stalled the activity response for 30 seconds. Chrome revealed the available conversation after the 15-second deadline at both desktop and phone widths, kept the composer accessible, and cleared the loading notice when the response arrived. ## Risks Slow initial history requests can delay the first conversation reveal by up to 15 seconds. If that deadline expires, late data can change the available conversation while a loading notice remains visible. The reveal waits only for runs in the initial comment window, and later refreshes do not conceal an existing conversation. The larger editable text can change line wrapping on phones. Mobile navigation and composer motion use shared tokens and respect reduced-motion settings. Mobile tab selection changes the control layout. Document links retain URL history while sharing the task reading position; other tasks and external links keep their normal navigation behavior. I checked `ROADMAP.md`. This is a fix for existing UI behavior. ## Model Used OpenAI GPT-6 in Codex. The runtime does not expose a more specific model ID or context-window size. The agent used reasoning, code editing, terminal tools, and Chrome and iOS Simulator testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
59015846ae |
fix(chat): keep dismissed task questions in the feed (#15229)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask users questions in task chats and Agent Chat. > - Agent Chat keeps unanswered questions as compact entries in the feed. > - Regular task chats still show a pending composer badge after dismissal. > - A page reload can also open a dismissed question form again. > - This pull request applies the same feed behavior to both chat types and saves dismissal per person and task. > - Users can continue the chat and return to the original question later. ## Linked Issues or Issue Description **What happened?** Dismissing a question in a regular task chat leaves a pending composer badge. The question can return after a page reload. The earlier feed behavior only applied to Agent Chat. **Expected behavior** A dismissed question stays in the feed. Its form and pending badge leave the composer. Reload preserves dismissal. Opening the feed entry restores the original question and draft answer. **Steps to reproduce** 1. Open a regular task chat with a pending question. 2. Select an option, then dismiss the form. 3. Reload the page. Check that the composer stays clear. 4. Open the question in the feed. Check that the draft is restored. 5. Submit the answer. Check that the answered receipt appears. **Paperclip version or commit** Reproduced on master at `a386a599983519eb1d399f8b770bfccdb2a74762`. **Deployment mode** Browser UI in local and authenticated instances. This change does not depend on the agent adapter. Related work: Refs #14613, which added the Agent Chat feed behavior. Refs #9141, which validates real answers on the server. Refs #11434, which tracks comment-driven changes to interaction state. This PR changes question presentation only. ## What Changed - Show compact unanswered question entries in regular task chats. - Exclude durable questions from composer pending counts and navigation. - Save dismissal in local storage per person and task. Merge the latest saved IDs so dismissals from another tab survive reload. A new question can still open its form. - Keep the original question pending and answerable. Keep approval and permission controls. - Run the question-history regressions in both chat modes. Add reload, new-question, user/task scope, and stale-tab coverage. - Share the real-component Storybook fixture. Add a regular task test drive and an interactive dismissal/answer scenario. - Update the planning-mode browser test to check a dismissed question in the feed after reload and on mobile. - Update the design rules, implementation spec, and preview instructions. ## Verification - 329 focused thread, composer, and interaction-card tests pass. - `pnpm -r typecheck` passes. The UI typecheck also passes after the review fix. - `pnpm build` passes for the full repository. - UI build, Storybook build, and token gates pass. The UI build and token gates were rerun after the review fix. - Greptile gives final commit `98f71f221` a 5/5 score. The current-head check passes, and there are no unresolved review threads. - All 56 current-head checks are terminal: 54 pass and two conditional Storybook jobs skip. The CI run includes the full test shards, build, typecheck, browser tests, aggregate verification gate, and package canary. - The stale-tab regression fails in both chat modes before the review fix and passes after it. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/planning-mode-visual-verification.spec.ts` passes against a throwaway local server. It checks dismissal, reload, task navigation, and desktop/mobile planning controls. - The duplicate local `pnpm test:run` attempt was stopped after the full CI test suites passed. It did not complete locally. - Manual browser test: select Green, dismiss, reload, reopen, submit the saved answer, and inspect the answered receipt. Also send a new message while the unanswered question remains in the feed. The test drive uses real UI components with fixture response callbacks. - To repeat the browser test, run `pnpm storybook`. Open **Chat & Comments → Task Chat Unanswered Questions → Test Drive**. ## Risks - Dismissal is a browser-local preference. It does not sync to another browser or device. Clearing local storage removes it. - When local storage is unavailable, dismissal lasts for the mounted thread only. - Unanswered questions can accumulate in the feed. They stay pending until answered or resolved through the existing API. - No database migration, API change, or change to approval permissions. ## Model Used OpenAI Codex, `gpt-6.1-sol`, with xhigh reasoning, repository editing, code execution, and browser control. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a386a59998 |
Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native agents receive task constraints and completion tools from Paperclip. > - Completion tools already define the procedure for reporting a result. > - Repeated procedure text adds instructions to each full task turn. > - The final reply must still explain a blocker and link a saved document. > - This pull request removes repeated procedure text and keeps these visible outcome requirements explicit. > - A document receipt supplies the exact link, and stricter evals check the persisted reply and browser navigation. ## Linked Issues or Issue Description Refs: #14961. Related: #14948 and #15007. **What happened?** Native task envelopes repeat completion procedure text. A reduced envelope needs explicit final-reply requirements. The `write_document` receipt also lacks a canonical document link. **Expected behavior** Keep the completion tools as the source of procedure details. Require one accepted completion result before the final reply. A blocked reply must explain the reason, owner and unblock action. A document reply must contain a working link to the saved document. **Steps to reproduce** 1. Run the native assigned-skill document case and native blocker case. 2. Inspect the run-attributed provider final and its persisted comment. 3. Check the blocker explanation or open the final reply's document link. ## What Changed - Remove repeated completion procedure text from the native task constraints and backend instructions. - Keep explicit blocker and document-link requirements in full task turns. - Return a company/task-scoped `documentHref` from `write_document`. Preserve the link in the idempotent mutation receipt. - Repeat canonical links for this run's current saved revisions in accepted completion feedback. Give blocked providers final-response guidance for the cause, owner and unblock action. - Keep internal document/comment anchors when Markdown issue links load cached issue details. - Add a manual six-cell comparison suite with strict source, build, default-instruction and budget admission. - Capture eighteen shared runnerd RPC projections and six direct OpenCode HTTP projections across start, resume and continuation phases, using scripted local transports and no provider execution. - Apply v3 checks only to the manual instruction comparison; preserve v2 checks for the existing native completion suite. Check the actual persisted blocker reason and exact saved-document link. Click the rendered document link and check the original content marker in the classic document card or the new document tab. - Forward exact OpenCode finishing calls through the controller. Wait for acceptance, keep accepted feedback and concrete rejection text, and reject malformed responses. Preserve ordinary dynamic-tool response handling. - Settle the completion decision and tool response before mapping a racing idle/error/abort event or handling explicit close/interruption. Reject a concurrent finishing call before controller admission. - Add a provider-free regression through real runnerd, the OpenCode proxy and a fake provider. Reject the first completion, accept the corrected report in the same turn, and propose one result. - Keep all original verdicts unchanged. Treat replay under new checks as separate diagnostics. ## Verification - `pnpm -r typecheck` and `pnpm build` pass locally. - Native document-authority tests pass, including company/run authorization and idempotent replay. - Native runtime-context, backend and measurement tests pass. - Final-answer calibration, protocol scoring, source-admission and catalog tests pass. Wrong reasons, absent links and wrong link targets fail. - `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six single-attempt local cells with the declared models. - Exported `prepareNativeInstructionPreflight` then `verifyNativeInstructionPreflight` pass on this clean committed source. They build locally and make zero provider calls. - Corrective live confirmation is incomplete. Source |
||
|
|
1c07b5903b |
feat: Chat leads the left nav, agent work beside chats, and a Combined Inbox + Task List flag (#15100)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The left nav is the main way people move between tasks, the inbox, and Agent Chat > - The nav has separate Inbox and Tasks rows that show overlapping work, and Chat is one row among many > - The side panel beside a chat shows the conversation's own artifacts, not the work the agent did > - People want Chat to be easy to find, and they want one place for their task views > - This pull request moves Chat to the top of Work, shows the agent's tasks and artifacts beside each chat, and adds an experimental flag that folds Inbox into Tasks > - The benefit is a shorter nav and a chat view that shows what the agent is working on. Both changes stay off until an operator enables them ## Linked Issues or Issue Description Refs #14706 (the secondary Agent Chat navigation this change builds on) Refs #14848 (reopen the last visited agent chat) **Subsystem affected** UI navigation (left nav, mobile tab bar), Agent Chat side panel, task list and inbox, and the company artifacts API. **Problem or motivation** Inbox and Tasks are two nav rows for overlapping work. Chat sits in the top group with no clear home. The chat rail lists only agents you already talked to, so you cannot see your other teammates there. The side panel beside a chat shows only the conversation's own artifacts. It does not show the tasks and files the agent made. **Proposed solution** With Agent Chat on, Chat leads the Work section and the rail lists every eligible agent. The chat side panel opens on the agent's tasks as cards, and the agent's artifacts are available from +. A new experimental flag, Combined Inbox + Task List, makes Inbox a set of views inside Tasks. **Alternatives considered** Rebuilding the inbox inside the task list. Instead, `/issues` hosts the existing Inbox component for inbox views and the existing task list for status views, so all inbox behaviour stays the same. **Roadmap alignment** Agent Chat (ROADMAP.md, "Agent Chat (including CEO Chat)"). All changes are behind experimental flags that are off by default. ## What Changed - **Agent Chat nav (streamlined shell):** Chat is the first row of Work, not a top-group row. Workspaces leaves the nav while Agent Chat is on. The mobile tab bar is Home · Chat · + · Tasks · Agents. The legacy shell keeps master's top-group Chat row. - **Chat rail:** `AgentConversationsSidebar` lists every eligible agent. The open chat is first, then conversations by recent activity, then the rest of the roster alphabetically. Terminated agents and agents you left are omitted unless you have history with them. The picker still marks only real conversations as "Open chat". - **Chat side panel:** a new default Tasks tab shows one card per task the agent created, was assigned, commented on, or acted on, newest first. It has the task list's filter popover and a sort control. **+ → Artifacts** shows the agent's artifacts as cards. Cards open in a new tab. Agent Chat off keeps the old Artifacts tab. - **Artifacts API:** `GET /api/companies/:companyId/artifacts` accepts `agentId`. The filter applies to documents, work products, and attachments by the agent each result is attributed to. The shared validator and the UI client carry the new parameter, and the OpenAPI entry picks it up from the shared schema. - **Combined Inbox + Task List flag (`enableCombinedInboxTasks`, off by default):** new card in Settings > Experimental. The Inbox row goes away and its badge moves to Tasks. A Views menu on `/issues` covers Mine, Unread, Blocked, Recent, Everything, All, Active, Backlog, and Done. Bare `/issues` opens the last-used view (default Mine). Links that carry `assignee`, `workspace`, `participantAgentId`, or `q` open All so the filter is kept. `/inbox/*` and `/issues/{all,active,backlog,done,recent}` redirect to the matching view. `/inbox/requests` stays its own page. - **Task detail breadcrumb:** the view key now decides the source, so quick-archive still works after a reload from an inbox view. - **Docs:** `doc/PRODUCT.md` and `doc/SPEC.md` describe the chat rail, the chat side panel, and the new flag. ## Verification - `cd ui && npx vitest run --no-file-parallelism src/components/chat src/components/task-side-panel/TaskSidePanel.test.tsx src/components/AgentConversationsSidebar.test.tsx src/components/Sidebar.test.tsx src/components/SidebarCompanyMenu.test.tsx src/components/Layout.test.tsx src/pages/AgentChats.test.tsx src/pages/InstanceExperimentalSettings.test.tsx src/lib/task-views.test.ts src/lib/issueDetailBreadcrumb.test.ts src/pages/Inbox.test.tsx src/pages/Issues.test.tsx src/App.test.tsx src/App.activity-routing.test.tsx src/components/MobileBottomNav.test.tsx src/components/CommandPalette.test.tsx`: 20 files, 356 tests pass. - `cd server && npx vitest run src/__tests__/company-artifacts-service.test.ts`: 13/13 pass, including the new agent-filter test across all three artifact sources. - The new rail test fails against the unmodified rail. - `pnpm check:token-gates`: all gates clean. - Manual: enable Agent Chat in Settings > Experimental. Open Chat. The rail lists all agents. Open a chat. The side panel shows the agent's tasks. Use **+ → Artifacts** to see the agent's artifacts. Then enable Combined Inbox + Task List. The Inbox row goes away, and Tasks shows a Views menu. - Snapshot baselines are intentionally not updated. See `doc/design/DECISION-SHEET.md`, "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks - With both flags off, the app behaves like master. The only exception is the API: it accepts a new optional query parameter. - With Agent Chat on, the rail can list many agents in a large company. It uses the agent list the app already loads, and search filters it. - The Tasks panel reads at most 200 recently updated tasks per agent and says so when it reaches the limit. The Artifacts panel reads at most 500 of the agent's artifacts. - Combined Inbox + Task List changes what bare `/issues` opens for people who enable it. Deep links with a task filter still open All. ## Model Used - Claude (Anthropic), model ID `claude-opus-5-5`, through Claude Code with tool use (shell, file edit, test runs). Extended thinking was enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: scotttong <squadbot000@gmail.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b17019e14d |
fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its adapters supply task context and access to Paperclip skills and tools. > - The default hire manual and shared prompts also repeat general work procedures. > - Those procedures overlap with stock provider instructions and the Paperclip skill. > - Existing E2E fixtures supply a QA manual, so they do not qualify the production default. > - This pull request reduces the generic instructions and adds real default-hire coverage. > - The benefit is less competing guidance, with inspectable evidence for preserved skills and task context. ## Linked Issues or Issue Description Refs: #14920. That merged change preserves native Codex base instructions. This PR covers the default manual, shared legacy prompts, operational skill guidance, and the narrowly approved ACP skill-discovery/session-environment repair for measured delivery and credential-persistence failures. **What existing behavior does this improve?** New non-CEO hires without a custom bundle and legacy task/chat startup and continuation prompts. **Current behavior** The shipped default manual contains 602 words. Generic task/chat prompts and ordinary resume deltas repeat work procedures already available through the harness and Paperclip skill. **Proposed behavior** The default manual contains only the eight-word company identity. Shared startup prompts retain identity and connection guidance. Ordinary resume deltas retain current work context without the generic execution contract. **Reason and benefit** Let the stock harness guide general work. Keep Paperclip-specific capabilities and independently test default hires, skills, ordered comments, and chat restart. **Breaking changes** New default hires receive less guidance. Existing saved manuals, explicit custom bundles, CEO templates, and specialized wake contracts retain their behavior. The obsolete includeExecutionContract option remains accepted for source compatibility. ## What Changed - Reduce the default hire manual to one sentence. - Reduce shared task/chat defaults and remove the generic ordinary-resume contract. - Keep connection guidance, auth, skills, custom prompts, and specialized wake context. - Add credential-free instruction-boundary gates and 26 explicit Product E2E cells across eight legacy/native profiles, including two focused Paperclip-storage cases. - Capture public hire receipts before providers run, then grade delivered prompts and independent task/chat outcomes. - Add an early legacy skill API recipe for saving a task document, checking the saved revision receipt and linking the document. Improve stock task/heartbeat skill-selection metadata and show a clickable Markdown UI-link example. Keep native tool completion separate. - Advertise bounded routing descriptions and exact successfully staged SKILL.md paths in legacy ACP Claude; keep full bodies on demand and preserve remote path rebasing. - Remove only the provider environment from copied persisted ACP session records, while loading current run credentials and preserving all other options/conversation state. - Regenerate both capability metadata inventories and reject stale manifests/inventories before provider admission. - Publish the original reduction and focused skill-repair comparisons, preserving all failures, automatic recovery, cost coverage and limitations. ## Verification **Behavioral qualification remains pending.** Original legacy ACP Claude loses the issue document only in the reduced cohort beneath an unchanged credential failure. A source-backed diagnosis finds that neither ordinary assignment reads the staged operational skill, while the runtime persists provider environment in session state. The new common repairs expose skill metadata/path and omit persisted env; strict document and credential guards stay intact. [Inspectable diagnosis and retained hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md). Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb` incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged #14961/#15007). All 222 affected adapter tests, adapter-utils/E2E typechecks, and final 96 variant/grader/retry calibrations pass. The frozen historical comparator is `c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths identical, exactly two production instruction paths plus nine declared unit expectations differ. The operational skill/discovery/environment repairs, selected model/profile/task/core grader/auth/permissions/retry policy are identical. Both actual launcher prepare→verify admissions pass with zero providers. [Immutable manifest and exact receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json). One original legacy ACP Claude cell per variant is authorized, with enforced single campaign attempts, 12-minute deadlines and company/agent 1,000-cent hard stops; every product recovery run/cost is counted. Actual live outcomes are pending. Current normal CI has one failed server shard and failed aggregate verify under diagnosis; other normal gates including typecheck/build/Rust/all eight browser shards pass. Fresh review completed successfully; the valid historical startup/resume masking finding was fixed with per-invocation task/chat checks and strict complete-snapshot capture, calibrated and resolved. Prior heads, failures and campaigns below remain historical evidence, not checks on this repair head. - Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52 current-head checks pass with two intentional Storybook skips, including repository typecheck/test/build and the browser shard. Fresh Greptile is 5/5 with zero unresolved review threads. Exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained receipt verification, zero providers/source errors. Fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck and 26-cell stock discovery pass. Canonical contract/inventory checks and the later issue-derived reference calibration are retained; that reference-only follow-up is not live-qualified by earlier frozen runs. - Prior full repository typecheck/build passed. The complete local Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three timing failures. All three affected files passed unchanged narrow reruns; original failures remain retained. Current-head CI now passes the full general checks; the original local failures remain retained. - The original 24-pair default-manual/shared-prompt comparison has two new overall classic Claude/OpenCode document-delivery failures plus an additional legacy ACP Claude document loss beneath an unchanged credential-guard failure (not closed by later runs), two newly passing OpenCode ordered cases, seven unchanged failures and 13 unchanged passes. Equal 15/24 totals do not establish behavioral equivalence. [Complete original report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md). - The skill-only repair holds the eight-word manual/shared prompts and merged #14920 fixed. All four matched profile configurations and 203 fixture/behavior files match. Candidate `abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill sources against baseline `bc83fe030234439ac51279502a28803958963e2e`. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547) and [baseline workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047) each pass 571 exact-source prerequisites before providers; all eight cells clean up successfully. Failed campaigns publish successfully and remain failed. - Repair pairs: Claude original Fail → Pass; Claude explicit Pass → Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff worsens beneath the unchanged failing UI-link grade: baseline gives a clickable API URL, candidate gives a code-formatted path without an anchor. The request's usable-link wording is narrower in the UI-only oracle. [Complete repair report and safe projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md). - The subsequent narrow stock metadata/link correction has two matched Pass → Pass cases, zero new machine failures/passes and no pending pairs. Both original-case handoff links remain deficient: candidate uses a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the preserved original oracle only requires a durable document. Both explicit clickable UI-link cases pass revision/content/link grading. All four exact-source 587-check gates, single assignment runs and cleanup pass. This does not establish fix causality because baseline also succeeds. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401) freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374) freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only comparison with reduced manuals/shared prompts held constant, not a repeat of the historical-manual comparison. Only original and clarified explicit classic OpenCode cases are selected, two per variant/four expected turns. 8,242 other tracked files and both profile hashes match; protected workflows admit each exact source before credentials. [Complete qualification report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md). Candidate original loads Paperclip/reference before saving publicly; baseline original loads it after writing locally, then saves publicly within the same assignment. Reported cost totals are $0.0107824490 candidate / $0.0107909015 baseline, with unmetered runtime. The later reference-only issue-derived link correction is provider-free calibrated and **not live-qualified** by these frozen runs; no further paid runs. - Retained tool calls show the repaired original OpenCode assignment loads only its assigned output skill before writing locally. Operational Paperclip is first loaded during automatic disposition recovery; its early recipe is visible then, but it never saves the missing document. Explicit candidate loads Paperclip and reads the new reference before saving successfully. All nine actual runs are counted. Reported LLM totals are $0.3802537209 baseline and $0.4918990161 candidate; local runtime is unmetered. - Initial setup, packaging, cancelled/missing-cell recovery, callback test and relative-output attempts remain retained. No completed provider failure was rerun. Frozen measurement branches are unchanged by later canonical metadata maintenance. - Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`, and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly for paid execution; it is excluded from `--all`. Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` replays this PR on merged hiring #14985 (`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four explicit custom-CEO-bundle checks, minimal generic manual boundary, and both suites. The combined fixture catalog and hiring calibrations pass 67 assertions; exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained-receipt verification, zero providers/source errors, fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E typecheck and 26-cell stock discovery pass. Fresh current-head CI passes all 52 checks with two intentional skips, and fresh Greptile is 5/5 with zero unresolved review threads. The prior source-plan browser failure is retained: a deterministic process fixture replayed its last `fixture:plan` command on `chat_task_completed`, writing revision 2 with identical body after the approval handoff. This was not paid provider execution. Rebased current-head CI passes the same assertion without an old-head retry or a change to that browser fixture. The merged hiring change was measured separately on immutable matched unions, with this reduced/shared/operational context and native completion guidance held constant. [Complete original two-profile report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md): [candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208) / [historical baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463), 705 provider-free prerequisites each. Both pairs are unchanged Fail → Fail on the exact-five count, with six core delivery checks passing all four cells; 28 actual successful runs include eight automatic completion wakes, zero retries, four successful cleanups. Source-read coverage is uncomparable, actual model charges unknown. Separately versioned provider-free accounting remains analytical work; original verdicts are preserved. This does not rerun or qualify the completed default-manual or native campaigns. ## Risks - Legacy ACP Claude's additional delivery loss is not closed by any later matched run and blocks the no-extra-failing-behavior merge criterion. Legacy document delivery may have relied on the prior manual/shared prompts. The early skill repair improves Claude in one trial; the later OpenCode pairs pass in both variants and cannot establish causality or robust recovery. Both original-case links remain deficient beneath the storage-only grade. The later issue-derived reference correction has only provider-free validation. Native finish/block descriptions must not be supplied to legacy agents. - The comparison holds merged native Codex fix #14920 constant; it cannot measure that fix's before/after task performance. - These bounded skill/context/chat workflows do not measure general coding quality. Unrepresented providers remain unqualified. - Saved manuals and old Codex sessions are not automatically migrated. Codex through ACP still has a separate base-instruction follow-up. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution, and delegated PR/eval work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (relevant suites and all three unchanged narrow reruns pass; complete-run timing failures retained in Verification) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green on the new repair head (prior-head checks retained above) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the new repair head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
78e0034498 |
fix(evals): account for hiring completion notifications (#15007)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Product E2E evals check real hiring and delegated task completion. > - The hiring fixture requires three requested CEO turns and two coder executions. > - The server can also wake the CEO when each delegated task completes. > - Two exact-five-run guards rejected these valid completion turns in all four retained cells. > - This pull request validates bounded completion turns in both guards. > - The benefit is accurate workflow grading while all actual runs and coverage failures remain visible. ## Linked Issues or Issue Description Refs: #14985, #14948, #14961. **What happened?** The original hiring comparison reports Codex Fail → Fail and Claude Fail → Fail. Each cell has seven successful runs. The five requested work turns are accompanied by two server task-completion notifications. All six other delivery checks pass. **Expected behavior** Require exactly three distinct user-requested CEO turns and one coder execution for each of two known tasks. Admit at most two strictly attributed server completion turns, including one turn that batches both tasks. Reject unknown, duplicate, failed, retried or extra-work runs. **Steps to reproduce** Inspect the retained four-cell report linked below. Each original result fails `five-successful-turns`. The same exact count was also enforced by the final chat-flow guard. ## What Changed - Add one typed lifecycle helper shared by the hiring scorer and the hiring-only final chat guard. - Validate public run ledgers, company/user/account identity, request attribution, task origins, completion deliveries, timing and replies. - Keep exactly five required work turns; declare seven maximum total turns for cost and timeout planning. - Count all actual runs, including notification runs and unexpected resets. Keep other chat count guards unchanged. - Version the hiring grader as v3 (turn accounting v2) and include the helper and chat guard in its definition digest. - Keep source-read, exact coder-body and all six other delivery checks unchanged. - Add 144 focused helper/scorer/settlement calibrations and separately versioned exact retained-input replay reports. - Retry complete bracketed observations, await both owed callbacks and attributed replies, and refresh the final guard consistently. - Reject unrelated completion writes and failed mutation attempts using exact canonical/native action IDs. Missing identity mapping is uncomparable action coverage. ## Verification - All 977 credential-free E2E support tests pass across 64 files, including 144 focused lifecycle/action/scorer/settlement calibrations. - E2E typecheck, ordinary plugin SDK and Runner TypeScript dependency builds, capability contract/inventory checks and the existing two-cell hiring discovery pass. - [Executable replay report](https://github.com/paperclipai/paperclip/blob/fed1729018cc100f5f4bbfb692777e49009c423b/doc/plans/2026-10-02-hiring-executable-accounting-replay.md) pins current code revision `e4077ade1818d98b9862ae79ee1d49a007dcf9c1`, v3 definition digest, exact source/input hashes and each original/new check. - The stricter replay verifies both Codex variants through both executable guards. ACPX Claude action attribution remains unresolved/uncomparable because provider execution IDs cannot be exactly joined to native request IDs; guards fail closed. No notification writes are observed. All six other outcomes and every original source/template coverage check stay unchanged. Original files and Fail → Fail machine verdicts remain preserved; zero providers are called. - Full attempts remain uncomparable in both profiles. Historical Claude also keeps its six-backtick exact-template mismatch. This grading repair does not prove model-performance equivalence. - The limited sidecar-v1 and initial executable-v2 passes checked notification-created tasks but could miss unrelated document writes. Those assessments remain preserved and do not prove harmless notifications. The stricter v3 replay is separate. - [Original measurement and separate sidecar](https://github.com/paperclipai/paperclip/blob/8eb517ca1497687237163bdef4dfc4d3332ea916/doc/plans/2026-10-02-hiring-template-live-comparison.md) retain 28 actual runs, eight automatic notifications, four successful cleanups and unknown actual model charges. No models are rerun. - The branch is replayed on master `59c07ede7`. Intervening master changes are UI-only; eval source bytes and replay verdicts match. The four-cell provider-free replay was repeated against the reachable code revision. - Initial-head normal CI retained browser failures in agent-run denial feedback and touch-picker scroll position. Those browser paths and imports were unchanged, but their cause was not established. The necessary review-fix head passes both browser checks; no blind rerun was requested. - Local full repository typecheck/test/build were not repeated. Exact-head normal CI passes the required repository gates, including typecheck, tests, build and browser shards. Fresh Greptile review completed on `fed1729018cc100f5f4bbfb692777e49009c423b` with 5/5 and zero unresolved threads. An independent rerun of the 144 focused helper/scorer/settlement tests passes on the unchanged head. **Merge readiness:** This PR repairs the evaluator. Its positive and negative calibrations pass, both guards reject missing action attribution, current-head CI and review pass, and there are no merge conflicts. The retained ACPX cells remain uncomparable because their action IDs cannot be joined. That coverage limit remains a separate follow-up; it does not require relaxing this grader or changing the old results. No model calls, production instructions, carrier changes, or historical regrades are part of this readiness update. ## Risks - Missing or inconsistent public lifecycle evidence fails the bounded helper. The focused calibrations reject plausible false positives and malformed observations. Unmatched action IDs fail closed and are reported as uncomparable rather than a model task regression. - Source-read evidence remains incomplete. This PR does not change provider event carriers or relax the coverage oracle. - The versioned count check differs from original v1 results. Reports retain both versions and exact input hashes. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution and delegated calibration work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip normal CI gates are green (exact head `fed1729018cc100f5f4bbfb692777e49009c423b`; fresh review tracked separately below) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (completed exact-head review; zero unresolved threads) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dd868ed125 |
fix(runner): share native completion tool guidance (#14961)
## Thinking Path > - Paperclip manages AI agents and their work. > - Native Runner agents report completion through finish and block tools. > - The providers receive different descriptions for those tools. > - Completion guidance belongs with the tools that enforce the result. > - This pull request shares the descriptions and refreshes retained catalogs. > - A separate native suite checks completion and blocking on production defaults. > - Legacy agents retain their separate skill and API paths. ## Linked Issues or Issue Description Refs: #14920, #14948, #14985. **Current behavior** Native Codex and MCP bridges describe finish and block differently. Retained provider sessions can keep old descriptions. **Proposed behavior** Native providers receive the same finish and block descriptions. The descriptions cover report selection, validation feedback, returned outcomes, approval gates and final-answer timing. Retained native sessions refresh from v13 to v14. **Reason and benefit** Put the completion procedure next to its native tool. Preserve stock base instructions, schemas, permissions and terminal semantics. This PR now stands alone on master. It contains no reduced manual, shared prompt or operational-skill changes from #14948. ## What Changed - Add canonical native finish and block descriptions. Use them in direct Codex and both native MCP bridges. - Advance the native tool contract to v14. Cover old-v13 refresh without replacing task identity or prior history. - Check authenticated tool catalogs, provider start/resume frames and serialized daemon catalogs. - Add an independent, explicit-only native completion suite. Preserve the original assigned-skill durable-document journey. Pair it with a concrete whole-task blocker across Codex, ACPX Claude and OpenCode. - Verify the actual public production default bundle and budgets before execution. Require independent durable disposition, native result/terminal receipts and observable provider-final ordering. - Correct the blocker browser oracle to accept the requested explanation. Keep exact owner/action/scope checks. Calibrate positive, missing and contradictory replies. - Preserve only actual `tool_call` terminal names (`paperclip_finish` / `paperclip_block`) in the native compatibility run-log projection. Require the same named call ID through its finishing result; retain all other redaction boundaries. - Admit verified hosted shallow checkout/build hydration and bind the selected runnerd to exact source/archive/binary provenance. Hosted cells truthfully reuse the existing trusted build; local admission executes Rust calibration. Forward only public source/run identifiers through both launcher preflight subprocess paths. - Enforce single attempts in the launcher for opted-in fixtures. Keep ordinary retry policy unchanged. Run exact-source, credential-free admission before credential loading. ## Verification - Frozen candidate: `d6e59e4712a3158ab4cd7d58deff1389b4578c21`, based on master `59c07ede72dc08b8aba149a01cc11e0b7a204621`; historical descriptions: `e74ed61a69fbdd8b3a8f15dd6456bc3140246e33`. Exactly the five original native production files and six unit tests differ. Both carry identical corrected fixtures, strict named finishing-call grader, closed compatibility carrier and admission. Defaults, profiles/models/auth/permissions and manifest bytes match. - Actual launcher `prepareNativeCompletionPreflight` → `verifyNativeCompletionPreflight` admission passes on both exact refs with zero providers: candidate 132 / historical 127 selected TypeScript assertions, 128 Node calibrations and one Rust normalization calibration each; E2E typecheck, manifest checks, selected binary provenance and six-cell discovery pass. Each has 257 explicitly skipped unrelated assertions, not coverage. The credential-free environment calibration exercises both real prepare/verify subprocess options with public hosted identifiers and rejects credential/ambient overrides. Complete actual launcher prepare→verify also passes on both frozen refs with explicitly synthetic hosted metadata/verified archives, separately labeled as calibration rather than a trusted GitHub run. Exact framed provenance parsing and mock source identity are calibrated without relaxing the real verifier. - [Complete matched qualification report](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-master-qualification.md), [immutable manifest](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-manifest.json) and [closed retained audit/hashes](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-results/comparison.json) are inspectable. All six candidate cells pass; historical descriptions pass five. Paired outcomes: **zero new failures, one new pass (Codex blocker), five unchanged passes, zero pending pairs**. [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37098728980) and [historical campaign](https://github.com/paperclipai/paperclip/actions/runs/37098815696) each execute six original attempt-1 native runs, with no campaign retry and successful cleanup. Their trusted workflow revision is `215586d127e97c9301d86e769a39a15c13298ca2`, separate from measured source. [Candidate public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098728980-1/index.html) and [historical public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098815696-1/index.html) retain declared screenshots. - Independent candidate evidence agrees with all original grades: 51 strict native checks, 12 served-default/budget checks and 21 original skill/document checks pass. The historical Codex blocker saves the correct whole-task blocker but omits the required marker from its actual provider final and identical saved reply. This is not semantic-summary fallback. Its original browser/matcher failure stays retained; the additional native snapshot/grade and workspace before/after digest were never written and are not fabricated by the separate API/PRP audit. Historical Codex completion has one failed finish followed by success within the same native run; the public receipt records no failure reason. All twelve runs and their usage remain counted. Reported model-cost subtotals are $0.00421482 historical/$0.00437391 candidate; Codex/Claude zero entries have unknown billing type, actual invoices are unverified and hosted execution cost is unmetered. One matched trial supports no extra failure within these six cases, not broad statistical or coding-quality equivalence. - Initial hosted `e18c2cf9` / `459455ac` and subsequent `0a9c5a7` / `00a761b` cohorts each stopped before providers in all twelve cells. The latter failed a mocked-receipt unit test under ambient hosted metadata; all source/build proofs passed. [All twelve later setup receipts](https://github.com/paperclipai/paperclip/blob/402ee94c52273ad58de355ae9a7d562dd22f8101/doc/plans/2026-10-02-native-completion-qualified-hosted-setup.json) are retained. [Exact failed setup receipts](https://github.com/paperclipai/paperclip/blob/27653eb1a8f8ce839776d760f4563f672e5a706c/doc/plans/2026-10-02-native-completion-master-hosted-setup.json) and the original manifest remain intact. Local sandbox-denied loopback and stale anchor-expectation attempts are retained separately; unchanged appropriate assertions were corrected/admitted before paid dispatch. Old anonymous OpenCode streams are not assigned inferred tool names or retroactively passed. - Full provider-free E2E support previously passed 927 tests in 67 files. Exact-head d6 normal CI run `37098409915`, attempt 1 passes full repository typecheck/build/tests, Runner Rust/static checks, all browser shards/aggregate and canary: 52 check-runs pass, four intentional skips, Snyk passes. Fresh Greptile check `111132956342` is 5/5 with zero unresolved threads. Source-specific deterministic tests do not substitute for the bounded live comparison. - Earlier native source `9138f570c341c251a5727c32d6615ce238bc8e03` is archived. Its [complete reduced-manual-context report](https://github.com/paperclipai/paperclip/blob/9138f570c341c251a5727c32d6615ce238bc8e03/doc/plans/2026-10-02-native-completion-live-comparison.md) remains intact, including original failures, grader limits and provider-free replay. It is not current-master-context qualification. ## Risks Changed tool text can change model behavior. The completed six-pair qualification shows no extra failing outcomes in this bounded trial; other tasks and repeated-run variance remain unmeasured. Observable final ordering does not prove provider feedback consumption. Public evidence can fail closed if a provider does not expose the required result sequence. This slice does not remove native fixed prompts or measure general coding quality. No database, schema, permission or legacy completion changes occur. ## Model Used OpenAI Codex, GPT-6 family, with code inspection, execution and tool use. The exact deployment ID and context-window size are not exposed in this session. They are unavailable rather than inferred from the model menu. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
862a5758ba |
fix(agents): reduce hiring templates to role descriptions (#14985)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - New agents receive role instructions from onboarding, the hiring skill, or a team package. > - These sources repeat harness procedures and impose generic work policies. > - They can crowd out the task and the harness instructions. > - This pull request reduces those sources to short role descriptions. > - It preserves configuration, skills, authentication, reporting lines, and approval controls. > - The benefit is less repeated instruction text with explicit coverage for the default hiring path. ## Linked Issues or Issue Description Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection. Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes. ## What Changed - Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets. - Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples. - Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog. - Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions. - Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents. - Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse. - Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison. Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing. ## Verification - PASS: 99 focused server tests and eight shipped-catalog tests. - PASS: catalog generation and validation for four shipped teams. - PASS: hiring skill validation. - PASS: `pnpm -r typecheck`. - PASS: `pnpm build`. - INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419. - PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests). - PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog. - PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck. - PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads. - PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge. The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence. ## Risks - New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome. - Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break. - Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review. - The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback. - Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
7a52dcdc74 |
fix: repair MCP validation and cancelled execution recovery (#14951)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The tool gateway gives agents access to connected services. Recovery controls what happens when a run stops. > - Generated tool names can exceed the provider limit after the MCP client adds its prefix. > - The same invalid definition can fail each automatic retry. A cancelled run can also hold saved messages without showing its cause. > - This pull request bounds tool names, stops configuration retries, and retains cancellation evidence. > - It shows the stopped run and admits saved input only after the existing safety checks pass. > - The benefit is a clear recovery path that preserves operator Stop and prevents duplicate message delivery. ## Linked Issues or Issue Description **What happened?** A long connected MCP tool name makes the provider reject the entire request. Automatic recovery repeats the invalid request. Separately, unexpected legacy cancellations can leave saved input behind a recovery hold. The notice does not identify the stopped run or its cause. **Expected behavior** Complete MCP names fit the provider limit. Tool-definition errors require configuration repair. Cancelled runs retain their source and reason. The recovery notice shows the cause and saved-message count. Verified unexpected cancellations can start a fresh turn through the existing admission checks. **Steps to reproduce** 1. Assign an App gallery connection with a long application key and tool name to a Claude agent. 2. Start a run. The provider rejects a name over 128 characters, including its MCP prefix. 3. For cancellation recovery, stop a legacy provider turn without an operator Stop request and send a user message while the recovery hold is active. 4. Inspect the recovery notice and the deferred message queue. **Paperclip version or commit** Rebased onto master at `cf8ad63c806685bfd7c48e3ed4a919d61a7c55f1`. **Deployment mode** Hosted or self-hosted server with legacy Claude or Codex execution. Related public work: - Refs #14017. That PR caps name segments. This PR preserves existing short names and uses stable hash aliases for long complete names. It also covers classification and recovery. - Refs #4510. That PR adds a cancellation-source column. This PR records bounded evidence in the existing run result, without a migration. - Refs #12552 and #4506. Those PRs suppress recovery after operator cancellation. This PR preserves operator intent and uses the existing continuation gates. ## What Changed - Bound gateway names with the full provider prefix in the 128-character budget. Retain the original upstream tool name for dispatch and permissions. - Classify invalid tool definitions as configuration failures before diagnostic redaction. Stop automatic retries and continuation attempts for that error code. - Persist cancellation source, expectedness, initiator, reason, and time. Preserve recorded Stop intent when adapter results arrive. Report unexpected started cancellations with closed diagnostic labels. - Show the run cause, saved-message count, and Inspect run link. Offer Continue for eligible unexpected cancellations. Require verified provider stop, empty tool inventory, ownership, and the existing pause, budget, approval, and dependency gates. Use the existing queue for single delivery. - Add regression coverage and update the execution, MCP gateway, and run-log documentation. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - `pnpm check:token-gates` passed. - Ran `pnpm test:run` and completed its workspace and serialized groups. Initial resource and timing failures passed on isolated reruns. All 149 serialized route suites passed. - Reran the changed server, adapter, and UI suites after the rebase. Coverage includes long-name upstream dispatch, configuration retry suppression, cancellation evidence retention, privacy labels, oversized run projection, and concurrent saved-message delivery. - `pnpm test:e2e tests/e2e/legacy-failure-continuation.spec.ts` passed all six browser scenarios. The recovery notice shows the run cause and inspection link, and each recovery entry point reaches one new response. - Added database-backed checks for active, removed, paused, unavailable, and disabled chat connections. The final continuation and recovery-notice suites passed 167 tests. Externally bound chats hide board Continue and show a usable next action. - All 55 GitHub checks passed on `42afbf1371dcaeb72646e3d8f65c19ff7cddf8de`. Two unrelated Storybook jobs were skipped by their normal conditions. Greptile reviewed that commit at 5/5 with no findings and no open review threads. ## Risks - Long tool names change to aliases. Existing short names stay compatible. The original connection and upstream name remain the dispatch authority. - Invalid tool definitions no longer get automatic retries. An operator must repair the configuration before a new attempt. - Continuation changes apply only to positively identified unexpected legacy cancellations with complete empty tool inventory. Operator Stop, unknown historical cancellations, outstanding tools, and unverified provider termination keep their holds. - No database migration. The added projection fields are optional. Cancellation reason and initiator IDs remain local run evidence; Sentry receives only closed source and initiator-type labels and expectedness. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository editing, shell execution, and GitHub tool use. The runtime does not expose the exact model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6c1a75da49 |
feat(connections): make AgentMail a default connection with inline setup (#14772)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents access to external services. > - AgentMail needs both a saved key and an inbox assigned to the agent. > - Chat requests offered a setup link instead of an inline card and could treat a saved key as complete. > - Inbox setup also hid address conflicts behind a generic server error and a separate review step. > - This pull request makes AgentMail a default connection, adds the inline card, reduces setup to two steps, and shows conflicts beside the address. > - Shared native dropdown styles also give every caret a consistent inset. ## Linked Issues or Issue Description **What happened?** AgentMail requests in chat did not show a usable inline connection card. Manual setup required extra screens, ignored saved account keys, and could trap new-address setup in a locked inbox dropdown. Agent selectors omitted the avatar from the selected value. A taken address could produce an HTTP 403 from AgentMail and appear as an internal server error. Native dropdown arrows also touched the right edge of their fields. **Expected behavior** Make AgentMail available as a default connection. Ask for the API key inline, with a direct link to its provider page. Default human access to the company and agent access to the requesting agent. Resume the agent only after an assigned inbox is active. Manual setup should ask for an agent and email address, then finish. Address checks should run as the user types. Taken addresses should show clickable alternatives. A domain dropdown beside the name should prefer a verified custom domain. Setup should suggest authorized saved AgentMail keys and show agent avatars in the picker and selected value. **Steps to reproduce** 1. Ask an agent to connect AgentMail when it has no assigned inbox. 2. Check that an inline API-key card appears and links to the provider's API-key page. 3. Open AgentMail setup, choose an agent, and request an address that is already taken. 4. Correct the inline error, refresh, and finish setup with the same request ID. 5. Inspect native dropdown carets in light, dark, disabled, and right-to-left states. Uses the bounded provider-error parser merged in #14768. Related work: #13256 introduced AgentMail; #14725 expanded connection search. ## What Changed - Stop recurring email queries for tasks that have no email thread. Share the query between the thread provider and activity view. Keep email-task updates and invalidation-based discovery. - Make AgentMail available without the experimental chat setting. Keep the catalog, setup and management routes, agent Channels tab, task email feed, receiving worker, and agent tools available by default. Other experimental chat providers stay gated. - Make the email address and copy icon a single clickable action with the shared Copied! confirmation. Add View inbox linking directly to the matching AgentMail console inbox, with the address encoded as one URL path segment. - Reorganize inbox Settings around the copyable email address, usage instructions, and receiving status. Move reconnect credentials into a disclosure and separate the Disconnect action. Add production Settings stories for active, paused, unassigned-address, revoked, webhook, long-address, mobile, and reconnect states. Show repair controls when the inbox has an error. Keep usage instructions tied to an active inbox with an address. - Add AgentMail channel intents and an inline key field with the direct API-key URL. - Keep setup and retry state tied to the interaction. Require an active inbox for completion. Preserve company and agent access checks. - Reduce manual setup to agent selection and email selection. Put the domain dropdown beside the address and default to a verified custom domain. Preserve explicit choices across reloads. Keep receiving settings under Advanced options. - Check the initial address and edits after a 350 ms pause. Abort superseded requests and ignore stale responses. Show clickable suggestions and retain known creation conflicts across reloads. - Add a company-scoped, manager-only address check using the saved credential. Search the visible inbox list instead of fetching an uncreated inbox: live AgentMail retains negative lookups that can break subsequent access-key creation. Unlisted addresses remain unknown; creation is authoritative. - Suggest labeled saved AgentMail keys in both manual setup and the inline card. Filter by company, provider, active credential, and current-user grants on the server. Prefer an account key and preserve the selected key or an explicit new-key choice across refresh. Use verified scope metadata and bounded concurrent checks for legacy keys. Never return secret values. - Catch an inbox-only key before the email step. Allow its existing inbox only after an explicit choice. Recover old locked drafts at the key picker. Save the replacement key before retiring an empty draft, then use a new setup URL so refresh preserves the switched account; stop if cleanup fails. Preserve already allocated addresses and their original accounts. - Use the shared AgentSelect in email setup. Show the canonical agent avatar in each option and the selected value, including other consumers of the shared component. Add regression coverage for legacy and current Lucide agent-mention icon formats. - Start each catalog Add connection with a fresh setup identity. Honor Finish setup's exact draft/account/address instead of resuming an unrelated browser draft. Return Cancel and Done to Connectors and Email settings to the inbox. Group the task/thread explanation in a How it Works card. - Route AgentMail catalog removal through the email inbox control API, including unfinished drafts. Refresh both the catalog and inbox views. - Render each inbox management tab separately. Access uses the saved account grants and agent controls; Conversations and Activity use the shared persisted email feed. Activity lifecycle actions use the email API. Reconnect returns to inbox Settings. Conversation failures show a retry instead of a false empty state. Email delivery recovery stays in the task. - Map documented provider address conflicts to a field error. Preserve actionable messages for other failures. - Preserve non-secret draft fields across refresh, scoped to the requested agent. Never save API keys in browser storage. Resume partial inbox creation with the original agent, address, and request ID. - Show an already-created address with explicit retry and new-address recovery instead of locked inputs. Preserve the original inbox and resumable draft when choosing another address. Distinguish runtime-key 404 errors and log safe provider status/operation/code. - Apply final agent access once within email setup authorization for a new account whose original installs are unchanged. Preserve later permission edits and reused account installs. Support in-place retry of progress loading. - Let a failed inline setup change keys after retiring an empty draft. Persist its replacement setup identity without storing secrets. Recover a server-saved account when refresh interrupts the save response, while preserving intentional account changes. - Render the production setup in Storybook and add error, recovery, and mobile states. - Inset native select carets in shared CSS. Preserve custom icons, listboxes, keyboard behavior, and forced-color controls. - Add browser regression coverage and an AgentMail Product E2E case with persisted-state and rendered-card evidence. ## Verification - Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `git diff --check` passed after the default-availability change. - All 485 focused tests passed. These cover setup, management, catalog and route gates, connection intents, email authorization, Cursor execution, and the OpenAPI contract. All 39 email integration tests run with the experimental chat setting off. - The shared polling change passed four behavioral tests, UI typecheck and build, and token gates. - `tests/e2e/agentmail.spec.ts` passed with the actual server setting off. This full-stack browser test uses simulated provider responses. It covers catalog entry, saved keys, editable address and domain controls, creation, conflicts, retry, all management tabs, clipboard feedback, the provider link, and task email rendering. - In the live local browser, Add connection reached the editable email step with the saved account key. The verified custom domain was selected by default. Both domain choices worked. The existing inbox Settings page remained available. Both active inboxes completed new mail checks with the setting off. No new provider inbox or email message was created for this pass. - Earlier live provider acceptance covered creation on a verified custom domain, Finish connecting on the reported draft, successful mail checks after refresh, and catalog removal of disposable draft and active connections. Clicking the email address copied the exact address and showed Copied!. View inbox opened the same inbox in AgentMail’s console. No email messages were sent. - Production setup and Settings Storybook builds and interactions passed. Settings states include active, paused, unassigned, revoked, webhook, long-address, mobile, and reconnect. Receiving and revoked-access stories had zero accessibility violations. - Full local `pnpm test:run` on an earlier revision completed with 14,709 passing, 87 skipped, and four transient failures. All four failed cases passed in focused reruns without product changes. That serial full local command was not repeated after each follow-up. The latest-head full CI suite is the final test gate. - CI found an obsolete browser assertion that hid every channel when the flag was off. Updated it to keep AgentMail and the Channels surface visible while preserving the GitHub chat route gates. All 11 provider browser tests passed locally after scoping the Channels selector to the agent sidebar. Two initial local attempts stopped at temporary Postgres initialization. The passing run used a separate disposable database on the existing local Postgres server; it was removed after the test. - Updated the remaining sidebar and aggregator discovery assertions for default AgentMail availability. Ordinary task fixtures now return no email thread. All 128 sidebar/task-page tests and all 42 aggregator tests passed locally. - Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI passed, with 54 successful checks including Snyk and two intentional Storybook skips. The CI run is https://github.com/paperclipai/paperclip/actions/runs/37020833647. A fresh Greptile review scored 5/5 with no unresolved threads. Live model evaluations and inbound/outbound email delivery were not run. ## Risks - AgentMail no longer needs experimental opt-in. Setup still requires a human to connect an account and assign an inbox. Inline setup creates an inbox after a human submits a new or saved key. Company access, agent access, inbox assignment, and completion checks remain enforced. - AgentMail read APIs cannot prove global address availability. The visible-list check is bounded to 100 entries and cannot see inboxes outside the key’s scope. The UI reports this limitation, suggests alternatives without claiming they are free, and keeps final creation conflicts inline. Lookup outages show an error without preventing the authoritative creation attempt. - Native select CSS affects the whole app. Custom-icon selects and multi-row lists are excluded. Forced-color mode keeps the browser caret. - Saved-key discovery uses stored verified scope metadata and checks authorized legacy credentials concurrently within a shared three-second deadline. Provider outages mark legacy choices unavailable; users can still enter another key. Final use rechecks authorization and provider access. - No database migration or transport default change. Live connection remains the default. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The exact served model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cad26c6bfb |
fix(tool-gateway): bound MCP discovery memory and concurrency (#14864)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents discover governed tools through the MCP gateway. > - A listing repeated policy and full run-row reads for each catalog tool. > - Parallel listings multiplied those allocations during run startup. > - One 900-tool baseline listing used 11,489 queries and about 1.7 GiB of extra heap in a fixture. > - This pull request shares reads within a listing and bounds whole listings across the process. > - The benefit is lower discovery memory use while execution still checks current policy. ## Linked Issues or Issue Description Refs #13115. Its on-demand target change affects the same listing loop. **What happened?** MCP discovery repeated roughly 13 reads per tool. Full run snapshots and repeated connection configurations caused large allocations. Per-listing bounds alone did not limit concurrent listings across gateways. **Expected behavior** Discovery reads shared inputs once per listing. The process bounds active and queued listings. Catalog payload size and policy evaluation still grow with the catalog. Tool execution checks current access rules. **Steps to reproduce** Create a company with a remote MCP connection, 900 catalog tools, large schemas, and a large run snapshot. Send concurrent tools/list requests using a run-bound gateway token. Run the committed benchmark for a deterministic reproduction. **Paperclip version or commit** Baseline: |
||
|
|
33a00d2f1e |
fix(ui): reopen last visited agent chat (#14848)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent Chat keeps one conversation for each agent and board user. > - The Chat sidebar entry opens the agent chooser each time. > - A user must then find and reopen the chat they just used. > - The browser already records recent agent chat visits by company and user. > - This pull request uses that record to reopen the last available chat. > - The chooser still serves users who have no available saved chat. ## Linked Issues or Issue Description Related: #14706 added the secondary Agent Chat navigation. **What happened?** The Chat sidebar entry opened the agent chooser, even after a user opened an agent chat. **Expected behavior** The Chat entry should reopen the last agent chat visited by the current user in the current company. **Steps to reproduce** 1. Enable Agent Chat and open a chat with an agent. 2. Open another page. 3. Select Chat in the sidebar. 4. Observe the agent chooser instead of the chat. **Paperclip version or commit** Reproduced on master at `0829d94af`. **Deployment mode** Local development, browser UI. The change also uses the same browser storage path in authenticated mode. ## What Changed - Use the existing recent chat record when the Chat landing route opens. - Check saved agents against the current roster and chat history before redirecting. - Keep the chooser when no saved chat is available, and show a retry state for load errors. - Add route tests and update the Agent Chat implementation spec. ## Verification - `pnpm exec vitest run ui/src/pages/AgentChats.test.tsx ui/src/lib/recent-agent-chats.test.ts` — 16 tests passed. - `pnpm check:token-gates` — passed. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/agent-chat-sessions.spec.ts --grep 'secondary chat navigation preserves layout'` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed on the final commit. - `pnpm -r typecheck` and `pnpm build` — passed earlier in this branch; latest-head CI completed all 47 jobs successfully. - `pnpm test:run` reported an unrelated native runtime test failure before it was stopped. That test and an unrelated external object refresh test passed in isolation. CI runs the same suites on the PR. - To check in the UI: open an agent chat, leave it, and select Chat. The same chat should open. Clear the recent chat record or use another company to see the chooser. ## Risks - The recent order is stored in the browser. Clearing browser storage returns the user to the chooser. - An existing chat ID is stored with its visit. If the chat is removed, the landing route skips that visit when history loads. Cross-tab storage removal clears the identity; failed writes retain an in-tab fallback. - The landing route waits for the agent roster and validates saved issue IDs against chat history when available. If history fails, an active agent chat can still open; roster or session failures show a retry action. - No database or API contract changes are required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-6 family. The runtime did not expose an exact API model ID or context window. It used reasoning, repository tools, shell commands, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
467125fafb |
feat(connections): one-screen connector setup with stated defaults (#14811)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents use Connections (the Apps catalog) to act in services like
Notion, GitHub, Google Workspace and Railway
> - Each connector asked the user to answer setup questions before it
went to the provider. Most of the questions already had the correct
answer selected
> - ROADMAP.md lists "simpler setup" for Apps and Connections as ongoing
work. This change continues that work
> - This pull request removes the questions that Paperclip can answer
itself. It states the defaults in one line and moves the choices behind
"Change" and onto the Permissions tab
> - The benefit is that most connectors take one click in Paperclip and
then the provider's own consent screen
## Linked Issues or Issue Description
No public issue exists. This is the description, from the enhancement
template.
**What existing behavior does this improve?**
The setup flow for tool connectors in the Apps catalog.
**Subsystem affected**
Apps and Connections: `ui/src/features/connections`,
`ui/src/pages/apps`, the `packages/shared` app definitions, and the
OAuth routes in `server/src/routes/tool-access.ts`.
**Current behavior**
Every connector opened with an Access step. The step asked who can use
the connection and which agents get it, and both answers were already
selected. 18 connectors also asked "How do you want to connect?" when
Paperclip could rank the methods. The Google apps and Postman also asked
"What should Paperclip be able to do?" before sign-in. The four gateway
connectors (Zapier, Arcade, Composio, Executor) used a separate two-step
wizard. Asana was pinned to a customer-owned OAuth app, so the user had
to register an app in Asana's developer console. The "Set all" control
on the Permissions tab changed only one action. After the user approved
access, Railway's consent page showed "you can close this window" and
did not return to Paperclip.
**Proposed behavior**
One screen per connector, with one primary button. The screen states the
defaults in one sentence, for example "Connects for everyone in your
organization, available to all agents". A "Change" link opens one
Advanced panel. When the provider's metadata allows dynamic client
registration, Paperclip registers a client itself. Connecting lands on
the Permissions tab. On that tab, "Set all" changes every action in the
group.
**Reason and benefit**
The user makes fewer decisions before the connection exists. Most
choices are easier to make after the connection, on the Permissions tab,
where a change has an immediate effect.
**Breaking changes**
None. No schema or API change. Existing connections keep their settings.
## What Changed
- **No Access step.** `ConnectionSetupFlow` no longer has the Access
step. The flow shows the resolved default above the primary button and
on the completion screen. The access controls moved into one Advanced
panel. The panel opens automatically only when a setting in it is
required.
- **A default method for every app.** The flow always picks the ranked
default method. Alternate methods are in the Advanced panel. The Google
and Postman capability choice is not asked before sign-in. The
write-capable method is the default.
- **Gateway connectors.** `RemoteMcpProductionSetup` (Zapier, Arcade,
Composio, Executor) no longer has its own Access step. Its commit path
and the main commit path use one helper, `askFirstCatalogEntryIdsFor`,
for server-suggested defaults.
- **Dynamic registration from live metadata.**
`canRegisterOAuthClientDynamically` now allows registration when the
provider advertises a registration endpoint, even if the catalog entry
lists only customer-owned clients. The Asana and Linear definitions and
catalog text match live probes. Asana issues clients for loopback
callbacks only, so a hosted deployment still needs an Asana app.
- **Connection setup states.** New
`packages/shared/src/connection-setup-state.ts` sorts each method into
`instant`, `authorize`, `paste` or `register`. The gallery card verb
("Connect" or "Add key") comes from this resolver and the instance's
ownership availability.
- **Generic MCP.** The generic path no longer asks "Does it need a key?"
first. A credential challenge from the server shows the key field.
- **Permissions tab.** Each action row shows its risk level. Each group
has a "Set all" control. The control sends one change for the whole
group. Before, each row's save started from the same render, so the
saves overwrote each other. The Zapier/Arcade/Composio/Executor setup
screen had the same defect.
- **OAuth callback interstitial.** A cross-site browser navigation to
`/api/tools/oauth/callback` gets a small same-origin "Finishing your
connection…" page. That page repeats the request, and the repeat does
the code exchange. Railway's consent page replaces itself after about
two seconds, and the code exchange plus tool discovery takes longer than
that. The interstitial uses only a meta refresh, because the OAuth code
is single-use. Requests without `Sec-Fetch-Site: cross-site` take the
old path.
- **Linear registers through its MCP server.** Linear pins the console
endpoints at `linear.app`. Pinned endpoints now replace discovery only
when the method cannot register, or when the connection has an
operator-entered client. So a Linear connection now finds the
registration endpoint at `mcp.linear.app`.
- **Own-OAuth-app recovery stays on the one-click screen.** When the
method also accepts a customer-owned client, the client fields are in
the Advanced panel. The panel opens after a failed sign-in. "Try again"
resumes the draft with the operator's client.
- **E2E specs** follow the one-screen flow. The Access-step clicks are
removed, the specs open **Change** before they pick agents, and they
expect GitHub's **Add key** verb.
- **Default permissions do not change.** New connections still allow
every action. The user can set actions to Ask first or Off on the
Permissions tab.
## Verification
- `cd ui && npx vitest run src/pages/apps src/features/connections
--no-file-parallelism`
- `cd packages/shared && npx vitest run src/app-definitions.test.ts
src/connection-setup-state.test.ts`
- `cd server && npx vitest run src/__tests__/tool-access-service.test.ts
src/__tests__/remote-mcp-connectors.test.ts`
- `pnpm check:token-gates`
- New tests:
- `PermissionsPanel.group.test.tsx` checks that "Set all" sends one
change for the whole group. It fails on the old code.
- `action-permissions.test.ts` checks the group update.
- `connection-setup-state.test.ts` checks the four setup states.
- A server test checks that a cross-site callback gets the interstitial
and does not use the OAuth state, and that the same-origin repeat
completes the connection.
- Manual check on a hosted staging deployment. GitHub, Google Drive,
Composio, Notion, PostHog and Railway each connected from one screen and
returned to the Permissions tab. On Railway, "Set all" changed all 65
write actions, and the change remained after a reload.
- Visual changes: snapshot baselines are intentionally not updated. See
the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot
verification demoted to dormant (Jul 13 2026)".
## Risks
- **Fewer confirmation clicks.** Organization-wide access is the
default, and the user does not confirm it on a separate step. This was
already the preselected answer. The flow shows the default before the
user clicks and again after the connection.
- **Google write scope.** Google apps now request the write-capable
scope by default. A narrower scope needs a new sign-in.
- **Dynamic registration from live metadata.** A provider can advertise
registration and then reject a redirect URI. Asana rejects hosted
callbacks, for example. In that case registration fails, and the
customer-owned client path remains available for recovery.
- **Callback interstitial.** The OAuth callback adds one same-origin
step for cross-site browser navigations. Browsers without `Sec-Fetch-*`
headers use the old direct path.
- Chat and bot connectors (Discord, Telegram, Microsoft Teams, iMessage)
do not change.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, used through
Claude Code with tool use (shell, file editing, browser automation) and
extended thinking. It wrote the code, the tests and this description. A
human product owner directed the work and tested it by hand.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: scotttong <squadbot000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
33f2b3a159 |
fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Connectors catalog lets people give agents tools or connect agents to conversations. > - GitHub put these two uses behind one card and an extra choice. > - People should choose the connection they need from the catalog. > - This pull request keeps GitHub for tools and adds GitHub Code Review Bot as a separate card. > - Each card opens its setup directly. Both use the existing connection code. ## Linked Issues or Issue Description **What existing behavior does this improve?** GitHub connector discovery and setup. **Current behavior** With chat connectors enabled, GitHub opens a menu that asks whether to use tools or create a bot. Saved tools and bots share the same catalog entry. **Proposed behavior** GitHub opens tool account access. GitHub Code Review Bot opens agent selection. Saved bots and drafts appear under the bot card. Chat-disabled instances show only GitHub tools. **Reason and benefit** The catalog names the two uses and removes an extra setup choice. The bot keeps the existing GitHub provider, credentials, endpoint IDs, setup steps, and runtime. **Additional context** Related work: https://github.com/paperclipai/paperclip/pull/12843 and https://github.com/paperclipai/paperclip/pull/14594 established GitHub account identity. This change preserves that tool flow. No duplicate catalog split was found. ## What Changed - Split the generated app definitions into GitHub tools and GitHub Code Review Bot. Reuse the existing GitHub logo and channel method. - Open bot setup directly, including old resume and reconnect links. - Put existing bot endpoints and drafts under the bot card. Hide duplicate internal chat applications. - Keep pasted GitHub URLs mapped to the tool connection. - Add seven Storybook states for the catalog, saved connections, disabled chat, both setup paths, mobile, and light mode. - Fix narrow-screen bot rows so the label cannot overlap status and setup actions. - Update catalog, route, browser, and API tests, plus the GitHub connector guide. ## Verification - [Hosted Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog): seven states built from this branch. The deployment passed its public-file verification. - All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`. Two optional Storybook jobs skip under their normal trigger rules; the manual Storybook deployment passes. The branch has no merge conflicts. - Greptile: 5/5 on the current head, with no review comments or unresolved threads. - `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm build-storybook` passed. The final Storybook fixture also passed UI typecheck and the hosted build. - Targeted catalog, URL matching, routing, grouping, brand, and chat UI contract tests passed. - GitHub provider browser tests: 2 passed. These cover direct tool setup and the bot setup and management lifecycle with provider responses mocked. - Embedded-browser test on an isolated local instance: opened both cards, selected an agent, saved a bot draft, and resumed the same endpoint under the bot card after a reload. - Storybook Tool Setup and Bot Setup assertions pass in the published preview. Chat Disabled assertions pass locally. Inspected mobile and light mode, including the draft-row layout and official GitHub marks. - Local full-suite limitation: `pnpm test:run` was not clean. A cross-company route assertion failed in the aggregate run and passed in isolation; a workspace-runtime test reached its 30-second hook timeout. Some isolated database reruns skipped when the embedded-PostgreSQL availability probe failed. The local aggregate was stopped after CI completed. The corresponding full CI suites pass all 360 tool-access tests and all 162 workspace-runtime tests. - No live GitHub authorization or installation was performed. The isolated instance correctly stopped at the cloud enrollment or public HTTPS prerequisites. ## Risks - Low scope: catalog presentation and routing change. There is no database migration or provider credential change. - Existing GitHub bot URLs now open bot setup directly. The tool route remains `/apps/connect?source=github`. - The bot remains behind the existing chat-connectors feature flag. Existing endpoints retain `provider: github`. - Channel applications are represented by endpoint rows. Regression tests cover legacy bot applications, tools, active bots, and drafts together. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and embedded-browser tools. The exact deployed model ID, context window size, and reasoning setting are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018993140f |
feat: let agents name prompt-only tasks (#14761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad55d0a281 |
fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use Apps through a gateway that checks identity, company access, and action policies. > - Personal pasted credentials can point to company secrets. Setup can show success while the gateway rejects every call. > - Several OAuth methods also omit the scopes needed for their supported write actions. > - This pull request gives setup, health checks, and invocation the same credential rules. Owners repair existing connections by reconnecting. > - New connections request reviewed permissions for their supported actions. Read-only choices remain available under Advanced. > - Agents can use the connections people give them, while existing consent, identity boundaries, and action restrictions remain enforced. ## Linked Issues or Issue Description Refs #14009 and #14008. This addresses the personal-credential defect. The separate GitHub organization-identity selection defect is outside this change. Related work: #13942 fixed part of new personal-key setup. #14200 independently fixes legacy personal reconnect and protects managed-agent profile credentials during removal. This PR covers that ownership invariant across key and secret-URL setup, reconnect, health, discovery, and invocation, and keeps owner reconnect as the repair path. #14059 tracks requested versus provider-asserted OAuth scopes; it remains separate work. I searched open PRs and issues for Zapier, Airtable scopes, connector writes, and personal credential failures. **What happened?** A Zapier secret URL saved through personal setup can become a company secret referenced by a user grant. Health checks bypass the gateway's ownership check, so the connection appears healthy but calls fail with `grant_credential_invalid`. Custom-header paths can also receive a duplicate `credentials.` prefix. Omitted OAuth scopes make write access depend on provider defaults. **Expected behavior** Personal invocation credentials belong to the selected user. Setup, health, and actual calls enforce the same rule. New connections request documented permissions for supported read and write actions. Existing tokens gain no permissions without provider consent. **Steps to reproduce** 1. Connect Zapier or a generic secret URL with the personal identity. 2. Allow an agent to use the connection and complete setup. 3. Invoke a tool through a run-scoped gateway. The legacy layout fails ownership validation despite successful setup. **Paperclip version or commit** The implementation started from `44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master at `94e8dec56`. **Deployment mode** Built from source. Regression tests use isolated PostgreSQL fixtures and controlled MCP transports. ## What Changed - Share credential writing, ownership validation, and canonical paths across initial setup, resume, reconnect, rotation, health, discovery, and gateway calls. Keep OAuth client-registration secrets separate from invocation credentials. - Existing personal connections with company-scoped credentials require owner reconnect with a fresh key or secret URL. Reconnect creates a correctly owned value and updates the existing grant and declarations. There is no automatic ownership backfill or new startup hook. - Preserve PostgreSQL timestamp precision when reconnect checks whether a grant changed. Previously, converting the timestamp to a JavaScript Date could reject reconnect with a false concurrent-change error. - Protect credentials used by other grants, connections, bindings, managed-agent profiles, routine triggers, or secret proposals from connection removal. - Review all 117 tool methods, including 84 OAuth methods. Record explicit scopes or documented provider-default exceptions with official evidence. Add Airtable's seven scopes, Hugging Face repository/job scopes, and other documented MCP permissions. - Prefer available write-capable methods. Put explicit read-only choices under Advanced. Explain pasted-key permissions and offer reconnect for missing OAuth consent. Preserve existing grants, policies, Google availability gates, and curated scope allowlists. - Reconnect generic secret URLs and custom headers using their stored credential fields. Refresh the catalog after setup, correct reconnect feedback and error guidance, and let Cancel exit invalid setup while Save & exit retains draft-saving behavior. - Apply ownership checks to the new GitHub repository/skill connection picker. Align the permission audit with the Google scope reductions merged on master. - Add run-scoped gateway, ownership, owner-reconnect, OAuth URL, insufficient-scope, UI, and catalog-wide regression coverage. Update the connector playbook and permission audit. ## Verification Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all CI/status gates (55 completed check runs, no failures or pending checks) and has a completed Greptile review at **5/5 with no outstanding findings**. GitHub reports the PR as mergeable/CLEAN. - **Embedded browser:** used the actual server and built UI from this worktree, a fresh isolated database, and local HTTP MCP fixtures. Completed personal bearer-key, secret-URL, and custom-header setup; reproduced the legacy ownership failure; reconnected through the owner’s form; and completed writes afterward. Read-back was verified for bearer-key and secret-URL connections. Public organization-wide setup appeared immediately in Browse without reload. Zapier URL validation/Cancel and Google’s enrollment gate were also exercised. - **Persistence and invocation:** verified user ownership, canonical `credentials.authorization` / `remote.url` / `headers.X-Api-Key` declarations, and unchanged connection/grant identity. The old company secrets retain their ownership. Separate HTTP calls through an actual run-scoped gateway session completed a write and read-back. - **Backend coverage:** the final gateway suite passes all 82 cases, including catalog Zapier and generic inline reconnect. It checks company/user isolation, canonical declarations, same-endpoint URL validation, fresh credentials, retained restrictions, and real gateway read/write execution using fixture transport. A timestamp with PostgreSQL microseconds covers the former false reconnect conflict. - **Local checks:** 368 catalog, gateway, repository, and UI tests passed before the final extra Zapier case; 49 GitHub skill access tests also passed. All three Apps browser regressions pass, including reconnect through the actual form and catalog visibility without reload. Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final patch, and token gates passed. Full tool-access service runs hit varying 15-second Google fixture timeouts; both affected cases and the updated reconnect assertion pass in isolation (3 tests). The complete test matrix passes in CI on this head. - **Verification limits:** no live provider account was available for Zapier/Airtable/OAuth consent or account-bound write proof. Public metadata and local fixtures do not establish provider consent. The original development database clone failed on a pre-existing missing `tool_connections_transport_check` constraint; browser acceptance used a fresh isolated database created by the normal CLI onboarding flow. ## Risks - Existing broken personal connections stay unusable until their owner reconnects. Health, discovery, and invocation return an actionable ownership error; startup does not rewrite credential ownership. - Scope changes affect new authorization requests. Providers may still require resource selection, account roles, paid plans, or app verification. Existing consent and action restrictions remain unchanged. - Shared credentials are retained rather than reassigned or revoked. Provider-default exceptions and unavailable live checks are documented in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`. - No new endpoint, database table, lockfile change, or CI workflow change is included. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell execution, web research, and browser tools. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cbd278dc03 |
fix(interactions): derive chat recipients and validate explicit users (#14742)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agents use saved questions to get human input and continue the same task. > - The standard question example recently told models to copy a user ID. > - A model can omit an identity prefix and create a question its intended recipient cannot answer. > - Agent Chat already knows the conversation owner, so the server can supply that identity. > - This pull request removes the blanket instruction and validates explicit recipients before saving. > - Ordinary questions stay simple, and explicit addressing remains available for decisions that need a particular person. ## Linked Issues or Issue Description Refs #14707, #14188. Related: #14238 handles legacy email recipients; this change prevents invalid recipients in new cards and retains exact ID matching. **What happened?** A model copied a Cloud user ID without its prefix into `addresseeUserId`. Creation succeeded. The intended user's answer then failed the exact recipient check. **Expected behavior** Ordinary chat questions use the saved conversation owner. A task may optionally name a specific recipient. The API rejects an unknown or unauthorized recipient before it creates a card. **Steps to reproduce** Create a chat question for a user whose ID is `paperclip-id:example`. Supply `example` as the addressee. Before this change, creation accepts the invalid recipient and the owner cannot answer. With this change, creation returns 422. Omitting the field saves the full owner ID and allows that owner to answer. ## What Changed - Remove `addresseeUserId` from standard question examples and remove the blanket requester-ID instruction. - Derive the recipient of ordinary chat questions from the persisted conversation owner. Reject conflicting explicit user IDs. - Keep explicit task recipients optional. Validate supplied user IDs with the existing board mutation policy, including company, viewer, and Cloud restrictions. - Preserve explicit agent routing, connector intents, confirmations, exact recipient checks, idempotent retries, and no-login local-board authority in local-trusted mode. - Update the blocker grader to accept an omitted recipient and verify the actual requester answered. - Add database and HTTP tests for prefixed identities, denied recipients, concurrent retries, saved answers, and response delivery. ## Verification - Database interaction service suite: 90 tests passed, including implicit local-board creation/answering and authenticated/Cloud denial. - Interaction HTTP route suite: 84 tests passed. - Affected interaction/native/connector/documentation suites: 231 tests passed across six files after valid-user fixtures were updated. - Resolver and interaction unit suites: 29 tests passed. - Product E2E unit/calibration suite: 793 tests passed; Product E2E typecheck and blocker catalog discovery passed. - Generated API-reference and capability contract checks passed. - `pnpm -r typecheck` and `pnpm build` passed. - Full local `pnpm test:run` did not finish green: its initial general-server pass had 14,416 passing assertions, one unrelated native-resume assertion failure on macOS, and three teardowns from an intermediate fixture cleanup fixed above. Separate broad local groups also encountered timeout/live-port failures under host load. Local UI (7,026), CLI (502), shared (817), and skills-catalog (20) tests passed; the complete final-head CI matrix is the broad verification gate. - After two CI cold-start readiness timeouts, a separate test-only commit gives the first exposure lifecycle fixture the existing normal 30-second readiness budget. Its real HTTP, ordering, and cleanup assertions remain intact; the targeted case and final Linux CI shard passed. Production deadlines are unchanged. - A separate OpenCode fixture failed twice on GitHub-hosted Ubuntu because its cached Node executable was group-writable; the same case passed on AWS runners. The fixture now qualifies its own Linux copy with mode `0500` and the actual copy digest. Host files and production security checks are unchanged. The focused macOS case passed; the new Linux-copy branch also passed on the final AWS-hosted Linux runner (1,125 passing Runner tests, 3 skipped). The final run was not on a GitHub-hosted runner. - Final-head [CI run 36762078176](https://github.com/paperclipai/paperclip/actions/runs/36762078176) passed for `116b968b24fa0a8c5724a7bf96e73a8dda5f0425`: 54 successful checks and two conditional Storybook skips, with no pending or failed checks. The 27 general/serialized test jobs reported 28,635 passing tests. Typecheck, build, Runner, browser E2E, and Canary gates passed. Greptile reviewed that exact head at 5/5; both review threads are resolved, with no open follow-ups. - No live provider replay is claimed by this PR. ## Risks - New explicitly addressed cards reject users who cannot mutate the issue, including viewers, inactive members, and invalid IDs. Callers that supplied invalid recipients must correct their request. - Existing addressed cards are not rewritten. Existing authorization checks remain strict. - Chat inference applies only to questions without an agent addressee. Connector intents and governed confirmations retain their own recipient paths. - No schema change or migration is required. ## Model Used OpenAI Codex, GPT-6 (exact serving variant and context window are not exposed in this environment). Used reasoning, tool use, code editing, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |