mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 16:35:27 +02:00
eefd3f22ea3038964a0eeb5f472fc6df7e5e11ff
1345
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
eefd3f22ea |
fix(dot): synchronize human assignment and preserve Unicode pages
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
155c825315 |
Merge Dot reviewer and stop-notice fixes into onboarding
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d0e7e06269 |
fix(dot): authorize admitted reviews and identify stop notices
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
80e3534cf6 |
Merge updated Dot base and current task authority into capabilities
Preserve private task filtering and task monitors, keep independent Dot opt-in, and regenerate the combined tool catalog and seeded protocol fixture. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2e40bca050 |
Merge current master into Dot provider and regenerate migration
Preserve current migration locking and native accounting behavior alongside the Dot provider. Regenerate the idempotent Dot migration after the published migration history. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5a1bdea4ca |
fix(dot): honor standalone rollout in imports and dev startup
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
72de4984a9 |
feat(dot): give Dot an independent experimental agent choice
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2626c8f9b4 |
fix(dot): enforce operator file grants and cache verified pages
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1f27dbcb5d |
feat(dot): add consented assigned-task attachment reading
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a7a244ab33 |
feat: add company decision models with permission and cost controls (#15473)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Optional product features need small, typed model decisions. > - Each company needs to choose which shared API connection pays for those decisions. > - Calls must retain user and task permissions, budget limits, and cost attribution. > - This pull request adds a managed decision service, setup UI, and request history. > - Features can check availability cheaply and keep their existing behavior when decisions are unavailable. ## Linked Issues or Issue Description **Subsystem affected** Server services, shared contracts, database accounting, Company Settings, and Costs. **Problem or motivation** Paperclip has no common decision-model service. Adding provider calls within each feature would duplicate credential access, permission checks, and billing rules. **Proposed solution** Let a connection manager configure one company decision model. Support OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test and metadata-only history. Default company-sponsored background decisions to on during setup, and preserve a saved off setting. **Alternatives considered** Per-feature credentials would duplicate existing connection management. Personal overrides and provider fallback chains add permission and billing complexity; they remain deferred. **Roadmap alignment** Reviewed ROADMAP.md and searched open PRs. This extends existing connection access and budget accounting. Product features that call the service remain outside this change. No matching decision-model service PR was found. ## What Changed - Add company settings, an internal `decisionModelService`, local availability checks, and trusted human, agent/run, and system contexts. - Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and ordered-score decisions for both providers. Bound requests and time; disable paid retries. - Add durable invocation metadata and agentless decision ledger charges. Preserve fractional cents, pricing evidence, dispatch identity, and unresolved billing holds. - Share the company accounting lock and apply company, agent, and project budgets. Settle charges once, retain unknown holds, and recover interrupted calls without resubmission. - Reuse connection setup and management UI. Add a Decisions view under Costs, production-component Storybook coverage, database migration, and service documentation. ## Verification - Passed 166 current-code tests covering the decision service/provider, setup component, Costs, OpenAPI, and every failure from the earlier broad run. Coverage includes native SDK wire formats, refusals, billed malformed responses, permission and secret-rotation races, identity changes, concurrent budget admission, unresolved holds, agent/task deletion, and stale setup feedback. - Passed 170 existing connection, cost, budget, heartbeat-accounting, and profile regression tests. - Passed repository typecheck, production build, and design token gates after integrating master. Verified the generated migration on a fresh test database and upgraded the populated preview database from the branch's earlier migration without losing settings or usage. - Ran the required full `pnpm test:run`: its general phase completed with 16,354 passed and 10 failures across five files while this branch was still being updated. Every reported failure passes in the current-code rerun; the serialized phase did not run after that failure. The full GitHub CI suite passed on `b1b856a88`: general and serialized tests, browser shards, runner checks, typecheck, production build, packaging/canary, and policy gates. [CI evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360). - Passed the full-shell Storybook setup-to-history interaction test again after integrating master. Greptile rates the final revision 5/5 with zero unresolved threads. - Walked through the running app: empty setup, add each provider, save, reload, run all three sample questions, inspect fractional charges in history, switch provider while sponsorship is off, disable, and reconnect. Checked mobile settings. These tests used the actual UI, vault, server, SDKs, and database with simulated upstream responses. - Live paid setup tests remain unverified: this environment has no authorized OpenAI/OpenRouter credentials available. No mocked test is presented as live provider evidence. Reviewer journey: Company Settings → General → Decision model. Add/select a shared API connection, save, run the billed sample, open View usage, then disable decisions and verify Run test is disabled after reload. ## Risks - The SDK decision interface is experimental. Pinned versions and wire-format tests limit upgrade drift. - The migration allows agentless service charges and reservations. Existing agent cost-reporting APIs still require an agent, and decision receipts stay separate from run reconciliation. - Timeouts can have unknown provider charges. Holds remain until an audited accounting correction resolves them. - OpenAI prices use a versioned Decisions rate snapshot; OpenRouter costs use provider receipts. Unknown pricing is retained as unknown. - A configured company authorizes background spending by default. Setup explains this, and managers can turn it off. - Live provider account/model availability still needs the two credentialed acceptance checks. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser tools. The exact serving revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1c44b7b56 |
Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs. Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
50b1f95e79 |
fix(tool-gateway): keep MCP connection healthy on oversized/malformed responses (#15462)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents call remote MCP tools through the Paperclip tool gateway > - The gateway keeps a health status for each remote MCP connection, and it hides the tools of a connection that has `healthStatus = "error"` > - One oversized, malformed, or invalid-JSON `tools/call` reply set the whole connection to `error` > - Health goes back to `ok` only after a successful call, but the hidden tools prevent that call, so the connection stayed locked until a person reconnected it > - This pull request makes these reply errors fail only the one call, and it starts a new MCP session for the caller on the next call > - The benefit is that one large or bad reply no longer disconnects a working connection for every agent ## Linked Issues or Issue Description No public issue exists. Related open PRs fix other health downgrades in the same function. They do not overlap with this change: - Refs #11910 (timeouts and JSON-RPC errors) - Refs #15325 (transport errors) **What happened?** An agent called a remote MCP tool that returned more than `MAX_REMOTE_MCP_RESPONSE_BYTES` (1 MB). The gateway returned 502 `mcp_remote_response_too_large`. It also set the connection to `healthStatus = "error"`. Then `connectedMcpConnectionFilter` hid all tools of the connection (except `per_user` connections), and `connections_search` showed the connection as `needs_user_action`. The `malformed_response` and `invalid_json` errors from `readMcpHttpResponse` did the same thing. `readMcpHttpResponse` can also cancel the reply stream before the end. The remote server can then close its MCP session. The gateway kept the cached `mcp-session-id` for up to 30 minutes and cleared it only after a 404. The next calls then failed with "Remote MCP session expired". **Expected behavior** A reply that is too large or not correct fails only that call. The connection stays healthy, and its tools stay visible. The next call uses a new MCP session. **Steps to reproduce** 1. Add a remote MCP connection with `mcpSessionRequired: true`. 2. Call a tool that returns a reply larger than 1 MB. 3. Look at the connection: `healthStatus` is `error`. 4. Start a new gateway session: the tools of the connection are not in the list. **Paperclip version or commit** `master` at `99a9de9940bf5974352d9dbfbb2f21e62e89689f` **Deployment mode** All modes. The fault is in the server tool gateway. ## What Changed - `server/src/services/tool-gateway.ts`: `too_large`, `malformed_response`, and `invalid_json` from `readMcpHttpResponse` now fail only the call. The caller gets the same 502 reason code as before. The gateway does not call `markRemoteConnectionHealth(…, "error")` for these errors. The `invalid_json` branch after `JSON.parse(body)` also does not change health now, so all `invalid_json` paths are the same. - The `mcp_remote_response_too_large` message now tells the caller to request a smaller result, for example a narrower query or a smaller page size. The error details now include `maxBytes`. - `server/src/services/mcp-http.ts`: new `forgetMcpHttpSession()`. It removes only the cached session for one scope and one credential set, and only while that entry still holds the session ID that failed. It does not touch other agents, sessions that are initializing, or a newer session that replaced the failed one. - The gateway calls `forgetMcpHttpSession()` after these reply errors, so the next call from that caller initializes a new session. The existing 404 "session expired" path now uses the same function. Before, both paths cleared all sessions and all pending initializations on the connection. That made a concurrent initialization by another agent fail with "MCP connection changed while initializing", which the gateway reported as a fetch failure and marked as a connection `error`. - `forgetMcpHttpSessions(connectionId)` is not changed. Disconnect and revocation still clear the full connection. Why `malformed_response` and `invalid_json` are also per-call errors: - Each error is about one reply body. The server was reachable, accepted the credentials, and sent HTTP 2xx. A reply can be too large or bad because of the tool and its arguments. That is not a fault of the connection. - The other malformed-reply checks in the same function (payload is not an object, or has no `result`) already throw `remote_mcp_malformed_response` and do not change health. Only the reader-level errors changed health. - Health recovers only after a successful call. An `error` status for one bad reply therefore locks the connection until a person reconnects it. - Other failures (HTTP errors, fetch failures, timeouts, JSON-RPC errors) still change health. This PR does not change them. #11910 and #15325 address some of them. ## Verification - `server/src/__tests__/tool-gateway.test.ts`: the recovery test now runs for three replies: oversized, invalid JSON, and a reply without the requested message ID. Each run uses a fake HTTP MCP server with `mcpSessionRequired: true` that gives `session-N` for each `initialize`. Each run checks that: - the call fails with the correct 502 reason code (and, for the oversized reply, the smaller-page hint and `maxBytes: 1000000`) - `healthStatus` stays `ok` - a new gateway session still lists the tool and can call it - the second `tools/call` uses `session-2`, not `session-1` - New test: agent A gets an oversized reply while agent B initializes a session on the same connection. Agent B's call completes, health stays `ok`, and the two calls use `session-1` and `session-2`. With the old connection-wide reset, this test fails: agent B gets 502 `mcp_remote_fetch_failed`. - `server/src/__tests__/remote-mcp-protocol.test.ts`: new unit test. `forgetMcpHttpSession()` keeps the session of a different identity, and a late failure from an old session does not remove the newer session. - Each regression check fails without its fix. Without the health change, the test fails on `healthStatus: 'error'`. Without the session reset, it fails with `['session-1', 'session-1']`. ``` cd server npx vitest run src/__tests__/tool-gateway.test.ts src/__tests__/tool-gateway-service.test.ts src/__tests__/remote-mcp-protocol.test.ts Test Files 3 passed (3) Tests 132 passed (132) ``` - I did not run the full server `tsc --noEmit` locally because the sandbox does not have sufficient memory. The CI typecheck covers it. ## Risks - Low risk. The change affects only the error path of remote MCP `tools/call`. - A remote server that always sends bad replies now keeps `healthStatus = "ok"`. Each call still fails with a clear 502 reason code and an audit record, so the failure stays visible. Only the connection-wide hiding of tools stops. - After one of these errors, the next call from the same caller sends one more `initialize` request. This adds one round trip. - A 404 "session expired" now clears only the session of the caller that got the 404. Before, it cleared the cached sessions of all agents on the connection. If the remote server restarts, each agent now gets its own "session expired" error one time and then initializes again. A 404 is about one session, and the old connection-wide clear also cancelled other agents' initializations. - #11910 and #15325 change the same catch block. The PR that merges last can have a small merge conflict. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, 1M-token context window. - Run as an agent in Claude Code with tool use (shell, file edit, and local test runs). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing documentation changes are necessary) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge This PR replaces #15418. It keeps the same commits on a branch name without an internal ticket ID, and adds a fix for the Greptile review on #15418. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
fd8c6b920a |
fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native tasks can pause while a human decides whether to allow a connection action. > - The next turn needs both the saved decision and complete accounting for the previous turn. > - Cancellation could discard final usage, and complete direct Claude API receipts could remain unpriced. > - A stale blocked or review report could also request approval again after the original card was declined. > - This pull request retains shutdown accounting and rejects approval waits bound to an already resolved action. > - The benefit is reliable continuation with the existing budget and approval controls. ## Linked Issues or Issue Description Refs #15420. Related: #15312 addresses requester ownership during dispatch. This change addresses receipt capture and final-response validation. ## What Changed - Retain usage events after native cancellation. Continue to reject late provider messages and work. - Drain same-turn accounting and terminal events for at most fifteen seconds after a durable governed wait. Keep incomplete accounting blocked. - Estimate complete, unpriced, direct Anthropic API receipts for the exact `claude-sonnet-5` model. Record the rate version and assumptions. Use the one-hour cache-write rate when the receipt lacks cache TTL. - Bind stale approval reports to exact interaction, action-request, or invocation IDs in the same company, task, agent, and run. Cover blocked, review, and response-wake reports. Keep independent reviews valid. - Fence checkpoint and result writes after a controller detaches for restart, including operations waiting for a database lock. Reject stale successful returns before certifying accounting. - Allow bounded subscription teardown only after retaining an actual provider terminal. - Journal the exact governed-wait trigger and disposition before provider interruption. Recover that wait independently of a later saved answer, replay retained accounting, and reject mismatched or unproven terminal evidence. - Give settling governed turns a bounded window before shutdown detaches their controller. - Keep fuzzy external app matches alongside installed capability matches instead of forcing an unrelated provider question for a generic query. - Return up to twenty exact active catalog tool names after an invalid request, after eligibility checks; still reject the request without granting access or creating an approval. - Require retained provider terminal proof before settling a governed wait, including when complete usage arrives before stream closure/error/timeout. Retain harmless numbered cancellation events so restart replay stays contiguous. - Isolate accounting-test OpenCode config from the host plugin directory. - Add regression coverage and document the accounting, restart and connection-search behavior. ## Verification Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`; later review fixes are verified separately and do not relabel those runs. - Repository typecheck and build pass. - Current runner runtime and cancellation suites: 191 tests pass. Six new regressions cover stream end/error/timeout without provider-stop proof and contiguous cancellation acknowledgement/request replay; all six failed before the fix. The existing bounded cleanup case now explicitly supplies terminal proof. All 31 adapter accounting tests pass with isolated fixture config. - New checkpoint-rebinding and approval-criterion suites: 65 tests pass; runner HTTP integration: 29 tests pass. Unchanged executor/control-plane suites: 640 tests pass; database-backed connection suites: 68 tests pass. - Current retained evidence verification covers 222 file hashes across all fifteen original result artifacts. The three-case [published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html) and all eight screenshot hashes verify. The [three-case recovery report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html) and all six screenshots also verify; the nine-result campaign did not publish. - Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node tests pass. Eval typecheck and catalog discovery pass. - Full local repository unit run was interrupted before the follow-up edits after three tool-access failures and one runner HTTP failure. Those failures pass in isolation; the 09a full run was interrupted after one rapid Slack callback-ordering failure and seven skill-service failures. All eight pass both isolated and with full-runner environment settings, and all 74 skill-service tests pass together; the subsequent full run reported two 15-second OpenCode accounting timeouts and was stopped with exit 130 to apply review fixes. The timeouts reproduce while copying this host’s 61 MB OpenCode config. All 31 tests pass after isolating config inside each fixture without increasing timeouts or changing assertions. A complete local full-suite pass is not claimed. Repository-wide CI also passes on the final review-fix head in [run 37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952). Earlier database-skipped diagnostics and the older ENFILE run are retained and are not full-suite passing evidence. - Three-case live campaign [37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525) passes all three original grades on frozen source 09a: 55/55 checks, eight succeeded run records, complete accounting receipts, matching checkpoint identities and no pending approvals. Campaign [37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353) adds nine original passes (173/173 checks, eighteen succeeded records) on the identical source. Its other three jobs failed before runner assignment or any step while GitHub could not load the paid environment; those original infrastructure failures are retained. Campaign [37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356) completes only those unstarted cells: all three original grades pass (57/57 checks, six succeeded run records). All fifteen exact cases now pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new overall failures and zero pending pairs. Total current evidence: 285/285 checks and thirty-two succeeded run records, complete accounting receipts, matching checkpoint identities, no pending approvals or retry records. Earlier campaigns retain forty-seven additional run records and two known same-run recovery attempts; actual provider-call counts and invoices remain unknown. The nine-result campaign skipped publication and its public URL returns 403; original artifacts remain retained. Previous ba9 campaign [37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840) completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline pass, while OpenCode resumes but times out searching for exact tool names and never creates the access card. That failure and incomplete cancelled-run accounting remain preserved; the new catalog error guidance targets this observed dead end. Campaign [37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312) remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker leak and unknown-criterion approval gap. The original baseline remains 7 PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS / 3 FAIL. All fifteen selected cases are qualified by their original grades in this bounded trial. These live grades belong to 09a. Its twelve governed-wait checkpoints retain matching same-turn terminal fingerprints, but passing artifacts omit detailed event journals; the later six adversarial regressions qualify the new terminal-proof and replay guards separately. Final-head repository CI passes. [Fresh Greptile review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780) is 5/5, confirms both findings are fixed, and reports no new actionable issues. All review threads are resolved and the PR has no merge conflicts. ## Risks - Governed cancellation drains accounting for up to fifteen seconds. Restart detachment gives a settling batch up to twenty seconds to finish. An incomplete receipt or unproven provider terminal still prevents successful qualification. - Claude prices are estimates, not invoices. The estimate assumes standard global API pricing and uses a conservative cache-write rate. Unsupported models, billers, and billing modes remain unpriced. - Approval identity matching must remain scoped to the current run and the requested approval. It does not authorize execution of a declined call. - Original baseline and final candidate have different merged master context. Exact-case outcomes are before/after observations, not isolated causal attribution to this repair. - No schema migration, fixture, oracle or grader change. Search-result guidance now treats fuzzy external matches as suggestions. Existing app authorization and provider-consent checks remain required. ## Model Used - OpenAI GPT-6 through Codex, with code editing, terminal tools, and test execution. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (191 current runtime tests and 31 accounting tests; interrupted full-suite history and CI coverage are disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5405b1f46f |
Synchronize chat delivery tests with completed drains (#15480)
Expose the already tracked and caught deferred-work promise to test schedulers while preserving default setImmediate behavior. Assert Slack delivery ordering while the primary lease is held, then join the drains; update the existing GitHub test to assert the resulting quiescent queue. Verified all 1,063 chat tests, full typecheck/build, negative synchronization control, independent review and exact-head CI. An unchanged browser shard retry passed and is documented. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
465140596f |
Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials. Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2a24849422 |
fix(dot): settle read failures and serialize workspace calls
Sync dependency locks and pairing API contracts. Return admitted idle assignments and preserve stable polling while work is queued. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
128c95837b |
Retain bounded orphan process-loss diagnostics (#15475)
Record bounded observer uptime, run and output ages, process-check observations and retry eligibility before orphan cleanup changes the evidence. Preserve existing recovery and reporting behavior and omit process IDs, raw paths and credentials. Verified focused helper, actual reaper and real SDK regressions, full typecheck/build, exact-head CI and independent review. An unchanged Cursor timeout passed its isolated retry; unrelated local Slack timing failure is documented separately. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
43b0aff54b |
Keep pre-dispatch secret configuration blockers out of Sentry (#15472)
Keep confirmed pre-dispatch missing-secret setup blockers out of Sentry while preserving failed runs and owner recovery actions. Runtime, provider and ambiguous failures remain reportable. Verified full exact-head CI, focused local regressions, workspace typecheck and independent review; Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ae6f95ed7a |
feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults. Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ebc99db2f5 |
feat(dot): expand governed Runner capabilities and finish pairing setup
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
06484b3c41 |
fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Recovery schedules bounded retries after a provider disconnects. > - A stopped run can still hold its environment lease while cleanup runs. > - Retrying before that lease is released cancels the new run before it starts and spends another retry. > - Restoring the task can also send the worker repair instructions from an already resolved recovery action. > - This pull request preserves the waiting retry and removes settled recovery instructions from later wakes. > - The task can continue after cleanup without an operator repairing the same incident again. ## Linked Issues or Issue Description **What happened?** A legacy conversation run disconnected while it was doing ordinary work. Its environment cleanup took longer than the retry delay. Two retries were cancelled before dispatch with `execution_reconciliation_required`. Those cancellations exhausted the failure budget. After an operator restored the task, the wake still told the original worker to repair the runtime and hand the task back to itself. **Expected behavior** Cleanup waits preserve the pending attempt. The same retry can continue after ownership is released, subject to all current gates. Once a recovery action is resolved or cancelled, subsequent task wakes omit its repair instructions. **Steps to reproduce** 1. Fail a legacy conversation run while its environment lease remains in `pending_cleanup`. 2. Schedule a bounded retry and run promotion before cleanup releases that lease. 3. Repeat the scheduler sweep. Before this fix, retries promote and then cancel without starting. 4. Resolve a stranded-task recovery action and build the restored task wake with that action ID. Before this fix, the wake still includes the settled repair instructions. **Paperclip version or commit** Reproduced with database regressions against `ceabc3bc880` on master. Related: #15019 restores a skipped assignment handoff after lease release. #15235 filters stale handoff evidence in the recovery sweep. This change preserves an existing scheduled conversation retry and corrects restored wake content. It does not create a new handoff wake. ## What Changed - Keep an unstarted legacy conversation retry on the same durable row while prior execution ownership remains active. Recheck after 30 seconds without increasing retry accounting. - Return a queued retry to scheduled state if it encounters that hold at the claim gate. Retain its issue claim and publish the status change. - Record one local lifecycle diagnostic per blocking run. Remove that wait marker on promotion. - Include recovery action metadata only while the referenced action is active or escalated. - Return an explicit `waiting` response and the saved schedule when Retry now meets cleanup. Show the wait inline without a false success or disabled button. - Add database, rendered-prompt, route, and UI regressions. Document the execution and run-log contracts. ## Verification - Five cleanup and restored-wake regressions fail against the original production code. The Retry now route and UI regressions also fail before their correction. - Related retry, dispatch, stale-queue, and recovery suites: 321 tests pass across seven files. All ten focused cleanup/restored-wake cases pass after rebase. The final dispatch adjustment passes all 46 adapter tests. - Retry now routes and affected UI suites: all 46 tests pass. The tests cover repeated clicks, the saved schedule, unchanged accounting, promotion after release, and no false success or error state. - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm check:token-gates` pass. - Full local `pnpm test:run` was started and then stopped after the final commit passed all GitHub CI test shards. No complete local full-suite result is claimed; CI supplies the complete test result for the final commit. - Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub checks are green or intentionally skipped. Apex review is 5/5 after two reviews, with no unresolved threads. The PR has no merge conflicts. ## Risks The wait applies only to unstarted legacy conversation retries. Native runs and non-conversation execution keep their existing recovery rules. Cleanup must actually release ownership before execution can resume. The wait does not fix a cleanup service that never finishes. Promotion and dispatch still enforce cancellation, reassignment, pause, budget, and reconciliation gates. No schema change is required. The Retry now response adds a `waiting` outcome; the shared contract and all three UI controls handle it. ## Model Used OpenAI Codex based on GPT-6, with repository analysis, tool use, and local code execution. The runtime does not expose the precise serving model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; complete final-head test coverage in CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9fb955e9a2 |
fix(codex): gate models on the Codex CLI floor before the ChatGPT backend rejects them (#15399)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Codex adapter and the native runner run Codex in local hosts and in managed sandboxes, and the model catalog lists the models an operator can select > - The Codex backend accepts a new model with ChatGPT sign-in only from a recent enough Codex CLI. An older CLI fails every turn with `The '<model>' model is not supported when using Codex with a ChatGPT account.` > - After #14942 added `gpt-6.1-sol`, an operator selected it for an agent in a managed Daytona sandbox. The sandbox image still shipped Codex 0.156.0. The connection test and runs failed with that sentence, which reads like an account problem > - Paperclip only checked the Codex compatibility window (`>=0.149.0 <0.161.0`), so 0.156.0 passed and the failure surfaced from the backend without a cause > - This pull request records the verified Codex CLI floor for each gated model, compares the installed `codex --version` with that floor in the environment Test and in the remote runner, and names the backend rejection when it still happens > - The benefit is a precise, actionable message ("gpt-6.1-sol requires Codex CLI 0.159.0 or newer; detected 0.156.0; promote a sandbox image with Codex 0.160.0") instead of an opaque 400, and no doomed hello probe ## Linked Issues or Issue Description Refs #14942 **Bug: selecting `gpt-6.1-sol` in a managed sandbox fails with a ChatGPT account error.** **Steps to reproduce** 1. Promote a sandbox image that ships Codex CLI 0.156.0. 2. Create a Codex agent with a ChatGPT sign-in connection, select `gpt-6.1-sol`, and run the environment Test or a task in that sandbox. **Expected behavior** The Test or the run tells the operator that the Codex CLI in the sandbox is older than the model needs. **Actual behavior** The Test reports `codex_hello_probe_failed` with `{"type":"error","status":400,"error":{"type":"invalid_request_error","message":"The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account."}}`. A run fails with "The selected model is not supported by the current ChatGPT connection." **Evidence** - Codex 0.156.1 and older are rejected for `gpt-6.1-sol` with ChatGPT sign-in: https://github.com/openai/codex/issues/49396 - Codex 0.159.0 and 0.159.2 are accepted: https://github.com/openai/codex/issues/49464 and https://github.com/decolua/9router/issues/4471 (same error, fixed by raising the client version identity from 0.155.0 to 0.159.0) - Codex 0.157.0 release notes: "Add GPT-6 Sol and Luna to the model catalog": https://github.com/openai/codex/releases/tag/rust-v0.157.0 - GPT-6.1 Sol is included for Plus, Pro, Business, Enterprise, and Edu with ChatGPT sign-in: https://learn.chatgpt.com/docs/models ## What Changed - `packages/adapters/codex-local/src/index.ts`: add `minimumCodexCliVersionForModel` (`gpt-6.1-sol` → 0.159.0, `gpt-6-sol` and `gpt-6-luna` → 0.157.0, with the sources above), `parseCodexCliVersionOutput`, `codexCliVersionAtLeast`, and `CODEX_CHATGPT_MODEL_REJECTION_RE`. Models without a verified floor return `null`, so nothing probes them. The agent configuration doc describes the floors. - `packages/adapters/codex-local/src/server/cli-version.ts` (new): run `codex --version` where the run would execute it (local, SSH, or sandbox), compare it with the model floor, and build the `codex_cli_version_compatible` / `codex_cli_version_incompatible` checks. Map the backend rejection sentence to `codex_hello_probe_model_rejected` with the detected CLI version and a hint that separates a stale CLI from a plan that does not include the model. - `packages/adapters/codex-local/src/server/test.ts` (CLI lane Test): run the version check before the hello probe when the model has a floor. Skip the hello probe when the CLI is too old (`codex_hello_probe_skipped_cli_version`). When the probe still fails with the backend sentence, report `codex_hello_probe_model_rejected` instead of the generic `codex_hello_probe_failed`. The new codes do not match the auth-failure patterns, so a managed connection is not invalidated. - `packages/adapters/codex-local/src/server/acp.ts` (ACP lane Test): for remote targets, run the same version check against the shared `codex` the ACP server spawns. - `server/src/services/native-runtime/native-session-executor.ts`: after the compatibility-window check, compare the remote Codex with the configured model's floor and fail with `runner_remote_provider_artifact_incompatible: <model> requires Codex <floor> or newer with ChatGPT sign-in, received <version> from the sandbox image; promote a sandbox image with Codex 0.160.0 or configure PAPERCLIP_RUNNER_REMOTE_CODEX_NPM_SPEC=...`. When a preinstalled Codex fails this check and an npm spec is configured, the existing fallback installs the pinned release. - Tests: adapter metadata, CLI-lane Test (too old, compatible, no floor, backend rejection), ACP-lane Test (too old, compatible, no floor), and the remote runner floor (nine version/model cases). - `doc/adapter-model-audit-2026-10-02.md`: record the floors and the reason. No catalog entry, default model, saved agent configuration, pin, lockfile, or image definition changes. ## Verification Local (Node 25.9, pnpm 9.15.4 via corepack): ```sh pnpm --filter @paperclipai/adapter-codex-local typecheck pnpm --filter @paperclipai/adapter-codex-local exec vitest run # 81 passed pnpm --filter @paperclipai/server exec vitest run \ src/services/native-runtime/native-session-executor.test.ts \ src/services/native-runtime/codex-runtime-compatibility.test.ts -t Codex # 84 passed ``` Also run locally: the full `native-session-executor.test.ts` file (533 passed) and the full Codex adapter suite (81 passed). Follow-up commit `9fd4d3b25` (Greptile P2: the ACP-lane version probe ignored the agent's configured env): the probe now receives the adapter's string-valued `env` entries, the same ones `buildCodexAcpConfig` hands to remote ACP runs, so a `PATH` override selects the same `codex` for the Test as for the run. New test covers a `PATH` + `CODEX_HOME` override and a dropped non-string entry. Re-run on that head: `pnpm --filter @paperclipai/adapter-codex-local typecheck` clean; full Codex adapter suite 503 passed (31 files). Not run here: `pnpm -r typecheck` for `server` (the direct `tsc --noEmit` was killed by the sandbox memory cap; the adapter package typecheck passes and CI covers the server), `pnpm build`, browser suites, and a live sandbox probe. No UI files changed, so token gates do not apply. Manual check for a reviewer: set an agent's Codex model to `gpt-6.1-sol`, point its environment at a sandbox whose `codex --version` prints `codex-cli 0.156.0`, and run the environment Test. The result contains `codex_cli_version_incompatible` with `Detected Codex CLI 0.156.0.` and no hello probe check. With Codex 0.160.0 the Test contains `codex_cli_version_compatible` and runs the hello probe. ## Risks - A floor is enforced for all authentication modes, because the runner does not know the credential method at verification time. An OpenAI API key on Codex 0.157.0 or 0.158.0 with `gpt-6.1-sol` is now rejected before launch. Paperclip pins Codex 0.160.0 everywhere, so this only affects installs that lag the pin, and the message names the fix. - The 0.159.0 floor for `gpt-6.1-sol` is the oldest stable release verified to work; 0.157.0 and 0.158.0 were not verified either way. If OpenAI accepts an older client, the floor can be lowered in one table. - The `codex --version` probe adds one short process run to the Test only for models with a floor. - Rollback: revert this pull request. No data or configuration migrates. ## Model Used Claude Fable 5.1 (`claude-fable-5-1`), Anthropic, with extended thinking, tool use, and web research, operating as a Paperclip agent through Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Bender (Fable) <noreply@paperclip.ing> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
f669194298 |
fix(runner): propagate configured environment to future turns (#15451)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runners start provider processes and enforce a separate tool environment policy. > - The server resolves task environment bindings at each run boundary. > - Fixed launch allowlists dropped custom variables after resolution. > - Warm process reuse and a blanket fingerprint exclusion also hid configuration changes. > - This pull request carries a bounded, server-selected list of task variable names through each process boundary. > - The benefit is that later turns and tool commands receive configured values while ambient host secrets stay excluded. ## Linked Issues or Issue Description **What happened?** A native Codex agent could not use configured task credentials. Adding the variables in Settings did not fix the next turn. The provider sanitizer, runner launch, Rust provider launch, and shell policy each used fixed allowlists. Effective config fingerprints also excluded all variables with the `PAPERCLIP_` prefix. **Expected behavior** Explicitly bound task variables must reach provider processes and tool commands. Added, changed, and removed values must take effect at the next run boundary. A running turn keeps its original configuration. **Steps to reproduce** 1. Start a native Codex task without a custom environment binding. 2. Add a fake `PAPERCLIP_PAGE_BUCKET` value and a fake Pages credential binding in Settings. 3. Continue the task and inspect the tool environment. 4. Before this fix, those values are absent even from a fresh runner launch. **Paperclip version or commit** Reproduced at `b31558064`. The fix is rebased on current master. **Deployment mode** Self-hosted server with a native process runner. Related changes: [the legacy Codex MCP environment fix](https://github.com/paperclipai/paperclip/pull/13321) and [ambient server-secret exclusion](https://github.com/paperclipai/paperclip/pull/12870). These affect different launch paths. This change preserves their credential boundaries. ## What Changed - Capture scoped task bindings after resolution, before managed provider credential injection. Mint and validate the names-only projection at native dispatch before host inheritance. Legacy adapters retain their previous environment limits. - Strip user-supplied projection markers from agent, environment, project, and routine config. - Carry selected values through the Codex, ACPX, OpenCode, runnerd, and Rust subprocess launch boundaries. - Add selected names to native and ACPX Codex shell include lists. Keep selected values out of command arguments, including selected bootstrap values. - Reject malformed projections, reserved authority and loader names, missing values, null bytes, and oversized input. - Replace a retained native process when projected values change. Keep unchanged processes reusable. - Fingerprint custom namespaced variables while excluding known generated runtime variables. - Document next-run behavior and add regression coverage. ## Verification - Red: the permanent reproduction failed at four launch/tool boundaries and the namespaced fingerprint check. Two control checks passed. - Green: the initial regression plus existing Codex environment and shell tests passed (39 tests). - Server config resolution, fingerprints, and native-session suites passed (653 tests), including addition, rotation, removal, and unchanged warm-session reuse. - Rust regression tests passed. A real shell child received added and rotated values, then lost them after removal. Unselected host variables stayed absent. - `pnpm -r typecheck` and `pnpm build` passed after rebase. The full `pnpm test:run` was attempted but could not complete: fresh embedded PostgreSQL databases fail during bootstrap on this macOS host. An isolated suite and a disposable native `initdb` probe reproduced the failure before test execution. `shmget` reports `No space left on device` because the host has exhausted shared-memory IDs. This is not disk exhaustion. The run was stopped after confirming the external setup failure. CI results will be recorded separately. - Runner boundary suites: 361 tests passed. Two process-launch errors during concurrent binary staging passed on isolated rerun. - Rust Codex provider and process supervisor integration suites: 97 passed, 2 intentionally ignored. - ACPX shell and selected-bootstrap argv regressions: 3 failed before the fix, then all 55 relevant tests passed. - Review compatibility regression: 129-variable and large-value legacy configurations failed before the correction and passed after moving native-only validation to dispatch. - Final review head `3f11e8d25`: repository typechecks and production build passed. CI completed its implementation checks; one general-server shard hit SQL `40P01` in `heartbeat-runtime-skills.test.ts` during its `beforeEach` table truncate (1,199 tests passed in that shard). The shard passed on its single rerun. All CI checks for this head are green. Greptile reviewed this head at 5/5 with no new actionable findings; the compatibility thread is resolved. - No live provider credentials or model calls are required by these tests. ## Risks - Configured task credentials now reach the tools they were configured for. The controller selects names only after existing scope and secret-binding authorization. - TypeScript and Rust validate the same bounded projection. Their reserved-name rules must stay aligned. - A changed projection replaces an idle provider process. Unchanged values preserve reuse. Active turns retain their original environment. - No database migration or API schema change. > This fixes existing runner configuration behavior. It does not add a new roadmap capability. ## Model Used - OpenAI GPT-6 (Codex), with reasoning, local code execution, and repository tools. The session does not expose a more specific API model identifier or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1640c5b6ab |
feat(ui): render HTML artifacts in a secure sandbox (#15447)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents deliver reports as task attachments and workspace files. > - The board already renders Markdown, but it shows HTML reports as source or rejects their preview. > - Reports can need inline scripts to build charts and tables. > - Artifact scripts must not read board cookies, storage, or the parent page. > - This pull request adds an opaque-origin HTML preview and view controls beside Download. > - Users can explore a report, inspect its source, and download the original file. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task attachment panel and workspace file viewers. **Subsystem affected** ui/ and the server workspace file preview service. **Current behavior** HTML attachments show source text. Workspace HTML files cannot be previewed. **Proposed behavior** Render self-contained HTML reports in a sandboxed iframe. Place eye and code controls beside Download. Preserve the original file for raw view and download. **Reason and benefit** Users can read and filter agent reports in Paperclip without giving artifact scripts access to the board session. **Breaking changes** HTML previews now open rendered. Scripts and styles must be embedded in the report. Remote resources, API requests, forms, popups, and host navigation are blocked. Workspace HTML content uses the existing bounded UTF-8 JSON response. **Additional context** Related closed proposals: #3293 and #4857. This change uses the current attachment and workspace viewers and adds no report-serving endpoint. The Artifacts and Work Products roadmap item is complete; this improves its existing preview behavior. ## What Changed - Add a shared HTML iframe renderer with `sandbox="allow-scripts"` and no `allow-same-origin`. - Install a restrictive CSP before artifact markup. Keep inline report scripts and styles. - Add shared rendered/raw icon controls beside Download in the attachment panel, workspace panel, and file sheet. - Return workspace HTML as bounded text inside JSON. Keep download and path-access protections. - Add Storybooks for reports, security probes, workspace viewers, a task journey, and mobile layouts. - Document the security boundary and browser test command. ## Verification - `pnpm -r typecheck` and `pnpm build` pass. - Token gates and the static Storybook build pass. - Four browser tests pass. They cover report filters, raw mode, task entry, workspace viewers, cookies, storage, host DOM, resource requests, forms, popups, and host navigation. - The focused UI suite passes 26 tests. The file-resource server suite passes 36 tests. - Local full-stack acceptance passes with an actual 52 KB HTML report. Its chart and filter render. Raw mode preserves the source. Download bytes match the original file. - The complete local UI suite passes: 698 files and 7,754 tests. - `pnpm test:run` was attempted. Its server group recorded two unrelated timeouts and eight connector failures. All failed cases pass in isolated reruns, including the complete 388-test connector suite. The remaining local run was stopped after the full CI test coverage passed, to release test database resources. The full local command did not complete cleanly. - The separate local CLI group passes 511 tests. Three worktree database cases fail to start embedded PostgreSQL because this Mac has exhausted its shared-memory allocation. One isolated rerun fails at the same database startup step. No CLI code was changed. The corresponding CI test jobs pass. - All 56 PR checks are clean on `c771255f1a2e86bd8825f22eafb5d14a5e342cce`. Greptile gives 5/5 with no actionable findings. The security scan passes. The branch has no merge conflict. - Review the **HTML artifacts** section in Storybook. Run `pnpm exec playwright test --config tests/html-preview/playwright.config.ts` for the browser checks. ## Risks - Never add `allow-same-origin` to this iframe. Its opaque origin is the cookie, storage, and host-DOM boundary. - Reports that depend on CDN assets need to embed their dependencies. - The frame can navigate itself. Sandbox restrictions remain after navigation. This is not full network isolation. - Inline scripts can consume browser resources. This change does not isolate CPU or memory use. - No database migrations or API shape changes are required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository editing, code execution, and browser tool use. The runtime does not expose a more specific model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — feature and UI checks pass; the full local run has the infrastructure limits recorded above. All CI test jobs pass. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cfedd7df8 |
fix(runner): restore task monitors and durable timed waits (#15446)
## Thinking Path > - Paperclip manages AI agents and their task execution. > - Agents need a durable way to return to work after a delayed check. > - The issue monitor scheduler already provides a one-shot wake for an assignee. > - Native runners reject generic execution-policy writes and had no bound monitor tool. > - Scheduling alone is insufficient because native completion also needs to accept a timed wait. > - This pull request adds an authorized monitor tool and connects it to completion and the existing scheduler. > - An agent can now schedule its next check, end the run, and resume on the same task. ## Linked Issues or Issue Description **What happened?** A native runner could not set its own task monitor. `call_api` correctly rejected execution-policy writes, while `schedule_wake` had no production binding. `paperclip_finish` also rejected monitor waits. **Expected behavior** A standard native run can set a one-shot monitor on its current task or another accessible task assigned to the same agent. After a confirmed schedule on the current task, it can yield. The scheduler later delivers `issue_monitor_due`. **Steps to reproduce** 1. Start a standard native task. 2. Ask the agent to check the task again later and end its current run. 3. Inspect available tools and try the generic issue execution-policy update. 4. Observe the missing native tool and the lifecycle-write denial. Related PRs: #14680 concerns monitor notes in the shared wake prompt. #11919 changes attempt-limit scope. This PR adds native scheduling and completion authority and retains the existing cumulative attempt bounds. It does not depend on either PR. ## What Changed - Add provider-neutral `set_task_monitor` with a default current-task target, future timestamp, required notes, existing bounds, and explicit clearing. - Check company, task visibility, ownership, runtime permissions, work mode, and active-run authority. Preserve review-only restrictions. Reject the reserved server-owned quota-recovery name before saving or accepting a native wait. - Commit the monitor, audit event, and retry receipt together. Retry receipts survive a successor run without re-arming cleared or consumed timers. - Permit `paperclip_finish` to yield to a persisted monitor. Recheck ownership and the schedule when committing final disposition. Release execution without an immediate continuation. - Preserve due monitors during native execution. Fence wake admission and consumption against replacement, clearing, reassignment, and completion. Preserve unrelated review policy. - Expose scheduled and consumed monitor instructions in task context. Update provider schemas, Rust validation, generated contracts, and execution documentation. - Add an opt-in live Codex smoke script with isolated data and explicit run/session/runner/process evidence. ## Verification - Repository `pnpm -r typecheck` and `pnpm build` passed after rebase. Server typecheck passed again after review fixes. All CI test shards pass on `3def77b1b`, including runner TypeScript/Rust, server, serialized server, workspace, and browser tests. All CI gates are green, including the canary dry run. Greptile is 5/5 on the same commit with zero unresolved threads. - The local monolithic `pnpm test:run`, started before the rebase, was interrupted after current-head CI test coverage passed. It is not counted as a standalone full-suite pass; the focused local regression suites passed. - Targeted server tests cover scheduling, replacement, clearing, policy preservation, cumulative bounds, cross-run retries, permissions, provider-neutral discovery, review restrictions, completion authority, and scheduler/finalizer races. - Runner contract/catalog/semantic tests and Rust terminal-tool tests cover the new operation and monitor completion. - Live Codex test passed twice (latest live run on `e0bcd63e6`) in a temporary database and workspace, with a 300,000 ms warm window. First run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at `2026-10-07T13:15:03.274Z`. Second run `97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`, received `issue_monitor_due`, and completed the same task. Exactly one monitor wake was recorded. - Both live runs used native session `7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner `e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session `01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same process start time. This proves warm reuse for that local Codex test, not only successful scheduling. - Reproduce the paid live test with `node --import ./server/node_modules/tsx/dist/loader.mjs server/scripts/smoke-native-task-monitor.ts --run`, with the installed Codex binary on `PATH` and a valid local login. ## Risks - The scheduler now defers monitor dispatch while the task has an active native run. A stuck run still depends on the existing recovery lifecycle. - Idempotency uses the existing run ledger; no table or migration is added. - Other providers share the tested tool and completion contracts. Only Codex received a live model test. - Existing `call_api` lifecycle restrictions remain enforced. Monitor waits do not bypass task blockers, reviews, or approvals. ## Model Used OpenAI Codex, GPT-6 family, with tool use, code execution, and TypeScript/Rust editing. The session does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3a726e676f |
feat: enforce private task permissions across execution and data (#10633)
Enforce private task and project access across direct reads, search, execution, files, plugins, live delivery, and sharing mutations. Preserve downward-only sharing, current responsible-user authorization, and audited emergency access. Bind historical draft assets with migration 0314. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b67db12d90 |
feat: add durable storage for private tasks (#14717)
Add private task and project ownership, downward access grants, and immutable run/workspace provenance. Apply migration 0313 with bounded batch commits and concurrent indexes. Keep the enforcement and sharing changes in their dependent PRs. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d0db8820db |
fix(connections): repair native baseline and approval continuations (#15420)
## Thinking Path > - Paperclip manages AI agents and the tools they may use. > - Connection setup separates provider preference from permission to use a tool. > - The first native connection baseline could not exercise its intended decisions. > - The browser used mutable task titles, and the provider fixture already granted access. > - Native provider-choice instructions also disagreed with the preferred question format. Schema rejection gave no field guidance. > - This pull request repairs those test preconditions and native guidance, then fixes restart/approval defects exposed by the corrected baseline. It also restores missing OpenCode tool-error evidence. > - The benefit is an inspectable baseline before any further instruction reduction. ## Linked Issues or Issue Description Refs #15407. The original 15-cell baseline remains 0 PASS / 15 FAIL. Ten cells stopped on stale titles, two Codex cells had schema denials, two OpenCode cells used already-granted tools, and one Claude cell returned no native result. No intended user decisions were submitted. The exact invalid Codex field and underlying Claude failure cause remain unknown. [Original campaign](https://github.com/paperclipai/paperclip/actions/runs/37562577199) · [Original report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37562577199-1/index.html) ## What Changed - Match the browser's task route and visible identifier instead of a title the agent can change. - Start native provider-choice fixtures with no agent tool access. Verify the public effective-access records. - In the positive case, select Arcade, then grant its exact HubSpot tool through the real access card. Require both saved decisions and exactly one observed call. - Return a canonical `providerQuestionSet` for native input and retain the equivalent legacy `providerQuestion`. - Keep invalid input rejected. Return bounded schema locations and required field names without submitted values. - Preserve Claude's exact session/content identity while allowing authenticated registered instruction-copy paths to rotate on a new run. - Reject duplicate approval reports for an existing exact tool-action card before they create another human review. - Wait for a recorded service-approval continuation within the existing deadline; retain missing or failed continuation grades. - Forward OpenCode tool activity through the runner facade, preserving bounded errors and execution-part identity without inventing host-call joins or exposing arguments. - Preserve original grades, costs, scope limits and diagnoses in the dated repair report. ## Verification - Eval typecheck passes. Support suite: 1,801 PASS, one intentional skip; Node checks: 128 PASS. - Connection/schema tests: 51 PASS. Real-server public fixture setup: one PASS with zero providers. - Browser support regression: five PASS, including renamed and wrong tasks. - Focused Rust safe-feedback test: one PASS. - Repository typecheck and build pass before the latest master replay. Post-replay connection/shared/real-server fixture checks: 52 PASS; eval typecheck passes. The browser review fix additionally passes all five browser checks and seven suite checks. - The full local repository run was interrupted incomplete after about 45 minutes, with five integration failures retained. All five pass in a separate targeted invocation (1,250 unrelated tests skipped). No full local-suite pass or root cause for the initial local failures is claimed. - Corrected frozen source `162cc90fdabe7f505b88ae095044531b82784c92`: **10 PASS / 5 FAIL** across the [passing Codex canary](https://github.com/paperclipai/paperclip/actions/runs/37575158761) and [remaining 14 cells](https://github.com/paperclipai/paperclip/actions/runs/37576261807). The canary passes all 17 checks. Claude's two provider-choice continuations fail on restart, Claude service approval exposes an early evaluator rejection, Codex service approval creates a duplicate approval, and OpenCode provider-second times out after both decisions with no HubSpot call. No original result is regraded. - Final ledgers count 31 actual runs: 27 succeeded, two failed, two cancelled during cleanup. All 15 cleanup/budget checks pass. The late Claude continuation is absent from its earlier workflow snapshot; it remains in the result/API/final ledger. Original evidence retains 279 hashes. Recorded LLM subtotal $0.04553787 is incomplete billing, not actual total cost; local runtime is unmetered. - New repair regressions reproduce the Claude attach failure, duplicate approval acceptance and dropped OpenCode tool events before their respective fixes. Nine Rust attachment checks, 127 ACPX host/adapter tests, 33 completion/control-plane checks, nine eval deadline tests, 59 OpenCode proxy/driver tests, one Rust tool-error/redaction check, and TypeScript/Rust composer parity pass. Eval typecheck, repository typecheck and build pass. Existing support coverage is 1,802 PASS plus 128 Node PASS, one intentional support skip; two additional deadline tests also pass. - New-source full CI/review and live canaries are pending. The next bounded selection is Claude provider-decline, Codex service-approve and one OpenCode provider-second diagnostic with repaired event evidence. No broader campaign or instruction-reduction qualification is claimed. - Initial corrected campaign [37574251834](https://github.com/paperclipai/paperclip/actions/runs/37574251834) was cancelled during shared build after review found the breadcrumb whitespace assumption. Its matrix job has zero steps and no provider execution. The real adjacent-span browser regression now reproduces the old failure and passes after the fix. ## Risks - The corrected baseline remains 10/15. The new restart/approval fixes require live qualification; OpenCode evidence forwarding does not itself establish or fix its prior behavioral failure. - The positive provider case now expects three runs, including separate access approval. Its new results are distinct from the original invalid fixture. - The old Claude missing-result cause and rejected Codex field are unknown. These repairs do not retroactively explain or erase either failure. - Path rotation must preserve prompt, custom instruction, skill/content identity and protected provider settings; regression checks reject stale or changed content. No connection authorization, JSON schema, budget, cleanup, or final-result requirement is relaxed. Historical Everyday prompts and gateway setup remain unchanged. ## Model Used OpenAI Codex, GPT-6. The exact deployment variant and context window are not exposed in this session. Used repository inspection, code editing, test execution and retained-evidence analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
799e4d556f |
fix: make accounting durable and synchronize cost reporting (#14997)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
426e9fac15 |
merge: integrate Dot mailbox and fence fixes
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
89080f28a5 |
fix(dot): order mailbox receipts and retain fence acknowledgements
Serialize mailbox writes and cursor reads with the binding lock. Permit paused connections to receive and acknowledge fences while task tools remain denied. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f712015f7f |
merge: integrate current MCP connection improvements
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
beb1d5598a |
merge: integrate current MCP connection improvements
Preserve dedicated Dot authority while integrating personalized assistant setup and filtering revoked grants. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a15d074f2f |
merge: integrate reviewed Dot adapter fixes
Preserve the saved experimental feature gates and empty-instance startup behavior. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
78f4c4f5e6 |
fix(dot): preserve completion and admission receipts
Normalize accepted completion reports across the control plane and Rust bridge, serialize Dot assignments, replay work admission receipts, and continue subscription polling. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
99a9de9940 |
fix(mcp): personalize assistant connections and hide revoked grants (#15411)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Assistant connections let people use their organization from another assistant. > - Setup begins with an invitation and returns to a list of connected assistants. > - Client-only labels obscure the authorizing person, while revoked rows clutter that list. > - This pull request shows the person’s avatar, names connections by owner and client, and hides revoked rows. > - It also shortens the copied invitation while keeping the approval instructions. ## Linked Issues or Issue Description Related: #14933 and #15380. **What existing behavior does this improve?** Invitation copy and assistant connection management, in the organization’s Connections screen and the account-wide management page. **Subsystem affected** Cross-cutting: shared MCP connection types, server profile projection, UI and Storybook. **Current behavior** Connections are labeled only with a client name such as “Codex.” Revoked connections remain visible. The invitation includes an extra sentence about agent identity. **Proposed behavior** Show the authorizing person’s avatar and use names such as “Dotta’s Codex connection.” Hide revoked rows after successful revocation and when loading retained revoked grants. Failed revocation leaves the connection visible. Remove the extra identity sentence from invitation copy. **Reason and benefit** Make connection identity clear and keep the list focused on usable connections. **Breaking changes** The connection response adds optional `user` metadata with name and image. Older servers remain usable. Names and revoked-row visibility change in the UI; OAuth client identity, authorization and audit retention remain unchanged. ## What Changed - Shorten the shared invitation text. - Project the authorizing person’s name and avatar through the user-scoped connection endpoint, without returning email or credentials. - Reuse the existing Identity component and owner naming conventions across the connection page, catalog card and account-wide list. - Hide revoked grants and remove a successfully revoked row from the shared cache, even if the subsequent refresh fails. - Update documentation, regression tests and production-page Storybook fixtures and revocation journeys. ## Verification - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook` and `pnpm check:token-gates` pass. Final UI type checks also pass. - Focused consent and connection UI tests: 31 pass, including company filtering, legacy metadata, custom client names, user-initial fallbacks, failed revocation, retained revoked rows and refresh failure after successful revocation. - Existing MCP regression suite: 6 pass; 71 database checks are skipped locally because embedded PostgreSQL cannot start on this machine. The added database check verifies user-profile isolation and retained revocation history; CI runs these checks. - Browser verification with Storybook fixtures: owner avatar and name render; revocation removes the selected row in both production pages, leaves other connections visible, and restores the empty state after the last revocation. - All 54 current-head CI checks pass, with two optional Storybook jobs skipped. CI includes database, browser, runner, typecheck, build and clean-install canary coverage. - Greptile reviewed commit `1368d79e1066b418712224378d89d64c2b11cb86`: 5/5, no actionable findings or unresolved threads. - The full local `pnpm test:run` was stopped after complete CI passed. Local database coverage remains unavailable because embedded PostgreSQL cannot start; no full local-suite pass is claimed. ## Risks Low risk. The additive profile field is optional for compatibility. Revoked grants are filtered only from management UI and retained for audit. Revocation failure does not hide an active connection. No authorization scopes, token handling or schema changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, command execution and browser verification. The exact deployment ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b4f0ed180e |
fix(dev): preserve empty instances through managed startup
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a9a20fb5c6 |
feat(security): add read-only customer-success inspection APIs (#15405)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a persistent identity and a verified active run for governed access. > - Customer-success inspection needs broad reads without tenant writes or secret access. > - Ordinary board login and database credentials give more authority than this task needs. > - This pull request adds a dedicated inspection API and strict managed-run authority. > - Cloud owns short grants, human approval, replay protection, and audit records. > - The benefit is inspectable access that an operator can disable immediately. ## Linked Issues or Issue Description Refs: #15352. This change reuses the persistent Ed25519 identity from that PR. **Subsystem affected** Server authentication, pure resource readers, and the shared wire contract. **Problem or motivation** One internal Paperclip agent must inspect customer onboarding work. It must not receive owner login or database credentials. Reads must not create customer sessions, memberships, activity, or read receipts. **Proposed solution** Add disabled-by-default run authority and versioned tenant inspection endpoints. Require strict instance-bound managed-run JWTs on the home instance. Require exact-operation, single-use Cloud permits on tenants. Execute a reviewed company-scoped catalog in read-only transactions. Cloud applies seven-day stack-age eligibility and human exceptions. **Roadmap alignment** This is access support for Cloud deployments and governed agent identities. Bot creation, scheduling, scoring, and reports are separate work. The maintainer requested this implementation. ## What Changed - Reuse existing public identity reads and managed private-key injection. Reject unprovisioned keys, paused agents, ended runs, legacy signatures, and wrong instances. - Mount `/api/customer-success/v1` before actor/session synchronization. Verify Cloud permits and consume them centrally before reading. - Add explicit company-scoped database readers and bounded instruction, skill snapshot, run log, workspace, and asset reads. Preserve existing redactions and file protections. - Add protocol, security, database immutability, and managed-agent qualification tests. Add deployment and rollback documentation. ## Verification - Full `pnpm -r typecheck` and `pnpm build` passed. Server typecheck passed after review fixes. - The broad local `pnpm test:run` recorded 14,277 passes and four failures in unchanged suites: two timeouts and two PR-metadata mock assertions. All three affected suites passed on isolated reruns (36 tests). The complete CI matrix passes at the final head, including every test lane, typecheck, build, runner checks, canary dry run, and the security scan. - Focused inspection, JWT, and existing identity tests pass. The catalog test compares every public database table before and after reads. - Inspection and route-contract tests: 22 passed. The coordinated test runs a real managed process agent against separate home/customer PostgreSQL databases and a PostgreSQL broker over HTTP. It proves wake through the existing controller, bounded binary file reads, single challenge consumption across replicas, concurrent grants with a two-connection pool, scoped SQL audits, append-only runtime auditing, one-year retention, and unchanged tenant data/files. - Run the coordinated test with `PAPERCLIP_INSPECTION_CLOUD_DIST` pointing at the sibling Cloud build. Normal unit runs skip that optional private integration. - Final-head Greptile is 5/5 with no unresolved findings. - No production deployment or customer inspection occurred. ## Risks - This adds an authentication boundary. Keep both feature flags disabled until coordinated staging and canary qualification. - Cloud support must deploy after this API. Unsupported tenants fail closed. There is no owner-login or database fallback. - Existing redactions remain the content boundary. Arbitrary pasted secrets in readable prose or files may remain. - Remote files and suppressed provider traces remain unavailable. Wake can cause normal startup/background writes; test those separately. - Disable Cloud policy first during rollback. Preserve existing identity material and Cloud audit history. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, and browser testing. The session does not expose a more specific deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; broad-run flakes are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
88dddb8298 |
feat(dot): add a saved experimental setting with prerequisites
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3023aa2f61 |
Merge current master and preserve dedicated Dot agent gateway
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a6306ba606 |
feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
faa8e452c7 |
fix(tasks): stop repeated reminders for historical questions (#15392)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task questions remain saved so a person can answer them later. > - The composer moves an old question into history after a newer human message. > - Native completion still treated every pending question as a required response. > - This caused agents to demand an old answer after the person moved work forward. > - This pull request shares the historical-question rule across context and task execution. > - Agents can finish verified work while the original question remains answerable. ## Linked Issues or Issue Description Refs #15229, Refs #14613, Refs #13130. **What happened?** An agent repeatedly asked a person to answer a question that had moved into feed history. Completion feedback explicitly told the agent to request a response. The saved pending row also blocked task completion. **Expected behavior** A question before newer human direction remains answerable in the feed. Its pending state alone must not require another reminder or stop completed work. A new input blocker, an approval, or a configured review stage must keep its gate. **Steps to reproduce** 1. Let a task agent create an ordinary question. 2. Dismiss the question and send a newer task message. 3. Let the agent finish the requested work and submit its completion report. 4. Observe a demand to answer the old question and a retained completion gate. ## What Changed - Add one company-scoped predicate for historical questions. Only later human comments count. Exclude agent attribution, run attribution, system notices, and untrusted source data. - Apply the predicate to completion feedback, native waits, finalization, commit validation, retry validation, blocked routing, and successful-run handoff. - Include question classification and guidance in heartbeat context and both native task-context tools. Add the guidance to fresh and resumed task prompts. - Replace automatic reminders for current ordinary questions with instructions to assess the real blocker, continue independent work, and withdraw obsolete questions through the existing API. - Preserve historical question rows during completion while cancelling their live native source runs through the existing post-commit and recovery paths. Let an authorized human answer them after completion without reopening work or creating a response wake. Preserve cancellation, current-input, approval, permission, credential, connection, and review gates. Add no dismissal storage or migration. - Document the rule in the execution contract and agent skill. Refresh generated capability source anchors. Add database-backed status, context, attribution, and governance regression tests. ## Verification - `pnpm build` passed. The server rebuild also passed after the lifecycle fix. Generated capability contract and inventory checks passed after the agent documentation update. - `pnpm -r typecheck` passed. Final `pnpm --filter @paperclipai/server exec tsc --noEmit` also passed after the last test additions. - Lifecycle and interaction regressions passed: 224 tests in 3 suites. Context and prompt tests also passed. Final historical-question cases passed (33 tests), native cancellation/recovery cases passed (6 tests), and the existing interaction/confirmation suites passed (72 tests). The full local `pnpm test:run` was attempted and stopped after more than two hours with unrelated fixture/hook timeout failures; it did not pass. All 52 successful GitHub checks are green on the latest commit, including the complete test matrix; no checks are pending or failing. Greptile is 5/5 and both review threads are resolved. - Regression cases cover the old-question/new-human-message sequence, final task status, answering after completion with no wake, live-run cancellation and crash recovery, current input blockers, both task-context tools, API context, timestamp precision, attribution boundaries, and protected gates. ## Risks - A later human task message makes an earlier ordinary question historical even if its input is still missing. The agent must identify the current blocker and ask only for information that still prevents work. - Browser dismissal remains a local preference. Dismissal without a later human message is not recorded by this change. - No schema change or data migration. Completion retains ordinary historical questions; cancellation still expires them. A completed task accepts historical answers only from an authorized human and creates no response-delivery outbox row. Governed requests retain their gates. ## Model Used OpenAI Codex, GPT-6, with repository editing, code execution, and browser diagnostics. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2d0c138122 |
Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also use assistants in Codex, Claude, and other MCP clients. > - The existing assistant connection can read work and create tasks or comments. > - It cannot edit tasks, exchange files, or manage normal agent and project settings. > - These operations must retain the person's permissions and Paperclip's execution rules. > - This pull request adds an explicit operation registry and separately consented configuration access. > - Assistants can manage work without receiving credentials or runner authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: assistant MCP, domain routes, consent UI, storage, and Product E2E. **Problem or motivation** A connected assistant cannot update tasks, maintain documents, attach files, or configure existing agents, projects, and skills. Users must leave the assistant for these routine actions. **Proposed solution** Add named tools and a restricted API registry to direct connections. Require separate configuration consent. Reuse domain routes and retry receipts. File uploads save the attachment when the byte transfer succeeds. **Alternatives considered** Arbitrary REST forwarding would expose administration and credential operations. Runner impersonation would bypass execution ownership. Separate upload completion calls add unnecessary client state. **Roadmap alignment** Checked ROADMAP.md and related MCP pull requests. This extends the human-authorized connection from #14933. It does not replace the runner or introduce agent impersonation. Companion Cloud routing and directory isolation: https://github.com/paperclipai/paperclip-cloud/pull/678. ## What Changed - Add task editing, finish/block, documents/revisions, deliverables, agent settings/instructions, projects/repositories, and skills/files. - Add an allowlisted API search/call registry with identical field restrictions, scopes, and retry identities. - Add unchecked configuration consent. Existing write grants retain their current authority. - Add hashed, expiring file transfer tickets and atomic upload receipts. No completion call is required. - Preserve company boundaries, human attribution, active native execution ownership, and execution review gates. - Add protocol/domain tests, consent stories, and eight paid Product E2E workflows. - Repair two CI fixture races: await cold route setup before assertions, and wait for asynchronously loaded connection copy. Both fixture suites pass (24 + 48 tests). ## Verification - Consent revision: one write-access checkbox controls requested work and configuration permissions in browser and device flows. All 16 consent tests, UI typecheck/build and token gates pass. Updated interactive stories cover default approval, opt-out and viewer restrictions. The paid browser helper uses the new exact label. Real GPT-5.4 Mini Product E2E passes 2/2 at `64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic retries, cleanup passed; $0.04149375 estimated assistant cost plus unpriced worker usage. Raw results, usage and source fingerprints are retained in the worktree. UI and Product E2E typechecks pass. - Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks pass; two optional Storybook checks skip. Greptile 5/5 on that head, no unresolved review threads. Final consent head `64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing and Greptile 5/5 with no unresolved threads. The unchanged Cursor sandbox test had one 10-second timeout, passed in local isolation, and passed its single CI rerun; the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37554106934). The existing chat retry-denial browser test had one visibility failure; its single rerun passes, and the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37542735691). - Full workspace `pnpm -r typecheck` and `pnpm build` pass at final runtime source `b2196fae1`. UI token gates pass. - 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests pass, including one-connection PostgreSQL OAuth and concurrent upload retries. - Paid Product E2E: all eight expanded cases qualified across Mini, Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell timed out before application startup. Final affected-case qualification passes 9/9 on all three models with grader v16, including the failed cell. Automatic retries disabled; failures, costs, source hashes and independent durable-state/file assertions are retained in [the verification record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md). - Actual Codex CLI, Claude Code and OpenCode clients completed local reads/mutations. Codex wrote a report, Claude updated it in a later conversation, and OpenCode uploaded/downloaded a file with matching SHA-256 and registered the attachment. Revoking the CLI grant rejects subsequent bridge initialization. - Butter staging is verified on final runtime `b2196fae1` ([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)). A fresh OpenCode workspace fetched the copied invitation, configured remote MCP, started OAuth and reached real consent with configuration unchecked. Invalid transfer tickets return 403 through Cloud. Human approval for the new persistent staging grant is pending; hosted task/file success is not yet claimed. The final transaction fix is deployed. - Full local `pnpm test:run` passed 15,614 general-server tests but stopped on two macOS timeouts. The heartbeat test passed in isolation; the existing 40,000-file Git stress fixture timed out again. Its Linux CI lane passes. Later local full-suite phases did not run after the timeout; this is not an all-green local full-suite claim. - Instructions and security limits are in `doc/public-mcp.md`; the saved plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`. ## Risks - This expands the experimental direct MCP surface. Explicit schemas and domain permissions must stay synchronized. - Migration 0311 adds transfer tickets and upload receipts. Expired orphan cleanup must not remove committed attachments. - Configuration requires a new consent request containing that scope; the single write-access choice controls it alongside work mutations. Refreshing an old grant does not add it. - The public directory keeps its original ten tools through the companion Cloud change. - Hosted consent/work proof remains the final delivery gate. The PR stays draft while approval of the new staging grant is pending; code checks and review are green. Merging is a separate action. ## Model Used OpenAI Codex (GPT-6, tool use and code execution). The exact serving model ID and context window are not exposed in this session. Paid evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001; claude-sonnet-4-6. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite macOS limitation disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
892b0b3606 |
fix: scope quota reports to authorized subscription accounts (#14994)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
eab93fd4a0 |
fix: checkpoint adapter usage and preserve unknown prices (#14991)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5201456bfc |
Fix Dot reconnect, disconnect, and pairing recovery
Wake the existing assignment after verified reconnection, fence and cancel matching agent work on OAuth revocation and replay, and restore durable pairing references into reopened forms. Synchronize protocol fixtures and board-only OpenAPI coverage with the merged gateway. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f77fcbf4bf |
feat(apps): add Telem.AI web search connection (#15379)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Research agents need current web information. > - The Apps catalog connects agents to remote MCP tools through the normal access rules. > - Telem.AI supplies web search and page reading through one API key. > - This PR adds its catalog entry, optional search settings, artwork, and setup guide. > - The connection keeps an operator's saved header policy when they reconnect. ## Linked Issues or Issue Description Refs #15302. Related catalog work: #13881. This PR continues #15302 by Yifei Ai (@aiwen324). Thank you for the connector and its review fixes. All seven original commits are preserved. GitHub denied the attempt to push to the contributor's fork, so this branch retains the repair commit and merges current master. Master now includes the same test fix. For a squash merge, keep the original author in the final commit message: ```text Co-Authored-By: Yifei Ai <aiwen324@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> ``` A search found no separate public Telem issue or competing Telem PR. The existing Apps path matches the roadmap. **Agent or provider** Telem.AI provides web search and page reading through a hosted MCP server. **Why this adapter is useful** Agents can use multiple search providers through one governed connection. Operators can set the search tier, auto routing, and provider lists. **How the agent is invoked** The remote MCP server uses Streamable HTTP at `https://mcp.telem.ai/mcp`. The API key uses an `Authorization: Bearer` header. See the [official MCP guide](https://docs.telem.ai/integrations/mcp/). ## What Changed - Add the Telem.AI definition, research entry, permission review, and generated registry entry. - Add four optional settings. Unset settings send no request header. - Add official light and dark artwork, source records, and a setup guide. - Forward company, issue, agent, run, project, and correlation IDs by default. Preserve a saved policy, including disabled forwarding, on reconnect. - Add catalog and connection tests. ## Verification Current head: `b4164477fb1b312a504789bf17b51c963244c0fd`. Merged master: `228f0e2807c5b59d2aa129cf2d80b9777ebabf07`. - Resolved five shared catalog conflicts after the Superagent connection merged. - Keep both providers in the research ledger, generated registry, generator, branding manifest, and connection guide index. - Correct the combined catalog totals: 52 self-serve candidates, 55 research entries, and 68 Apps entries. - The published Git tree exactly matches the tested local resolution. - Catalog and Apps UI suites: **295 tests pass** after the catalog count fixes. - Connection service suite: **387 tests pass** in the full run. Its only failure was the old catalog count. That test passes on a focused rerun after the fix. This gives **388 passing service tests** across the two runs. - Total focused coverage: **683 passing tests**. The first runs exposed four fixed-count assertions that needed the combined totals. - Shared package build and plugin SDK compile pass. Token gates and whitespace checks pass. - Generation with `--definitions-only` reproduces the Telem definition and registry. The unrelated AgentMail and Linear drift remains excluded. - Local UI typecheck ended with exit 137 at the container memory limit. Full local typecheck, test, and build are not claimed. Earlier runs also recorded missing Cargo and Node development headers. - GitHub reports a clean merge state against master `228f0e280`. - All 54 checks are complete: **52 passed and two Storybook checks skipped**. No check failed or remains pending. - [CI](https://github.com/paperclipai/paperclip/actions/runs/37545412406) passes on this head. This includes typecheck, build, tests, browser shards, Runner checks, and Canary Dry Run. - [Greptile](https://github.com/paperclipai/paperclip/pull/15379#issuecomment-6024519537) is **5/5 on this head**. There are no review threads, open P2s, recommendations, or follow-ups. - [Superagent](https://github.com/paperclipai/paperclip/runs/112548142930) passes. - Final recovery checks confirm all seven original commits and current master remain in history. Token gates and whitespace checks pass. - The final recovery run makes no source change. It verifies the published repair and retains the local check limits below. - [Commitperclip](https://github.com/paperclipai/paperclip/actions/runs/37545407943) passes with no failures. Its only informational note asks the merger to keep the author trailer above. - All seven original contribution commits remain in history. The diff against master contains the same 14 Telem files. It adds no dependency, lockfile, schema, or workflow change. - The managed GitHub CLI capability was missing in this run. The installed GitHub connection applied the base files, merged master, then restored the tested combined catalog. No history was rewritten. The previous head `2d32a0094` passed all remote gates and had Greptile 5/5. Those results do not verify this new head. The original PR reports live setup, discovery, settings headers, gateway calls, and context-header forwarding. This repair does not repeat those account-bound checks. The permission record still marks maintainer live qualification as outstanding. ## Risks - Telem.AI receives the six context IDs by default. The saved header policy controls forwarding. Search use is billed to the account that owns the key. - All agents on a connection share its search settings. - The merge uses master's route-test setup unchanged. The company-boundary assertions remain intact. - No schema, dependency, or workflow change is included. - Live provider evidence is attributed to the original contributor. Maintainer live qualification remains outside this CI repair. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Original contribution: Anthropic Claude Opus 5.5 (`claude-opus-5-5`), 1M-token context, through Claude Code with shell, editing, and test tools, as disclosed in #15302. - CI repair and review: OpenAI `gpt-6-astra`, through Codex with reasoning, shell, editing, and GitHub tools. The runtime does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Yifei Ai <aiwen324@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
228f0e2807 |
fix(ui): make pull request review waits visible (#15375)
## Thinking Path > - Paperclip lets people supervise agent work and review its outputs. > - A task can wait for a human to review a pull request while an agent schedules checks. > - The PR work product and detected external links use separate display paths. > - A private PR can disappear from the useful sidebar view when its provider lookup fails. > - The monitor countdown says when the next check runs but omits the saved review request. > - This change shows the saved PR and requested review in the task's existing surfaces. ## Linked Issues or Issue Description **What happened?** A task kept checking whether a private GitHub PR had merged. Its work product requested board review, but the Properties sidebar used detected external objects instead. The PR URL was stored only in metadata, which the PR refresh path did not read. The composer showed a generic monitor countdown with no PR link or review action. **Expected behavior** Show saved PR links even if provider access fails. During a GitHub monitor wait, make outstanding review requests visible with the next check time and a status-check action. Stop requesting review when a PR merges or closes. **Steps to reproduce** 1. Register a PR work product with `reviewState: needs_board_review` and its URL in `metadata.url`. 2. Schedule an external-service monitor for GitHub. 3. Open the task with external-object lookup unavailable or unable to access the private repository. 4. Inspect Properties and the composer wait strip. Related work: #14469 added rich artifact cards. #8759 addresses attention on monitored task blockers. This change uses the existing work products and monitor action; it adds no blocker or approval mechanism. ## What Changed - Show saved PRs in Properties, including metadata-only links, independent of external-object availability. - Deduplicate equivalent GitHub PR URLs while preserving the saved navigation link, and put explicit review requests first. - Keep provider status and freshness on the combined PR row; keep private PR links usable when lookup fails. - Show the saved PR review request above the composer and in the existing monitor banner during a GitHub monitor wait. - Label the existing monitor action `Check status` and show check failures inline. - Suppress review prompts for merged, closed, or archived PRs even if their review flag is stale. - Refresh PR metadata using `metadata.url` and the existing `repository` alias. - Share saved work-product reads across the thread, Properties, and Artifacts. Refresh GitHub in a separate query and enrich only matching PR versions. - Start a fresh saved-row request on live invalidation so late provider responses cannot hide new artifacts or changed review requests. - Cover stalled GitHub lookups in both panels so refresh latency cannot hide saved work. - Scope monitor-check mutation state to the task so failures and late responses do not leak across navigation. - Refresh GitHub status when the displayed run finishes, including when saved PR rows have not changed. - Clarify that PR review and external release handoffs need a saved human-input interaction with an agent assigned for continuation; a flag, monitor, or handoff comment alone does not create that card. ## Verification - 386 tests pass across eleven affected UI and server suites on `184c7d8746`. Regressions cover saved links, provider status, cold-cache loading, panel reopening, monitor errors during navigation, and live updates during provider refresh. - Three live-update regressions and two run-completion regressions fail before their fixes and pass afterward. These tests use the real API client's GET coalescing and abort handling. - Full repository typecheck and build passed on the merged parent `1a49112dd2`. The latest UI changes pass UI typecheck and token gates; CI also verifies the latest build. Capability contract/inventory checks pass. - Storybook build and browser checks passed before the query race fix. Browser checks cover desktop and 390px mobile review waits, plus video, mixed-file, and empty artifact galleries after background refresh. - The earlier full local `pnpm test:run` was stopped under disk pressure. It reported failures outside the changed suites; an isolated skill-cache run reproduced three existing macOS permission failures. The full local suite was not rerun for this follow-up. - Latest-head CI passes on `184c7d8746`, including the browser shards and canary dry run. Apex is 5/5 with no actionable findings and no unresolved threads. All eight reported findings are addressed. ## Risks - This uses saved PR review state. When GitHub access fails, the saved state can remain stale until the agent updates it. The link remains visible and the status check remains available. - `Check status` wakes the existing monitor owner. It does not merge a PR, accept an approval, or mark the task done. - No schema, permissions, or scheduler behavior changes. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository analysis, code editing, and test execution. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — 277 affected tests - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cbd21b3e7 |
feat(apps): add Superagent connection (#15394)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents reach external services through the Apps catalog. Each catalog entry is a reviewed `AppDefinition` that connects a provider's hosted MCP server to Paperclip's shared vault, grants, policies, gateway, and audit trail. > - Superagent (`superagent.sh`) is a security platform. Its hosted MCP server lets agents read and triage security findings, start red-team reports, run Contributor Trust and dependency update jobs, and score web pages, email, files, skills, MCP repositories, and packages before an agent trusts them. > - Superagent is not in the catalog. Operators must use the generic "Connect your own MCP server" flow. That flow has no branding, no key guidance, and no warning about billable or destructive tools. > - The connector playbook supports this provider with existing definition fields. The server accepts an organization API key as a bearer header. It publishes no OAuth authorization-server metadata, so browser sign-in is not possible. > - This pull request adds the Superagent definition, its artwork, the research and permission-review ledger rows, documentation, deterministic tests, and one reviewed risk-classification rule. > - The benefit is a branded, governed Superagent connection with clear key guidance and a warning about tools that cost credits or delete data. ## Linked Issues or Issue Description **Problem or motivation** Teams that use Superagent for PR security, red teaming, and agent guardrails want their Paperclip agents to read findings, return structured reports, and score content before they use it. Superagent is not in the Apps catalog. Operators must paste the MCP URL and an `Authorization` header into the generic remote-MCP flow. That flow gives no branding and no provider guidance. It also does not tell the operator that the key reaches the whole organization, or that some tools consume credits or permanently delete findings. **Proposed solution** Add a catalog-only Superagent connection that follows the connector playbook. It has one method: a customer organization API key (`sk_live_...`), sent as an `Authorization: Bearer` header to `https://www.superagent.sh/mcp`. The field helper text explains that Superagent keys are not scoped. The method warning tells operators to set billable and destructive actions to Ask first before agents run unattended. All discovered tools stay governed by the normal per-action policies. **Alternatives considered** A browser sign-in method was not added. The server's protected-resource metadata names `https://superagent.sh` as its authorization server, but that origin publishes no `oauth-authorization-server` or `openid-configuration` document, so Paperclip cannot discover OAuth endpoints. A plugin was not needed because the connection needs no custom UI, tables, workers, or webhooks. Relying on the generic risk classifier was not enough. Several Superagent mutations (`triage_finding`, `scan_*`, `restore_agent_builtin_rule`) use names that it reads as reads, so a narrow reviewed Superagent rule was added instead. **Roadmap alignment** This extends the existing self-serve remote-MCP connection catalog. It does not overlap planned core work. ## What Changed - Added the `superagent` row to `packages/shared/src/self-serve-mcp-research.json` (API-key auth, risk tier S4). - Added the `superagent` provider to `scripts/ingest-app-definitions.mjs` (category, key placement and placeholder, console links, guidance, description). Regenerated `packages/shared/src/app-definitions/superagent.json` and the generated registry. - Added the `superagent/mcp-api-key` permission review to `doc/connections/tool-method-permission-reviews.json`, with key-permission text and evidence links. - Added Superagent's official mark (`ui/public/brands/apps/superagent.png`, the 460×460 avatar of the official `superagent-ai` GitHub organization) and the brand manifest entry. - Added gallery copy for the Superagent card. - Added a reviewed Superagent rule to `classifyRisk` in `server/src/services/tool-access.ts`. Only `list_*` and `get_*` tools, and tools that Superagent marks read-only, are reads. `delete_*` and `revoke_agent_client` are destructive. All other tools are writes, so billable and rule-changing tools can be set to Ask first. - Put the Ask-first advice in the API-key helper text, because the key form shows helper text and not method warnings. - Added `doc/connections/SUPERAGENT.md` (transport and auth, why there is no OAuth, administrator setup, capabilities and policy, manifest, brand provenance, validation hook). Linked it from the connections README and the permission audit. - Tests: definition shape, store visibility and artwork, URL recognition, the bearer header on discovery with the key kept out of connection config, read/write/destructive classification of fixture tools, the Superagent risk rule (including `triage_finding` and `restore_agent_builtin_rule`), the visible Ask-first advice and API-key gating of the connect form, and the pinned catalog counts. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts packages/shared/src/app-definitions-url.test.ts ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx`: 317 passed. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts`: 385 passed. - `node scripts/check-app-brand-assets.mjs` and `node --test scripts/app-brand-validation.test.mjs`: passed. - `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter @paperclipai/server typecheck`, and `pnpm --filter @paperclipai/ui typecheck`: clean. - Manual: in a local instance, open Apps → Browse and confirm the Superagent card and icon. Open `/apps/connect?source=superagent`. Confirm the single API-key method, and confirm that Connect enables only after a key is entered. - Live metadata probe on 2026-10-06: an unauthenticated `initialize` on `https://www.superagent.sh/mcp` returns 401 with `resource_metadata="https://www.superagent.sh/.well-known/oauth-protected-resource"`. That document returns 200. No authorization-server metadata exists at the named issuer. ## Risks - Low risk to existing providers. The change is additive catalog data plus tests. The generated registry only gains one import. The new risk rule runs only for Superagent connections. - A Superagent key reaches its whole organization. Some tools consume credits (`create_*_report`, `triage_finding`) or delete data permanently (`delete_finding`). Every action starts Allowed under the current product default. The key helper text tells operators to set these actions to Ask first. - The permission-review ledger records live proof as not run. No Superagent account was used. The lifecycle checklist in `doc/connections/SUPERAGENT.md` needs a documented pass before the entry is fully qualified. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with extended thinking and tool use (shell, file editing, web fetch, browser checks). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1be3f69703 |
Merge master and reconcile Dot with the released MCP gateway
Reuse scoped browser and device consent, verified client metadata, signed event admissions and warm-standby gates. Preserve the dedicated agent resource and Runner authority. Regenerate the Dot-only migration after the merged gateway history and cover issuer, company, viewer and token isolation. Co-Authored-By: Paperclip <noreply@paperclip.ing> |