mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 23:08:38 +02:00
ed6abbf158b73086ca0237bec23ea72ec89e297f
4916
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ed6abbf158 |
build(deps): bump actions/cache/restore from 5.1.0 to 6.1.0 (#15163)
Bumps [actions/cache/restore](https://github.com/actions/cache) from 5.1.0 to 6.1.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/actions/cache/releases">actions/cache/restore's releases</a>.</em></p> <blockquote> <h2>v6.1.0</h2> <h2>What's Changed</h2> <ul> <li>Bump <code>@actions/cache</code> to v6.1.0 - handle read-only cache access by <a href="https://github.com/jasongin"><code>@jasongin</code></a> in <a href="https://redirect.github.com/actions/cache/pull/1768">actions/cache#1768</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/actions/cache/compare/v6...v6.1.0">https://github.com/actions/cache/compare/v6...v6.1.0</a></p> <h2>v6.0.0</h2> <h2>What's Changed</h2> <ul> <li>Update packages, migrate to ESM by <a href="https://github.com/Samirat"><code>@Samirat</code></a> in <a href="https://redirect.github.com/actions/cache/pull/1760">actions/cache#1760</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/actions/cache/compare/v5...v6.0.0">https://github.com/actions/cache/compare/v5...v6.0.0</a></p> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/actions/cache/blob/main/RELEASES.md">actions/cache/restore's changelog</a>.</em></p> <blockquote> <h1>Releases</h1> <h2>How to prepare a release</h2> <blockquote> <p>[!NOTE] Relevant for maintainers with write access only.</p> </blockquote> <ol> <li>Switch to a new branch from <code>main</code>.</li> <li>Run <code>npm test</code> to ensure all tests are passing.</li> <li>Update the version in <a href="https://github.com/actions/cache/blob/main/package.json"><code>https://github.com/actions/cache/blob/main/package.json</code></a>.</li> <li>Run <code>npm run build</code> to update the compiled files.</li> <li>Update this <a href="https://github.com/actions/cache/blob/main/RELEASES.md"><code>https://github.com/actions/cache/blob/main/RELEASES.md</code></a> with the new version and changes in the <code>## Changelog</code> section.</li> <li>Run <code>licensed cache</code> to update the license report.</li> <li>Run <code>licensed status</code> and resolve any warnings by updating the <a href="https://github.com/actions/cache/blob/main/.licensed.yml"><code>https://github.com/actions/cache/blob/main/.licensed.yml</code></a> file with the exceptions.</li> <li>Commit your changes and push your branch upstream.</li> <li>Open a pull request against <code>main</code> and get it reviewed and merged.</li> <li>Draft a new release <a href="https://github.com/actions/cache/releases">https://github.com/actions/cache/releases</a> use the same version number used in <code>package.json</code> <ol> <li>Create a new tag with the version number.</li> <li>Auto generate release notes and update them to match the changes you made in <code>RELEASES.md</code>.</li> <li>Toggle the set as the latest release option.</li> <li>Publish the release.</li> </ol> </li> <li>Navigate to <a href="https://github.com/actions/cache/actions/workflows/release-new-action-version.yml">https://github.com/actions/cache/actions/workflows/release-new-action-version.yml</a> <ol> <li>There should be a workflow run queued with the same version number.</li> <li>Approve the run to publish the new version and update the major tags for this action.</li> </ol> </li> </ol> <h2>Changelog</h2> <h3>6.1.0</h3> <ul> <li>Bump <code>@actions/cache</code> to v6.1.0 to pick up <a href="https://redirect.github.com/actions/toolkit/pull/2435">actions/toolkit#2435 Handle cache write error due to read-only token</a></li> <li>Switch redundant "Cache save failed" warning to debug log in save-only</li> </ul> <h3>6.0.0</h3> <ul> <li>Updated <code>@actions/cache</code> to ^6.0.1, <code>@actions/core</code> to ^3.0.1, <code>@actions/exec</code> to ^3.0.0, <code>@actions/io</code> to ^3.0.2</li> <li>Migrated to ESM module system</li> <li>Upgraded Jest to v30 and test infrastructure to be ESM compatible</li> </ul> <h3>5.0.4</h3> <ul> <li>Bump <code>minimatch</code> to v3.1.5 (fixes ReDoS via globstar patterns)</li> <li>Bump <code>undici</code> to v6.24.1 (WebSocket decompression bomb protection, header validation fixes)</li> <li>Bump <code>fast-xml-parser</code> to v5.5.6</li> </ul> <h3>5.0.3</h3> <ul> <li>Bump <code>@actions/cache</code> to v5.0.5 (Resolves: <a href="https://github.com/actions/cache/security/dependabot/33">https://github.com/actions/cache/security/dependabot/33</a>)</li> <li>Bump <code>@actions/core</code> to v2.0.3</li> </ul> <h3>5.0.2</h3> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/actions/cache/commit/55cc8345863c7cc4c66a329aec7e433d2d1c52a9"><code>55cc834</code></a> Merge pull request <a href="https://redirect.github.com/actions/cache/issues/1768">#1768</a> from jasongin/readonly-cache</li> <li><a href="https://github.com/actions/cache/commit/d8cd72f230726cdf4457ebb61ec1b593a8d12337"><code>d8cd72f</code></a> Bump <code>@actions/cache</code> to v6.1.0 - handle cache write error due to RO token</li> <li><a href="https://github.com/actions/cache/commit/2c8a9bd7457de244a408f35966fab2fb45fda9c8"><code>2c8a9bd</code></a> Merge pull request <a href="https://redirect.github.com/actions/cache/issues/1760">#1760</a> from actions/samirat/esm_migration_and_package_update</li> <li><a href="https://github.com/actions/cache/commit/e9b91fdc3fea7d79165fceb79042ef45c2d51023"><code>e9b91fd</code></a> Prettier fixes</li> <li><a href="https://github.com/actions/cache/commit/e4884b8ff7f92ef6b52c79eda480bbc86e685adb"><code>e4884b8</code></a> Rebuild dist</li> <li><a href="https://github.com/actions/cache/commit/10baf0191a3c426ea0fa4a3253a5c04233b6e18f"><code>10baf01</code></a> Fixed licenses</li> <li><a href="https://github.com/actions/cache/commit/e39b386c9004d72a15d864ade8c0b3a702d47a37"><code>e39b386</code></a> Fix test mock return order</li> <li><a href="https://github.com/actions/cache/commit/b6928203372a8571ff984c0c883ef3a1adfb0c06"><code>b692820</code></a> PR feedback</li> <li><a href="https://github.com/actions/cache/commit/60749128a44d25d3c520a489e576380cf00ff3f1"><code>6074912</code></a> Rebuild dist bundles as ESM to match type:module</li> <li><a href="https://github.com/actions/cache/commit/5a912e8b4af820fa082a0e75cfd2c782f8fbfe0e"><code>5a912e8</code></a> Fix lint and jest issues</li> <li>Additional commits viewable in <a href="https://github.com/actions/cache/compare/caa296126883cff596d87d8935842f9db880ef25...55cc8345863c7cc4c66a329aec7e433d2d1c52a9">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>canary/v2026.1008.0-canary.19 |
||
|
|
1881894973 |
fix: prevent inbox archive deadlocks during completion (#15615)
Take a company-scoped parent lock before writing archive state and reuse the caller transaction. Preserve newer archive timestamps and attribution when a request waits, so completed tasks remain archived. Verified with real PostgreSQL race, rollback, company isolation, concurrency, timestamp and visibility tests, independent review, full CI and Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3d1d5294d9 |
Drain idle plugin workers before automatic sleep (#15599)
## Thinking Path > - Paperclip runs autonomous work through agents, schedules, and plugins. > - Automatic idle sleep must preserve accepted work and unfinished cleanup. > - Every enabled plugin currently blocks sleep, including unused providers. > - An enabled flag or package version cannot prove that a worker is idle. > - This change adds a live worker drain under the existing owned hold. > - Idle workers can allow sleep while unknown work stays protected. ## Linked Issues or Issue Description Refs #15522. Related: #15391 adds plugin readiness for agent admission; this change concerns instance sleep and does not replace that contract. **Subsystem affected** Plugin worker lifecycle and automatic idle sleep. **Problem or motivation** A workspace with no pending work cannot sleep when any plugin is enabled. Removing that check alone would lose accepted RPCs, background tasks, or cleanup after a caller timeout. **Proposed solution** Require a live `onIdleDrain` handshake from each worker. Close admission in both processes for the exact owner and expiry. Count accepted work until completion. Continue checking durable work separately. ## What Changed - Add bounded worker holds, exact-owner release, automatic expiry, and an abort signal for plugin-owned background work. - Count host and worker requests through their real completion receipts. Keep timed-out work counted. Check active notifications and terminal routes. - Accept enabled plugins only when their current workers provide matching runtime receipts. Missing workers, old SDKs, crashes, invalid replies, and unknown cleanup still prevent sleep. - Let unused Daytona workers opt in. Once a worker contacts the provider, it remains a blocker for that process lifetime. This restriction avoids treating its existing timeout and terminal-close behavior as a cleanup receipt. - Document the plugin author contract. No schema or user-facing API is added. ## Verification - All hosted CI checks passed on |
||
|
|
4fe45bf8fe |
refactor(heartbeat): extract run retrieval and session state (#15601)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Heartbeat runs retain task sessions, usage totals, and bounded run data. > - After the preparation extraction, the main service still has over 24,000 lines. > - Run reads and session operations share one database and have no dispatch dependencies. > - This pull request moves that group into one module and preserves its callers. > - The benefit is a smaller run engine and one place to maintain saved run and session state. ## Linked Issues or Issue Description **What existing behavior does this improve?** It improves the structure and test coverage of heartbeat run retrieval and session state. **Current behavior** `server/src/services/heartbeat.ts` has 24,033 lines on the merged base. It mixes run projections, session queries, resume rules, compaction, and usage helpers with execution and recovery. **Proposed behavior** Move that group into `server/src/services/heartbeat/run-state.ts`. The main file loses 1,770 lines and ends at 22,263 lines. Keep current behavior, public exports, and service methods. This continues #15568, #15573, #15578, and #15591. **Reason and benefit** One domain extraction makes useful progress toward a few manageable files. Run reads and saved session state can be reviewed together without searching the run engine. A search found no duplicate open extraction. Related session changes include #15142, #14661, and #14670; this PR does not apply their proposed behavior changes. **Breaking changes** None. Existing public helpers retain their identity. Service method names and results stay the same. ## What Changed - Move bounded run projections, encoding checks, and run/event reads into `heartbeat/run-state.ts`. - Move task session persistence, explicit resumes, reset rules, compaction, and usage/billing helpers into the same module. - Bind database operations through `createHeartbeatRunState(db)`. Construction does no database work. The encoding cache stays local to each service instance. - Keep execution order, status transitions, session-goal recovery, cost accounting writes, dispatch, and cancellation in `heartbeat.ts`. - Preserve all 140 exports. A syntax-tree comparison confirms that 43 moved function bodies, 22 declarations, and nine service method bodies are unchanged. The copied private string helper is unchanged. Surrounding orchestration is unchanged apart from the factory binding. - Add eight tests for legacy export identity, database binding, encoding cache isolation, scoped resumes, cumulative usage, compaction, and stale or cancelled conversation session writes. - Document the module boundary in `doc/DEVELOPING.md`. ## Verification - Before and after extraction, eight existing suites pass: 235 tests. They cover workspace/session helpers, runtime state, cost accounting, ledger attribution, task reset, timer reset, run lists, and run privacy. - The new module suite passes: eight tests, including five real PostgreSQL cases. Total focused coverage: 243 tests across nine suites. - Commands: `pnpm exec vitest run server/src/services/heartbeat/run-state.test.ts` and `pnpm exec vitest run server/src/__tests__/heartbeat-runtime-state.test.ts server/src/__tests__/heartbeat-cost-accounting.test.ts server/src/__tests__/heartbeat-ledger-billing-code.test.ts server/src/__tests__/heartbeat-task-session-reset.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/heartbeat-timer-wake-session-reset-pf4.test.ts server/src/__tests__/heartbeat-list.test.ts server/src/__tests__/heartbeat-run-privacy-routes.test.ts`. - `pnpm --filter @paperclipai/server typecheck` passes. - `pnpm -r typecheck` and `pnpm build` pass. - The full local `pnpm test:run` did not complete and was stopped with SIGINT after it reported an OAuth scope health-check failure and a workspace-runtime timeout. All six OAuth variants and the workspace case pass in isolation. A single rerun of the 413-case tool-access suite passed the OAuth case but hit a different Notion callback timeout; that Notion case also passes in isolation. The full local suite is not claimed as passing. - Greptile is 5/5 with no findings or inline comments on head `2ca0908cfc112dd6ae26e2565cb3611b5bec52ba`. Its exact-head check completed successfully. - CI remains blocked on hosted-runner assignment after more than ten minutes. [`ci / Select trusted runner`](https://github.com/paperclipai/paperclip/actions/runs/37828454507/job/113487287178) is queued for `ubuntu-latest` with no runner assigned. All seven completed review/security checks pass; two optional Storybook checks are skipped. The full CI matrix has not started. It must complete before merge. ## Risks - Wiring errors could bind reads to the wrong database or share the encoding cache. Direct factory tests cover database and cache isolation. Real PostgreSQL suites cover query and transaction behavior. - Session writes must retain their issue row locks and generation fences. The original function bodies are unchanged. New database tests reject stale and cancelled conversation writes and clears. - Existing SQL_ASCII safeguards and bounded projections remain in place. No schema, API contract, migration, lockfile, or workflow changes are included. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8e59efc50b |
fix(ui): keep sidebar plugin outlets quiet when the server is unreachable (#15588)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The board UI sidebar shows plugin contributions: plugin nav slots,
plugin launchers, and the plugin panel at the bottom
> - These outlets read `GET /plugins/ui-contributions` through React
Query
> - When the server is unavailable, the refetch fails. Each outlet then
replaces its content with a red "Plugin extensions unavailable: Failed
to fetch" or "Plugin launchers unavailable" box
> - React Query still has the last good data, so the error box removes
useful content and adds noise in a persistent part of the UI
> - This pull request adds an option to hide the outlet error, and the
sidebar uses it
> - The benefit is a quiet sidebar during a server outage. The sidebar
keeps the last loaded plugin items. Other pages still show the error
inline
## Linked Issues or Issue Description
No public issue exists. I searched open and closed issues and PRs for
related work and found none.
**What happened?**
When the server is unavailable, the sidebar shows red error boxes in
place of plugin items. The text is "Plugin extensions unavailable:
Failed to fetch" or "Plugin launchers unavailable: Failed to fetch".
**Expected behavior**
The sidebar does not show errors while the server is unavailable. It
keeps the last loaded content.
**Steps to reproduce**
1. Start Paperclip with at least one plugin that contributes a `sidebar`
slot, a `sidebar` launcher, or a `sidebarPanel` slot.
2. Open the board UI and let the sidebar load.
3. Stop the server.
4. Wait for the next refetch of `/plugins/ui-contributions`, or focus
the window.
5. See the red error boxes in the sidebar.
**Paperclip version or commit**
`master` at
|
||
|
|
53105d5830 |
feat(apps): add Gauge connection (#15596)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents reach external services through the Apps catalog. Each catalog entry is a reviewed `AppDefinition` that connects a provider's hosted MCP server to Paperclip's shared vault, grants, policies, gateway, and audit trail. > - Gauge (`withgauge.com`) measures how AI answer engines mention and cite a brand. Its hosted MCP server also gives SEO, analytics and ads reports, and runs a content pipeline that can publish to a connected CMS. > - Gauge is not in the catalog. Operators must use the generic "Connect your own MCP server" flow. That flow has no branding, no organization guidance, and no warning about live CMS publishing. > - The connector playbook supports this provider with existing definition fields. The server supports dynamic client registration with a reviewed `mcp` scope, and it accepts an organization API key as a bearer header. > - This pull request adds the Gauge definition, its artwork, the research and permission-review ledger rows, documentation, and deterministic tests. > - The benefit is a branded, governed Gauge connection with browser sign-in, an API-key option, and a clear publish warning. ## Linked Issues or Issue Description **Problem or motivation** Marketing and growth teams use Gauge to track their brand in AI answers and to run content workflows. They want their Paperclip agents to read visibility, keyword and traffic data, and to prepare content. Gauge is not in the Apps catalog. Operators must paste the MCP URL into the generic remote-MCP flow, which gives no branding, no method guidance, and no warning that content tools can publish live. **Proposed solution** Add a catalog-only Gauge connection built from the connector playbook. Browser sign-in uses Gauge's dynamic client registration and requests only the reviewed `mcp` scope. The user selects one Gauge organization on the consent screen. An optional method sends a customer-created organization API key as an `Authorization: Bearer` header. Both methods warn the operator to set publish actions to Ask first. Every discovered tool stays governed by the normal per-action policies. **Alternatives considered** A plugin was not needed because no custom UI, tables, workers, or webhooks are involved. The identity scopes that Gauge also advertises (`openid`, `profile`, `email`, `organizations`) are not requested, because Gauge selects the organization on its consent screen. A Gauge-specific `classifyRisk` rule was not added, because the tool names cannot be seen without an account. The generic rule already classifies `publish`, `create` and `update` tools as writes. **Roadmap alignment** This extends the existing self-serve remote-MCP connection catalog and does not overlap planned core work. ## What Changed - Added the `gauge` provider to `scripts/ingest-app-definitions.mjs` (category `analytics`, API-key placement, guidance, warnings, description) and regenerated `packages/shared/src/app-definitions/gauge.json` and the generated registry. - Added the Gauge row to the self-serve MCP research ledger with `dcr_or_api_key` auth and risk tier S3. - Added permission reviews for `gauge/mcp-oauth` (explicit scope `mcp`, from Gauge's live authorization-server metadata) and `gauge/mcp-api-key` (provider key), with evidence links. - Added Gauge's mark (`ui/public/brands/apps/gauge.png`, the avatar of Gauge's official GitHub organization) and the brand manifest entry. - Added gallery copy for the Gauge card. - Added `doc/connections/GAUGE.md` (endpoints, scopes, administrator setup, capabilities and policy, manifest, brand provenance, validation hook) and linked it from the connections README and the permission audit. - Tests: Gauge definition shape, store visibility and artwork, URL recognition, reviewed scope, bearer-header placement, the connect form's sign-in default and API-key gating, and the pinned catalog counts. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts packages/shared/src/app-definitions-url.test.ts ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx server/src/__tests__/tool-access-service.test.ts` — 738 passed. - `node scripts/check-app-brand-assets.mjs` and `node --test scripts/app-brand-validation.test.mjs` — passed. - `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter @paperclipai/server typecheck`, `pnpm --filter @paperclipai/ui typecheck` — clean. - `node scripts/ingest-app-definitions.mjs --definitions-only` produces no Gauge drift. - Manual: on a local instance, open Apps → Browse and confirm the Gauge card and icon. Open `/apps/connect?source=gauge` and confirm that sign-in is the default and that the API-key method is under Advanced. - Live metadata probe on 2026-10-08: an unauthenticated `initialize` on `https://app.withgauge.com/mcp` returns 401 with `resource_metadata`. Both `.well-known` documents return the recorded endpoints and scopes. Dynamic client registration succeeds. The authorize endpoint accepts `scope=mcp` and rejects an unknown scope with HTTP 400. ## Risks - Low risk to existing providers: the change is additive catalog data plus tests. The generated registry only gains one import. - Gauge content tools can publish to a connected CMS, including live, and a bulk keyword update replaces each prompt's keyword list. Both methods warn the operator, and the API-key helper text tells operators to set publish actions to Ask first. - Gauge API keys are not scoped and reach the whole organization. Paperclip cannot narrow an issued key. - The permission-review ledger records live proof for both methods as not run. The lifecycle checklist in `doc/connections/GAUGE.md` still needs a documented pass with a Gauge organization. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with tool use (shell, file editing, web research). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0111689d0a |
test(ui): cover monitor-only task policies (#15587)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators open tasks to read progress and edit reviewers and approvers. > - Stored execution policies can contain a monitor and omit the stages array. > - These policies previously crashed the task properties panel. > - The policy reader now applies schema defaults through the fix in #15593. > - This pull request adds regression coverage for monitor-only policies with external service metadata. > - The tests check that the panel renders and participant edits keep the existing monitor. ## Linked Issues or Issue Description **What happened?** A task with a monitor-only execution policy could crash with `Cannot read properties of undefined (reading 'find')`. The stored policy omitted `stages`. **Expected behavior** The task shows no reviewers or approvers. Adding both kinds of participant keeps the monitor and its external service metadata. **Steps to reproduce** 1. Load a task whose execution policy contains an external service monitor and omits `stages`. 2. Render its properties panel and inspect both participant controls. 3. Add reviewers and approvers. Check that the resulting policy keeps the monitor. Refs #15593, which provides the policy normalization used by these tests. ## What Changed - Test that both stage types return empty participant lists for a monitor-only policy. - Test adding reviewers and approvers while keeping the external service monitor. - Test retaining a monitor when no review or approval stages exist. - Render the properties panel with omitted stages and check its participant and monitor controls. ## Verification - The helper and properties panel suites pass on current master with these tests: 108 tests. - `pnpm check:token-gates` passes. - All CI gates pass on the updated PR commit, including build, types, Runner verification, full test suites, browser suites, security scans, and canary checks. - Greptile gives commit `2634bd1d` 5/5 with no actionable findings or unresolved review threads. - The full local workspace suite was not repeated for this test-only rebase. CI verifies the full suite, build, types, Runner checks, browser suites, security scans, and canary run. ## Risks - Low risk. This change adds tests against the existing policy behavior. - The fixture uses example external references and test participant IDs. - Existing documentation remains accurate for this regression coverage. ## Model Used - OpenAI Codex, based on GPT-6, with tool use and code execution. The exact serving model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0d57b9b98e |
refactor(heartbeat): extract run preparation (#15591)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service prepares context and configuration before agent runs. > - The workspace extraction left the main file at about 27,000 lines. > - Preparation helpers form another large boundary with explicit inputs and one database dependency. > - This pull request moves those helpers into one module and preserves their callers. > - The benefit is a 24,000-line orchestration file and one place to maintain run preparation. ## Linked Issues or Issue Description **What existing behavior does this improve?** It improves the structure and test coverage of heartbeat run preparation. **Current behavior** `server/src/services/heartbeat.ts` mixes context and configuration preparation with execution, queueing, and recovery in 27,048 lines. **Proposed behavior** Move preparation into `server/src/services/heartbeat/run-preparation.ts`. The main file loses 3,048 lines and ends at 24,000 lines. Keep the existing imports and behavior. This continues #15568, #15573, and #15578. **Reason and benefit** A larger domain extraction makes useful progress toward a few manageable files. Context, identity, environment, and tool preparation can be reviewed together without searching the run engine. **Breaking changes** None. Existing public helpers and the configuration-incomplete error class retain their identity. ## What Changed - Move wake payloads, comments, attachments, skill mentions, adapter environment configuration, and MCP/tool setup into `heartbeat/run-preparation.ts`. - Bind issue context, pinned routine snapshots, organization rows, and responsible-user resolution through `createHeartbeatRunPreparation(db)`. - Keep factory construction free of database work. Keep queueing, dispatch, retries, cancellation, and execution order in `heartbeat.ts`. - Preserve all 140 existing exports. A syntax-tree comparison confirms that the 47 extracted function bodies and 25 moved declarations are unchanged. The private string normalizer is also unchanged. - Add six tests for legacy export identity, independent database binding, company scope, pinned routine configuration, wake author selection, company-owner fallback, and credential preflight. - Document the preparation boundary in `doc/DEVELOPING.md`. ## Verification - Passed: 183 focused tests across ten suites, including the six new tests. Command: `pnpm exec vitest run server/src/services/heartbeat/run-preparation.test.ts server/src/__tests__/heartbeat-context-summary.test.ts server/src/__tests__/heartbeat-agent-session-message.test.ts server/src/__tests__/heartbeat-runtime-mcp-servers.test.ts server/src/__tests__/heartbeat-runtime-skills.test.ts server/src/__tests__/heartbeat-project-env.test.ts server/src/__tests__/heartbeat-responsible-user-invariant.test.ts server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/run-secret-redaction.test.ts server/src/__tests__/low-trust-red-team-routes.test.ts`. - One runtime-skills case reported a Postgres deadlock in its after-test `TRUNCATE` cleanup. A focused rerun passed both runtime-skills cases. The other nine suites passed in the combined run. - Passed: `pnpm -r typecheck`. - Passed: `pnpm build`. - Passed: all 54 GitHub checks on `8d0cb9a251b10be5735fa13035e8a31e50b132e7`. The two optional Storybook jobs were skipped. The complete CI test matrix is green. - The full local `pnpm test:run` was started and then stopped after the complete CI test matrix passed. It had no failures reported before stopping and did not complete locally. - Greptile: 5/5 on the same commit, with no inline comments or actionable findings. ## Risks The extraction crosses context, identity, credential, and tool-access module boundaries. Existing company filters, secret redaction, low-trust rules, and error classes stay intact. The new factory captures only the service database and does no database work during construction. Existing integration tests cover responsible-user authority, credential boundaries, MCP tokens, and quarantine. New tests cover the factory wiring and legacy identity. A syntax-tree comparison confirms that orchestration changed only to bind the extracted loaders. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9f4acc8f14 |
build(deps): bump @modelcontextprotocol/sdk from 1.30.0 to 1.31.0 (#15361)
Bumps [@modelcontextprotocol/sdk](https://github.com/modelcontextprotocol/typescript-sdk) from 1.30.0 to 1.31.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/modelcontextprotocol/typescript-sdk/releases">@modelcontextprotocol/sdk's releases</a>.</em></p> <blockquote> <h2>1.31.0</h2> <h2>Upgrade notes</h2> <ul> <li>Stored OAuth tokens and client information now include an <code>issuer</code> field. Storage that rejects unknown fields needs to allow it.</li> <li>Pass <code>expectedIssuer</code> when constructing <code>ClientCredentialsProvider</code>, <code>PrivateKeyJwtProvider</code> or <code>StaticPrivateKeyJwtProvider</code>. Constructing them without it is deprecated.</li> </ul> <h2>What's Changed</h2> <ul> <li>[v1.x] Bind stored OAuth credentials to the authorization server that issued them by <a href="https://github.com/maxisbey"><code>@maxisbey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2888">modelcontextprotocol/typescript-sdk#2888</a></li> <li>chore: bump version to 1.31.0 by <a href="https://github.com/claude"><code>@claude</code></a>[bot] in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2890">modelcontextprotocol/typescript-sdk#2890</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.1...1.31.0">https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.1...1.31.0</a></p> <h2>1.30.1</h2> <h2>What's Changed</h2> <ul> <li>[v1.x] fix(server): read HTTP request bodies with a size limit and bound JSON-RPC batch length by <a href="https://github.com/maxisbey"><code>@maxisbey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2717">modelcontextprotocol/typescript-sdk#2717</a></li> <li>fix(auth): preserve resource URI without trailing slash (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1968">#1968</a>) by <a href="https://github.com/MukundaKatta"><code>@MukundaKatta</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1972">modelcontextprotocol/typescript-sdk#1972</a></li> <li>chore: bump version to 1.30.1 by <a href="https://github.com/claude"><code>@claude</code></a>[bot] in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2848">modelcontextprotocol/typescript-sdk#2848</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/MukundaKatta"><code>@MukundaKatta</code></a> made their first contribution in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1972">modelcontextprotocol/typescript-sdk#1972</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.30.1">https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.30.1</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/4b0051f400219f8d8855f9a5433c6df35f15a639"><code>4b0051f</code></a> chore: bump version to 1.31.0 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2890">#2890</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/51ad4f03190dc5c84ab8b1a25f9b78b277be0dc7"><code>51ad4f0</code></a> [v1.x] Bind stored OAuth credentials to the authorization server that issued ...</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/289ac2c3af7e1536160e80414296b175171a1a87"><code>289ac2c</code></a> chore: bump version to 1.30.1 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2848">#2848</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/12b425678a76cd54b0452a2ccf1e5dc7740f73ef"><code>12b4256</code></a> fix(auth): preserve resource URI without trailing slash (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1968">#1968</a>) (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1972">#1972</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/a9f6eb709b85459d01e8f2a9e881fef2621756c1"><code>a9f6eb7</code></a> [v1.x] fix(server): read HTTP request bodies with a size limit and bound JSON...</li> <li>See full diff in <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.31.0">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>canary/v2026.1008.0-canary.17 |
||
|
|
317ed367d4 |
Handle incomplete task policies and preserve review limits (#15593)
Apply the reviewed change for Handle incomplete task policies and preserve review limits. Validation: local typecheck/build and focused regression tests, passing exact-head CI, Greptile 5/5, and independent source review. The PR records full-suite evidence and any local environment limitations. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.16 |
||
|
|
3d0e74e743 |
Keep proven external connection failures out of Sentry (#15590)
Apply the reviewed change for Keep proven external connection failures out of Sentry. Validation: local typecheck/build and focused regression tests, passing exact-head CI, Greptile 5/5, and independent source review. The PR records full-suite evidence and any local environment limitations. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.15 |
||
|
|
0ac194450a |
fix: make Copilot provider-pack wrappers portable (#15586)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Container images include a movable provider pack for agent execution. > - pnpm executable wrappers can contain the temporary build directory. > - New optional Copilot packages add wrappers that the build does not replace. > - This pull request gives those installed wrappers relative executable paths. > - Image publication can finish while the existing path check remains enforced. ## Linked Issues or Issue Description Refs #15572 and #15560. The small portability loop comes from cryppadotta's larger Copilot runtime PR #15560. This separate fix repairs image publication without waiting for that feature's qualification and runtime changes. That PR can remove its duplicate loop after this lands. **What happened?** The standard Docker build stops with `Provider pack shim copilot-linux-x64 retains its temporary build path`. The failure occurred before and after #15522. See [the failed master build](https://github.com/paperclipai/paperclip/actions/runs/37802065316). **Expected behavior** Installed native Copilot wrappers resolve their pinned executable after the provider pack moves. Optional packages that are absent do not gain a command. The builder still rejects wrappers with temporary paths. **Steps to reproduce** 1. Run the provider-pack build from the affected master revision on Linux x64. 2. Let `pnpm deploy --prod` install the optional Copilot package. 3. The wrapper scan rejects its temporary `NODE_PATH`. **Paperclip version or commit** Master `3367b75ccce34d02f355cda1f1ed3fe0b34cf93d`. **Deployment mode** Docker and provider-pack builds. ## What Changed - Replace installed Copilot platform wrappers with relative native executable launchers. - Move the existing executable wrapper writer into an importable helper. Retain the same launch behavior for Node, Claude and OpenCode. - Test relocation, argument handling, exit status, absent optional packages and missing executable packages. Register the tests in the existing Runner test preparation command. - Hash the helper in Daytona image identity and prove that helper changes invalidate the image cache. - Document the packaging rule. Keep dependencies, provider qualification and the final temporary-path check unchanged. ## Verification - 17 Node packaging tests passed across the new wrapper tests, provider-pack release tests and candidate selection tests. - Six existing bundled remote-provider-pack tests passed using the Runner Vitest configuration. - Nine Daytona image identity tests and Runner E2E typecheck passed after the cache-input correction. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - Real `pnpm deploy --prod` reproduction on macOS ARM64: the original Copilot wrapper contained the temporary path. After the rewrite and directory relocation, the actual executable returned Copilot CLI 1.0.88 with exit code zero using only `/usr/bin:/bin` in `PATH`. - Syntax checks, `git diff --check` and the pre-push secret scan passed. - The full local `pnpm test:run` reproduced the same five skill/connector fixture failures observed earlier in this workspace. It was stopped after current-head clean-checkout CI passed; later local phases were not run. This local run is not claimed as passing. Focused packaging tests, workspace typecheck/build and all hosted CI passed. - [Hosted Docker verification passed](https://github.com/paperclipai/paperclip/actions/runs/37804902678): Linux AMD64 and ARM64 image builds, multi-architecture publication and the process-reaping smoke check. This run tested `779d94989c54d9abbeba0838194c951186af66a6`; the only later changes are Daytona cache identity and its regression test. Provider-pack build code is identical. Local Docker did not respond within the bounded probe. - Current head `db2c19e270b8d5a7bab5db39d3de0e4761cf57c3`: 54 successful checks, two conditional skips, no failures and no merge conflicts. Apex is 5/5 with no unresolved comments. ## Risks The helper uses each installed package's exported executable. An installed wrapper with a missing package still fails the build. This change does not execute Copilot during image construction, alter dependency pins, change runtime admission or weaken the temporary-path check. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use and code execution. The exact serving model identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; the full local-suite limitation is documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.14 |
||
|
|
75918997a8 |
refactor(heartbeat): extract workspace preparation and resolution (#15578)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service prepares workspaces and dispatches agent runs. > - Its main file still has more than 30,000 lines after two small extractions. > - Workspace preparation forms a larger boundary with one database dependency. > - This pull request moves that code into one workspace module with focused tests. > - The benefit is a smaller orchestration file and one place to maintain workspace policy. ## Linked Issues or Issue Description **What existing behavior does this improve?** It improves the structure and test coverage of heartbeat workspace preparation. **Current behavior** `server/src/services/heartbeat.ts` mixes workspace preparation, run execution, and scheduling in 30,634 lines. **Proposed behavior** Move workspace preparation into `server/src/services/heartbeat/workspaces.ts`. Keep the public imports and run behavior stable. The main file loses about 3,540 lines in one extraction. Related extractions: #15568 and #15573. **Reason and benefit** A domain-level extraction makes useful progress toward a few manageable modules. Future workspace changes can be reviewed without searching the whole run engine. **Breaking changes** None. The existing public exports and workspace validation error class retain their identity. ## What Changed - Move managed checkout preparation, workspace validation and reuse, referenced project resolution, and session/workspace config freshness into `heartbeat/workspaces.ts`. - Bind run workspace resolution to the database through `createHeartbeatWorkspaceResolver(db)`. - Keep the checkout single-flight map at module scope so all callers share pending materialization. - Keep all 140 existing `heartbeat.ts` exports. The 84 moved function bodies and 53 other moved declarations are unchanged in a syntax-tree comparison. - Add seven tests for legacy export identity, database binding, issue selection, session fallback, concurrent checkouts, same-name repositories, and retry after failure. - Document the module boundary in `doc/DEVELOPING.md`. ## Verification - Passed: seven focused workspace suites, 264 tests. Command: `pnpm exec vitest run server/src/services/heartbeat/workspaces.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/heartbeat-referenced-projects.test.ts server/src/__tests__/heartbeat-remote-referenced-projects.test.ts server/src/__tests__/heartbeat-workspace-branch-containment.test.ts server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts server/src/__tests__/heartbeat-project-env.test.ts`. - Passed: `pnpm -r typecheck`. - Passed: `pnpm build`. - Passed: all 54 GitHub CI checks on `14dec775c7fa83716432f73740bd697106fdef68`. The two optional Storybook jobs were skipped. The full CI test matrix is green. - The full local `pnpm test:run` hit a 15-second timeout in the first `public-mcp.test.ts` case. A focused rerun passed all 80 public MCP tests. The long local run was stopped after all CI test shards passed; it did not complete locally. - Greptile: 5/5 on the same commit, with no inline comments or actionable findings. ## Risks Module initialization and shared checkout state are the main extraction risks. Legacy exports retain the same function and class objects. The checkout map stays outside the database factory. Existing integration suites cover company boundaries, workspace containment, reuse, and referenced project authorization. New local Git tests cover concurrent materialization and failure cleanup. The orchestration body is unchanged except for the resolver factory binding. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.13 |
||
|
|
3367b75ccc |
fix: fence accepted work and cleanup before idle sleep (#15522)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A host can stop an idle instance to reduce unused compute. > - Zero active runs do not prove that requests, saved work, or cleanup are complete. > - A client can disconnect while its request still writes, and failed cleanup can remain only in memory or on disk. > - This pull request holds new admission and checks accepted work, cleanup, and persisted work under an owned drain. > - The host gets an empty report only while those checks remain valid. ## Linked Issues or Issue Description Refs #13413. That companion PR uses the same owned-hold vocabulary for runtime services. This PR covers HTTP requests, scheduler work, accounting, cleanup, and the instance work inventory. It does not include the preview gateway changes. **What existing behavior does this improve?** The instance task-drain API reports process counters. It does not prove that a host can safely stop the instance for idle sleep. **Current behavior** A quiet run set can coexist with an unfinished request, cleanup after a completed run, future work, or a failed accounting write. **Proposed behavior** Provide a bounded owned idle hold. Block new ingress, track accepted handler promises, inspect durable and local work, and return `none` only when the same hold stays quiet through the checks. Keep normal deployment drains compatible. **Reason and benefit** Hosts can identify eligible idle instances without treating a disconnect or failed cleanup write as completed work. ## What Changed - Add `purpose: "idle"`, a bounded TTL, unique owners, and owner-checked release to task drain. Existing holds cannot be replaced by another API request. - Gate HTTP ingress before parsers, auth, webhooks, and MCP. Gate new WebSocket upgrades. Track async handlers in nested Express routers and error middleware until they settle, even after the response or client disconnect. - Count accepted live-event WebSocket authentication through settlement, even after disconnects. Count detached built-in agent, managed-home and runtime-service startup reconciliation after readiness. - Keep health and control mutations tracked. Count control-request authentication separately from the read-only report, including concurrent user/company/membership writes. - Pause new scheduler admissions during idle holds. Count work already in flight, including database backups, and reject scans whose work generation changes. Periodic backups block sleep without a host wake schedule. - Inspect accounting and orphan-cleanup spool directories without skipping temporary or malformed entries. Retain orphan tokens through queue splices, flush failures, and buffer overflow. Keep failed usage capture counted until its database failure fence is written. - Check persisted work across companies in a bounded read-only transaction. Include deferred agent-file cleanup and saved watchdogs whose watched issues are complete. Enabled plugins and unsupported retained work remain blockers. - Require the exact idle owner and completed startup tracking before returning an empty report. Keep reports free of tenant details. - Document the hosting protocol, retry behavior, conservative blockers, and the remaining external provider-stop race. ## Verification - Current head: `8d9a599b732c2047b8c671d4799065d7c90a3567`, rebased on master `941a3fa991aeb97eb1ac390c65b7973b5f6de1ad`. Both heartbeat helper extractions are preserved. GitHub confirms no merge conflicts. - Focused heartbeat renderer/run-log, drain, control-auth, admission, route and PostgreSQL inventory checks: 234 tests passed in 10 suites after the rebase. - The POST task-drain contract includes the expected `409` conflict response. The 84 OpenAPI and instance-settings route tests and server typecheck passed after that final documentation fix. - Final accepted-upgrade/startup regression run: 104 tests passed in five suites, including success and failure after readiness or disconnect. Final server typecheck also passed. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm install --frozen-lockfile` and the PR diff secret scan passed. No dependency or lockfile changes are added by this PR. - The prior verification covered HTTP disconnects and early responses, async error handlers, concurrent authentication writes, signed bootstrap, saved watchdogs, backup promises, orphan cleanup, spools and failed accounting fences. Those tests passed. - The last full local `pnpm test:run` stopped in the general-server phase with 16,658 passing tests and five skill/connector fixture failures caused by an ancestor workspace skill directory. That full local run preceded this rebase and has not been repeated for the import conflict. Full current-head CI passed: 53 successful checks and two conditional skips, with no failures. - All eight review findings are fixed and their threads resolved, including accepted upgrade authentication, detached startup writes and the POST conflict contract. Current-head Apex review is 5/5 with no new findings; all eight review threads remain resolved. - No live provider stop or production change was performed. ## Risks - The hosting controller must use the owned protocol and hold external admission through its final validation and provider stop. A legacy drain cannot authorize idle sleep. During an idle hold, new requests receive 503 with `Retry-After: 1`; the host must handle queueing or retry before enabling this path. - This is a single-process protocol. An unexpected restart after the last validation can race an external provider stop. The host must serialize deploy/wake/sleep operations and bind the validation to the instance it stops. Multiple replicas need shared fencing. - Some retained state conservatively prevents sleep, including every enabled plugin. This PR does not promise that every inactive instance becomes eligible. - The HTTP adapter uses Express 5 router layers. Real Express tests cover nested routes, errors and disconnects. New routes must register before tracking is installed. Detached work must have durable state or explicit work tracking. - Unrecoverable in-memory cleanup debt keeps the instance awake until reconciliation. This change does not make such debt survive an unplanned process crash. No schema migration or provider configuration change is included. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use and code execution. The exact deployment variant/model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; the full local fixture limitation and clean full CI result are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
71af2fbc3b |
fix(slack): teach agents how people connect their accounts (#15576)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack connections let people talk to an agent with their own Paperclip permissions. > - Each bot has a saved command that starts account linking. > - The Access page explains this flow, but Slack agent turn context omitted it. > - An agent could guess the command or confuse channel membership with Paperclip access. > - This pull request gives agents the saved command and account confirmation instructions. > - People can ask the agent how to join without changing access controls. ## Linked Issues or Issue Description **What happened?** Slack agent context explained messages, tools, and questions. It did not explain how a teammate can connect an account. The saved slash command can differ from the current agent name. **Expected behavior** Give the saved bot command, such as `/research_ops connect`. The teammate runs it themselves. They sign in through a private confirmation link. A new company member must receive admin approval before account confirmation. The connection manager can copy instructions from Access → Invite people. **Steps to reproduce** 1. Configure a Slack bot with a custom slash command. 2. Rename its assigned agent. 3. Inspect a fresh or resumed Slack task prompt. Before this change, it contains no account invitation instructions or saved connect command. **Paperclip version or commit** Base: `d6df12cef69fcaf2d2fe66a393168931d5b8b4e7`. **Deployment mode** Applies to local and hosted Slack chat connections. Deterministic tests used a local isolated database. The new model probe has not run against a live Slack bot. Related work: Refs #15413 and #13638. This fixes agent guidance for their existing account-linking flow. It adds no new membership system or invitation endpoint. ## What Changed - Read only the saved public slash command from the company-scoped conversation endpoint. Supply it only to the assigned agent for Slack turns with an active or verifying connection. - Validate the command with the shared Slack configuration schema. Refer to Access → Invite people when it is missing or invalid. Never guess from the current agent name. - Explain personal account confirmation, link expiry, and company membership approval on fresh and resumed turns. Distinguish channel invitations from Paperclip access. - Add deterministic guidance regressions, a manual invitation model probe, and setup documentation. ## Verification - Passed: 63 tests in `heartbeat-context-summary.test.ts` and `heartbeat-chat-task-link.test.ts`. - Passed: four real-heartbeat regressions in `heartbeat-slack-invitation.test.ts`. They check the saved command after an agent rename, a persisted resumed session, full and compact prompts, missing-command fallback, inactive endpoints, and a different assigned agent. They also reject caller-supplied command fields and exclude other setup metadata. - Passed: the isolated `chat-channels.integration.test.ts` case `discovers a Slack connect identity without starting work or granting access`. It covers the private link, duplicate connect requests, and nonmember access requests without a membership grant. - Passed: `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/server build`. - Passed: `pnpm test:slack-connector --list` and `git diff --check`. - Passed: repository `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. Typecheck required local IPC access for the migration check. - Passed: final-commit CI, with 54 successful checks and two optional Storybook checks skipped. CI includes all server, chat, workspace, serialized, Runner, and browser test shards, typecheck, build, and canary dry run. - The full local `pnpm test:run` was started. It was stopped after full CI passed; it had no final local summary. The focused invitation checks passed locally. Do not count the interrupted local run as a full-suite pass. - Greptile gave the exact final commit `f37e52bde621ef7f5b2bb345074d53345f8e4ee9` a 5/5 score. Both review findings were fixed and their threads resolved. - The new `invite-person` model probe is manual. These deterministic results do not establish a live model or Slack acceptance pass. ## Risks - Model guidance cannot prove that a person joined. The existing account-linking and membership checks remain authoritative. - Legacy rows without a saved command use the Access page fallback. Invalid command text is excluded from the prompt. - The query selects only the public command. It does not expose registration secrets, tokens, or personal confirmation links. - No schema changes, new provider requests, permission grants, or telemetry changes. ## Model Used - OpenAI GPT-6 through Codex. The exact backend variant and context window are not exposed in this session. Used reasoning, repository inspection, code editing, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.12 |
||
|
|
d6df12cef6 |
fix(ui): add icons to chat connection choices (#15574)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connectors let people chat with agents and let agents use external tools. > - Slack setup asks the user to choose between those two uses. > - Both choices currently use text alone. > - This pull request adds an agent avatar and a hammer to the shared chooser. > - The icons help users identify each choice at a glance. ## Linked Issues or Issue Description **What existing behavior does this improve?** The chat-or-tools choice in connector setup. Related setup work: Refs #15413. **Subsystem affected** ui/ — React board UI and Storybook. **Current behavior** Both choice cards show a title and description with no identity icon. **Proposed behavior** Show the existing agent avatar beside chat. Show a hammer beside tools. Align both icons and keep the existing navigation. **Reason and benefit** Users can tell the two choices apart with less reading. **Breaking changes** None. The shared chooser applies the icons wherever that choice appears. Only Slack currently offers both paths in the catalog. Other chat-only providers keep their direct setup path. ## What Changed - Add the shared agent avatar to the chat choice. - Add a decorative hammer icon to the tools choice. - Add desktop and mobile Storybook stories that render the production page. ## Verification - Token gates pass. - Focused catalog, routing, and UI contract tests pass, including the new chooser accessibility and navigation test. - Checked the production chooser in Storybook at desktop and mobile widths. - Local repository typecheck, build, and Storybook build pass. - All 56 checks pass on commit `0f0673f1bb5cd340084f3a5229983da049f1c8db`, including the full CI test suite and browser shards. Greptile reports 5/5 with no findings. - Started the full local test command. Stopped that duplicate run after the full CI suite passed. The focused local tests completed successfully. - In Storybook, open Connections → Slack → Automatic setup → 00 · Choose connection. The mobile variant is in the same group. ## Risks Low risk. This changes card layout and decorative images. Existing button text, click handlers, credentials, and permissions stay the same. The avatar is generic because the user has not selected an agent yet. ## Model Used OpenAI GPT-6 through Codex. Used code inspection, edits, shell tools, and browser tools. The session does not expose the exact runtime model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
941a3fa991 |
Pin Copilot native dependencies with the maintained lockfile refresh (#15572)
Pin the three optional GitHub Copilot 1.0.88 native packages for Runner and server using the maintained lockfile workflow. Synchronize the package contract and bound initial render readiness in the deliberately throttled browser fixture. Current-head CI and focused checks pass. Co-Authored-By: Dotta <cryppadotta@users.noreply.github.com> Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
87a7312cfd |
refactor(server): extract heartbeat run-log formatting (#15573)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Heartbeat orchestration records agent output and run events. > - The heartbeat service still contains more than 30,000 lines after the first extraction. > - Log formatting has fixed limits and does not write run state. > - These helpers form a small next step toward a more manageable heartbeat service. > - This PR moves them into the heartbeat folder without changing their function bodies. > - Direct boundary tests and existing caller tests check the output. ## Linked Issues or Issue Description Refs: #15568. **What existing behavior does this improve?** This improves the structure and test coverage of heartbeat run-log formatting. **Subsystem affected** server/ — orchestration services. **Current behavior** `heartbeat.ts` contains excerpt handling, payload size limits, and log chunk formatting beside run orchestration. **Proposed behavior** Move these helpers and their constants into `server/src/services/heartbeat/run-log.ts`. Keep the same function bodies, output, limits, and public exports. **Reason and benefit** Run-log formatting has its own small module and focused tests. This keeps the second extraction small enough to review on its own. **Breaking changes** None. The existing public import path and persisted output stay the same. Related run-log PRs: Refs: #6373, Refs: #8841. Those PRs change redaction behavior. This PR only moves existing formatting code. ## What Changed - Move 109 lines of helper functions and six constants into `heartbeat/run-log.ts`. - Move `appendExcerpt` and retain the existing public exports for `boundHeartbeatRunEventPayloadForStorage` and `compactRunLogChunk` in `heartbeat.ts`. - Add 17 direct regression cases for size and depth limits, cycles, shared references, immutable input, image omission, redaction order, excerpt tails, and UTF-8 boundaries. - Document both heartbeat extractions in `doc/DEVELOPING.md`. ## Verification - Before extraction, four focused files passed with 55 tests. - After extraction and the two new excerpt tests, the same four files passed with 57 tests. - Run `pnpm exec vitest run server/src/services/heartbeat/run-log.test.ts server/src/__tests__/heartbeat-run-log.test.ts server/src/__tests__/heartbeat-list.test.ts server/src/__tests__/redaction.test.ts`. - The original and extracted function blocks match byte for byte after adding the export keyword to `appendExcerpt`. - Existing tests still import the public helpers from `heartbeat.ts`. - `node scripts/check-module-boundaries.mjs` and `git diff --check` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - The full suite passed in CI. The duplicate local `pnpm test:run` was stopped after CI passed; it did not complete locally. - All CI gates passed on commit `4e65c666e3b5b46162b3a528b37fba8902a1d200`: 54 passing checks, two skipped Storybook checks. - Fresh Greptile review of the same commit: 5/5, no actionable or inline findings. The PR has no merge conflicts. ## Risks - Low risk. Moving code can cause an import or build error. - The same redaction and adapter utility modules remain in use. - Database writes, current-user redaction, live event delivery, and run state remain in `heartbeat.ts`. - The event schema and payload output do not change. - No database, API, Telemetry, Observability, or UI contract changes are required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
51653becc2 |
refactor(server): extract heartbeat task markdown rendering (#15568)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Heartbeat orchestration supplies each run with task context. > - The task Markdown renderer is inside a service with more than 31,000 lines. > - The renderer formats task data and does not write run state. > - It is a small first step toward a more manageable heartbeat service. > - This PR moves the renderer without changing its function body or public export. > - Direct tests and existing caller tests check the output before and after the move. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the structure and test coverage of heartbeat task context rendering. **Subsystem affected** server/ — orchestration services. **Current behavior** `buildPaperclipTaskMarkdown` occupies 352 lines inside `heartbeat.ts`. The service also owns run execution, scheduling, recovery, and cancellation. **Proposed behavior** Move the renderer to `heartbeat/task-markdown.ts`. Keep its function body and the existing `heartbeat.ts` export unchanged. Leave orchestration in place for this first PR. **Reason and benefit** Prompt formatting has its own small source file and direct regression tests. Reviewers can check this one extraction before a later refactor. **Breaking changes** None. The input type, rendered output, and existing import path stay the same. Related renderer changes: Refs: #4732, Refs: #14030. These PRs change prompt behavior. This PR only moves the current renderer. ## What Changed - Move the 352-line renderer to `server/src/services/heartbeat/task-markdown.ts`. - Import and re-export it from `heartbeat.ts`. Remove imports used only by the renderer. - Add seven direct regression cases. Cover empty context, ordered comment-only wakes, input preservation, nested code fences, ancestor limits, attachment-only wakes, and rejected plans. - Keep the renderer and its direct tests in `server/src/services/heartbeat/`. Document this folder as the home for relevant later extractions in `doc/DEVELOPING.md`. ## Verification - The focused suite passed before and after extraction, and after moving to the heartbeat folder: four files, 85 tests. - Run `pnpm exec vitest run server/src/services/heartbeat/task-markdown.test.ts server/src/__tests__/heartbeat-context-summary.test.ts server/src/__tests__/heartbeat-chat-task-link.test.ts server/src/__tests__/codex-local-execute.test.ts`. - The extracted function matches the original function byte for byte. The folder move only adjusts imports. - `node scripts/check-module-boundaries.mjs` passed. - `git diff --check` passed. - The initial extraction passed local `pnpm -r typecheck` and `pnpm build`. - The folder update passed `pnpm --filter @paperclipai/server exec tsc --noEmit` and `pnpm --filter @paperclipai/server build`. - The initial extraction passed the full CI suite. Its duplicate local `pnpm test:run` was stopped after CI passed; it did not complete locally. - The folder update passed all CI gates on commit `e21e589b379d7ca2fae16c1dcd91cf2f604fd03f`: 54 passing checks, two skipped Storybook checks. Three test shards passed after one retry following simultaneous runner shutdowns; those failures had no failed test assertions. - Fresh Greptile review of commit `e21e589b379d7ca2fae16c1dcd91cf2f604fd03f`: 5/5, no actionable or inline findings. ## Risks - Low risk. A moved module can change import resolution or expose an import cycle. - Existing caller tests still load the compatibility export from `heartbeat.ts`. - The same guidance constants and public task URL resolver remain in use. - No run-state writes, transactions, locks, shared process state, or cleanup paths moved. - No database, API, telemetry, or UI contracts changed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.11 |
||
|
|
24a15c5afa |
test(evals): verify durable waiting across continuation checkpoints (#15564)
## Thinking Path > - Paperclip manages agents and durable tasks across provider runs. > - Human questions must preserve task ownership and stop work until a real answer arrives. > - The continuation eval checks the final work, but its lifecycle oracle misses several broken waiting states. > - A correct final answer can hide a stale execution lock or a lost intermediate answer receipt. > - This pull request checks each wait and retains every question and run identity. > - A source audit records which waiting operations belong to native runtimes and which still require legacy API calls. ## Linked Issues or Issue Description **What existing behavior does this improve?** The existing Product E2E continuation oracle for human question and approval journeys. **Current behavior** A pending interaction can pass the waiting check even when its task has the wrong status, a stale lock, or a retry. Only the first pending question and first run receipts are checked at the final checkpoint. **Proposed behavior** Require a healthy wait on the same task and assignee. Preserve all intermediate run receipts and every pending question's answered identity. Allow a paused native provider question only when its pending runtime request identifies the running native run and execution lock. Related: #15548 and #15554 cover earlier bookkeeping slices. #15544 changes production continuation summaries; this PR changes the eval oracle and does not overlap that fix. ## What Changed - Check waiting state, execution ownership, and all answer/run identities in the existing lifecycle oracle. - Add negative calibrations for broken records and positive coverage for both semantic waits and paused provider questions. - Retain task activity at each checkpoint for inspection of successful persisted mutations. - Verify 1,000-cent company and agent budget limits before continuation work. Disable automatic cell rerolls. - Document runtime ownership, existing coverage, and the remaining instruction decision. Keep production instructions unchanged. ## Verification - `pnpm test:e2e:runner:unit`: 1,867 Vitest tests pass, one skips; 128 Node tests pass. - `pnpm test:e2e:runner:typecheck`: passes. - `pnpm -r typecheck`: passes. - `pnpm build`: passes. - Negative calibration: 27 added cases fail against the prior oracle and pass with these checks. - Four existing local continuation cells are selected for a separate bounded live canary. Live results are pending; no behavioral pass is claimed here. - The full local `pnpm test:run` suite was not repeated because embedded PostgreSQL was unavailable in the preceding workspace verification. Required Linux PR CI must pass before readiness. ## Risks The stronger oracle can expose existing product or fixture defects. A paused provider question and a terminal semantic wait have distinct valid states. The four-cell canary does not qualify approval/review, dependency unblock, crash races, remote execution, or general task quality. Task activity records successful persisted writes, not failed API attempts; repeated progress comments are not automatically defects. Historical eval grades remain unchanged. No production scheduling, prompt, tool, schema or migration changes. ## Model Used OpenAI Codex, GPT-6-based, with tool use and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f47614046d |
fix(slack): upload agent avatars directly during Cloud setup (#15566)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Slack setup creates a dedicated bot for an agent. > - The setup should upload that agent's avatar. > - The current upload asks Slack to fetch an image from the board origin. > - Cloud requires a tenant session at that origin, so Slack receives HTTP 401. > - This change renders the PNG on the server and uploads the file directly. ## Linked Issues or Issue Description Refs #15413. **What happened?** Automatic Slack setup left the app and bot with default icons on Cloud staging. An unauthenticated request to the exact avatar URL returned HTTP 401 with `tenant_session_required`. Local setup did not expose this Cloud ingress requirement. **Expected behavior** New Slack bots receive the assigned agent's 512-pixel avatar with the Paperclip dark background. **Steps to reproduce** 1. Create a Slack app through automatic setup on a Cloud tenant. 2. Complete installation. 3. Inspect the bot avatar in Slack. The previous URL-based upload cannot fetch the image without a tenant session. ## What Changed - Render the assigned agent's preset PNG with the existing bounded worker pool. - Send PNG bytes as multipart `file` data to `apps.icon.set` instead of passing a board URL. - Keep the temporary token in the Authorization header. Let fetch set the multipart boundary. - Recheck management permission and credential-lease ownership after rendering. Close the worker pool during chat service shutdown. - Cover actual PNG dimensions, uploaded bytes, failure recovery, and secret-safe responses. Update deployment documentation. ## Verification - Passed: 84 focused tests across automatic Slack registration and on-demand agent avatars. - Passed: full repository build. - Passed: full repository typecheck. All 54 current-head GitHub checks passed, including the complete test matrix, all browser shards, build, typecheck, canary, and security checks. Greptile completed on `79e79f6fc` with 5/5 and no actionable findings or open review threads. - The Cloud fetch failure was reproduced without browser credentials. No Cloud access rule was changed. - A real Slack upload with this new path still requires deployment and a fresh automatic setup. Existing apps retain the manual avatar-upload fallback. ## Risks - The renderer can time out or Slack can reject the upload. Both failures preserve the saved app and leave installation usable. - The renderer adds a bounded, lazy worker pool to Slack registration. Shutdown closes it. - No migration, bot permissions, credential retention, or Cloud authentication behavior changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser inspection. The runtime does not expose a more specific authoring model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.10 |
||
|
|
2f0c485dec |
fix(skills): ship the completion helper with the installed skill (#15554)
## Thinking Path > - Paperclip manages work for AI agents. > - Legacy agents use the Paperclip skill to save task status and comments. > - The skill names a script relative to the task workspace. > - That script exists only in the Paperclip source repository. > - Agents in other workspaces can hit a missing command or search for it. > - This PR ships the helper inside the skill and uses the installed skill path. > - The repository command remains available through a forwarding wrapper. ## Linked Issues or Issue Description Fixes #9527. Refs #15548 for the preceding runtime checkout guidance. Related: #6052 addresses LF line endings for the repository helper; this change addresses helper delivery and path resolution. ## What Changed - Bundle the existing issue update helper with the Paperclip skill. Preserve its HTTP checks, echoed-status check and two-attempt limit. - Resolve the command from the installed skill directory. Use a verified PATCH when that path is unavailable, without searching the filesystem. - Keep the repository command as a wrapper that works from any directory. - Test shell execution and exact status/comment payloads through both provider skill-home layouts, including paths with spaces. - Add helper sources and existing verification tests to stock-harness admission. Record an absent historical helper explicitly. Add the missing declaration for the admission fingerprint export. ## Verification - Complete directly affected source suites: 30 tests pass. They cover skill delivery, preserved multiline comments and links, authentication headers, empty responses, mismatched status, transient retries and definitive rejections. - Product E2E typecheck passes. Support suites: 1,835 Vitest tests pass, one is skipped; 128 Node tests pass. - Full local build and workspace typecheck pass. - Full local repository tests are not claimed as passed. Embedded PostgreSQL was unavailable in this worktree during the preceding task; Linux CI will run the repository gates. - The authorized matched Codex/Claude comparison is pending. It uses the existing assigned-skill case and original oracle, one initial attempt per profile and variant. - CI and a fresh Greptile review are pending. Keep this PR in draft until readiness gates complete. ## Risks - Correct path resolution depends on the harness supplying the installed skill path. The instructions use verified PATCH when that path is unavailable. - The helper still requires Bash, curl and jq. Its existing retry and response-verification behavior is unchanged. - Tests use the shared skill-directory symlink mechanism and an HTTP fixture. Real provider completion behavior still requires the bounded live comparison. - This fix does not redesign native completion, legacy recovery or ambiguous transport handling. ## Model Used OpenAI Codex, GPT-6 family. The exact model build and context window are not exposed in this session. Used code editing, shell tools and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.9 |
||
|
|
71cd0a2621 |
fix(skills): honor the current run harness checkout (#15548)
## Thinking Path > - Paperclip manages work for AI agents. > - The runtime claims eligible assigned tasks before it starts an agent. > - The wake tells the agent when the runtime already holds that claim. > - The legacy skill still requires another checkout in every case. > - This PR makes the skill honor the current task and run claim. > - Manual checkout and server ownership checks remain in place for other cases. ## Linked Issues or Issue Description **Where is the issue?** `skills/paperclip/SKILL.md`, in the scoped wake procedure and Step 5. **What's wrong?** The wake can say that the harness already checked out the issue. The skill still tells the agent that it must call checkout. These instructions conflict. **Suggested fix** Skip the second checkout only when the runtime wake explicitly confirms the claim for this issue and run. Retain manual checkout when that statement is absent or the agent selects another task. Refs #14948 for the existing shared prompt reduction. ## What Changed - Honor the explicit runtime claim in the scoped wake procedure and Step 5. - Keep context reads, status writes, deliverable handling and conflict rules. - Add checks for normal and resumed wake text and excluded automatic claims. - Retain successful checkout HTTP activity for legacy stock-task evals. Bind each receipt to the exact company, task, agent and run. Keep this observation separate from the original task grades. ## Verification - Checkout observation calibration: nine tests pass. - Focused skill, wake and database ownership tests: in progress. - Full repository build, typecheck and tests: in progress. - Planned live comparison: the existing assigned-skill document case on legacy Codex and Claude. One attempt per variant and profile. No automatic retries. The baseline and candidate share the observation code and task oracle. - Live results are pending. This draft does not claim behavioral qualification. ## Risks - Agents may misread prompt guidance. The API still enforces ownership; the text grants no new authority. - The exception is specific to the current issue and run. It does not remove ordinary legacy completion writes or authorize another task. - Activity measures successful checkout HTTP calls. Failed attempts require separate run-log inspection. Missing or mismatched observations cannot count as zero calls. - One trial per profile cannot establish general reliability, speed or cost trends. ## Model Used OpenAI Codex, GPT-6 family. The exact model build and context window are not exposed in this session. Used code editing, shell tools and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.8 |
||
|
|
d66acb7ac1 |
feat: automate Slack bot app setup and installation (#15413)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors give each agent a customer-owned bot and task-backed conversations. > - Manual Slack setup requires app creation and copying durable credentials. > - Operators need a shorter setup that an assisting agent can use safely. > - This pull request creates the app through Slack's Manifest API and installs it through OAuth. > - Durable registration state supports recovery without creating another app. > - A four-screen wizard, automatic avatar upload, and OAuth account linking reduce setup work. > - Connector settings and per-turn tool guidance support daily use after installation. ## Linked Issues or Issue Description **Subsystem affected** Native Slack bot setup, company secret storage, chat connector management, and agent tool guidance. **Problem or motivation** New Slack bots require manual app creation and copying a signing secret and bot token. Interrupted setup can create duplicate apps. The setup and management screens contain unnecessary controls. Agents also need guidance for native questions, files, thread replies, and governed Slack actions. **Proposed solution** Use a temporary app-configuration access token to create a customer-owned app. Save durable secrets in the vault. Bind OAuth to the initiating actor, company, endpoint, registration revision, scopes, and configured origins. Preserve manual and existing-app recovery. Link the installing user's account, send a welcome DM, and advance from saved server evidence. Keep request-URL recovery instructions available if automatic connection detection waits. **Alternatives considered** The Slack CLI adds installation requirements. Socket Mode changes transport. A shared Paperclip-owned app changes app ownership. These alternatives are outside this change. **Roadmap alignment** This extends existing chat connectors and secrets capabilities. Related public work: #14037 and #13954 cover Slack MCP prerequisites and user OAuth. No duplicate bot-registration PR was found. ## What Changed - Share one reviewed manifest builder between automatic registration and manual setup. - Add replay-safe migration 0318 and company-bound registration state with vault references and uncertain-creation recovery. - Add registration, installation, callback, and resume APIs with short-lived, single-use OAuth state. - Save installation credentials before downstream checks and preserve bot identity constraints. - Reduce automatic setup to four screens. Keep advanced app details, manual recovery, and existing-app setup. - Upload the agent avatar with the Paperclip dark background. Link the OAuth installer's account and send setup DMs. - Show agent and connector-owner avatars. Simplify settings, access, and conversation screens. - Discover joined Slack channels and enable them by default. Start a task from a bare mention and admit same-thread follow-ups. - Refresh Slack tool guidance each turn. Add native-form, file, approval, and delivery regressions plus manual model probe definitions and sanitized acceptance records. - Update deployment/database docs, OpenAPI, redaction, removal cleanup, production Storybook stories, and provider browser tests. - Merge current master and move the registration migration after its latest migration without rewriting published commits. The completed Slack success view intentionally has a single centered **Done** action and no **Save & exit**, as explicitly requested by the product owner. `DESIGN.md` records this exception; unfinished setup steps retain the aligned wizard footer. ## Verification - Passed after the master merge: repository typecheck, full build, Storybook build, design-token gates, module-boundary gates, and migration generation. - Passed: all 352 focused Slack deterministic tests and all 14 affected provider browser tests. Browser tests use controlled provider fixtures and a separate throwaway instance. - Passed on current head `c5d01e0e2`: the complete GitHub test matrix (general server, chat, all workspaces, serialized server, and Runner), all eight browser shards, typecheck/release registry, build, canary dry run, security checks, and policy gates. There are 52 passing checks and no pending or failing checks. - Greptile completed on the exact current head with 5/5 and no actionable findings or open review threads. - Local repair verification passed 93 focused tests, including same-app reinstall after revocation and rejection of consent started before revocation, the AgentMail browser journey, and repository typecheck. Local build and Storybook build also passed. The redundant local full-suite rerun was stopped after the complete current-head CI matrix passed. - Real Slack setup and agent replies were exercised in the authorized isolated test drive during the setup iteration. - The ten additional model probes were attempted with legacy `codex_local`, `gpt-5.6-sol`: five passed, two failed, and three were partly verified. Native runtime is not qualified. See `server/src/services/connectors/slack/evals/2026-10-08-acceptance.md` for evidence and limits. - Passing model probes cover native forms, downloaded file bytes, bare mentions with thread replies, explicit posts/reactions, and saved approval denial. - The controlled uncertain-write probe found wrong delivery-check IDs. The canvas fallback attempt used an invented tool name. Search pagination/native search, a private-source denied-tool receipt, and distinct board/webhook origins remain unqualified. Reviewer path: enable Chat connectors, start Slack chat setup, select an agent, enter an app-configuration access token, and approve Slack installation. Send a message to the bot and confirm that setup advances to success. Inspect settings and allowed channels. See `doc/connections/SLACK-AUTOMATIC-SETUP.md` for deployment and recovery. ## Risks - Slack app creation has no provider idempotency guarantee. A timeout after dispatch stays uncertain until the operator checks Slack. - OAuth needs a stable public HTTPS board origin. Webhook ingress may use a separate configured HTTPS origin. Workspace policy can delay installation. - Migration 0318 can replay safely on instances that applied the earlier development migration. - OAuth installation now links the installer to the initiating Paperclip user. Identity checks and company access rules still apply. - Joined channels now enable bot responses by default. Linked-user authorization and per-action approval rules still apply. - Model behavior has the documented delivery-check and canvas fallback failures. A passing CI run does not establish that every model probe passed. - Removing the connection does not delete the customer's Slack app. No new first-party telemetry is added. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser verification. The runtime does not expose a more specific authoring model ID or context-window size. The live bot probes used OpenAI `gpt-5.6-sol` through `codex_local` in legacy mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d0f69670db |
fix(runner): recover saved execution prompts after upgrades (#15518)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs save an immutable execution context for restart recovery. > - The context includes the prompt text, revision, and content hashes. > - The parser required that saved prompt to match the current release. > - A server upgrade could reject a valid saved run before provider recovery. > - This pull request validates and preserves the saved prompt snapshot. > - Routine prompt changes no longer need a catalog of past strings. ## Linked Issues or Issue Description **What happened?** A hot restart selected a dead native runner for same-run recovery. Reading its saved v5 execution input failed with `input.runtimeContext.prompt must match the fixed Paperclip prompt revision`. The new controller accepted only v6. **Expected behavior** Recovery uses the saved prompt and validates its content hashes. It preserves the same run and provider session without starting a duplicate turn. **Steps to reproduce** 1. Start a native run and save its execution input and provider checkpoint. 2. Change the fixed execution prompt in the server release. 3. Stop the runner and recover the saved run with the new controller. 4. Observe that the old parser rejects the saved prompt before provider recovery. **Paperclip version or commit** The v5-to-v6 prompt change was introduced in #15446. The defect also reproduces on current master before this fix. **Deployment mode** Source-built server with the native runner. Related work: #15446 added task-monitor guidance. The held prompt-size experiment in #15489 changes prompt wording but does not add recovery compatibility. ## What Changed - Read the prompt text and revision from the saved execution snapshot. - Treat the revision as non-empty metadata and preserve the exact saved bytes. - Validate the prompt SHA-256 and the aggregate context digest. - Keep fresh-run builders on the current prompt constants. - Test arbitrary saved prompts, malformed fields, altered text, stale hashes, and aggregate drift. - Test recovery parsing for Codex input versions v3-v5 and OpenCode, ACPX Pi, and Dot v6 inputs. - Extend the real-process restart suite with both the incident's v5 wire fixture and a prompt unknown to this release. - Run the restart recovery suite in the existing Rust-equipped PR lane, where its runner and fake-provider binaries are built. Verify complete, non-overlapping test coverage for PR, release, and local callers. - Document recovery from saved snapshots without a historical prompt catalog. ## Verification - Red: the new contract regressions fail against the catalog-based parser with the original prompt-validation error. - Green: 53 focused contract and materialization tests pass. - Red: the real-process unknown-prompt regression fails with master's original parser at the saved-input recovery read after process loss. - Green: all 15 real-process restart tests pass locally on the final branch. The saved-prompt cases keep the run and provider session, replace the PID, and record one `turn/start`. - Local repository `pnpm -r typecheck` and `pnpm build` passed after rebase on `89f09dad723766e5351953f0731b9aa5daada28d`. The 53 focused tests also passed on that head. - Red: the new test-roster checks fail against the old CI placement. - Green: all 26 test-scheduling checks pass after moving the restart suite. - CI ran all 15 restart recovery tests with no skips on final head `6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`. [Runner test job](https://github.com/paperclipai/paperclip/actions/runs/37764782334/job/113271464966). - Greptile scored 5/5 on that exact head with no actionable findings. - The complete CI matrix passed on final head `6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`: general/workspace tests, serialized server suites, both runner Vitest lanes, Rust and static checks, all browser shards, typecheck, build, and the canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37764782334). - There are no unresolved review threads or merge conflicts. - Reproduce focused tests with `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/contracts/runtime-context.test.ts src/contracts/native-execution.test.ts src/drivers/runtime-context-materializer.test.ts`. - Reproduce restart tests with `pnpm --filter @paperclipai/paperclip-runner build:rust` followed by `pnpm exec vitest run server/src/services/native-runtime/native-runner-restart-recovery.integration.test.ts`. They use temporary PostgreSQL, real runner processes, and a fake Codex provider. They do not use paid inference. ## Risks - The parser now accepts internally consistent saved prompt text that is absent from the current source. Inputs must come from trusted server persistence. Content hashes verify consistency; they do not authenticate authorship. - Existing execution-schema, ownership, checkpoint, provider, permission, and session-compatibility checks still apply. - This change validates the saved base prompt. It does not make all additional code-generated instruction strings versioned. - The process-level recovery proof uses Codex. Other provider coverage verifies the shared input parser and retained provider configuration. ## Model Used OpenAI Codex, GPT-6 family, with repository inspection, code editing, and test tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.7 |
||
|
|
bf9dbd18a8 |
fix(runner): recover incompatible Codex models before launch (#15519)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner prepares and saves each task's provider configuration before launch. > - A sandbox image can have a supported Codex CLI that is too old for the selected model. > - The current version check stops that task even when the image can run a similar model. > - This pull request selects a compatible model before it saves a fresh execution. > - The task continues with a visible warning, and recovery uses the saved effective model. ## Linked Issues or Issue Description - Refs #15053. This fixes the same CLI and model mismatch for fresh remote Runner tasks. It does not change subscription onboarding. - Refs #14721. This adds a bounded startup choice for Codex CLI compatibility. It does not add user-defined fallback chains or quota failover. - Related PR: #15399 added the model-specific CLI checks that this change preserves at launch. ## What Changed - Check the preinstalled remote Codex CLI before saving a fresh execution that uses a model with a verified CLI minimum. - Select the closest compatible older model in the same class, then the stable Runner default. Consider each candidate once. - Save the effective model before checkpoint selection. Keep the requested agent and task settings unchanged. - Add a task warning and a local `runner.model_fallback` run-log event with both models and the CLI version. - Share executable discovery with the launch verifier. Preserve explicit artifact and install settings, saved executions, and existing artifact checks. - Add regression tests and document the selection and warning behavior. ## Verification - Red: the three new provider-configuration regression cases failed before the fix. They kept the incompatible requested model. - Red: both rejected-probe cases failed before the review fix. They now defer to launch verification. - Green: targeted provider configuration, remote preflight, task warning, and launch-verifier tests pass (680 tests). - `pnpm -r typecheck` passes. - [Full CI](https://github.com/paperclipai/paperclip/actions/runs/37713759744) passes on `eb273fc32661526d90a864b42c73a7da61dd62da`: 47 successful jobs, including all general and serialized Vitest groups, browser shards, Runner verification, typecheck, build, and the canary dry run. - Local extended verification: three serialized shards pass (115 suites). HTTP route tests hit intermittent 15-second timeouts. The authorization suite passes on rerun (130 tests). The document suite passes on both master and the PR head in the same isolated setup (6 tests). The local general run was stopped after the equivalent full CI matrix passed. - `pnpm build` passes. - Server typecheck and build pass again after the probe-error review fix. - `git diff --check` and `pnpm check:module-boundaries` pass. - Example: request `gpt-6.1-sol` on Codex 0.158.0 to use `gpt-6-sol`; on 0.156.0, use `gpt-5.6-sol`. The warning names the requested model, effective model, and CLI version. - This PR has not been deployed to staging. ## Risks - A fallback can have different capabilities. The task warning makes the substitution visible. Later runs can use the requested model after the image CLI is updated. - Preparation adds two remote commands for models with a verified CLI minimum. Failed or invalid version probes retain the existing launch checks. - This only handles known CLI and model mismatches before a fresh provider launch. Authentication, capacity, and artifact failures keep their existing behavior. - No database migration is required. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. This session does not expose the exact deployment model ID, context window size, or reasoning setting. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.6 |
||
|
|
b5342febe5 |
fix(runner): require current-turn completion after connection continuations (#15514)
## Thinking Path > - Paperclip manages agents, tasks, permissions, and execution budgets. > - Native tasks can resume after a connection decision in the same provider conversation. > - Each turn still needs an accepted completion report. > - Compact continuation messages did not explain that reports from earlier turns cannot finish the new turn. > - The connection evaluator could also grade before the final reply was stored or reject valid unavailable-access wording. > - This PR clarifies the current-turn report requirement and fixes those observation boundaries. > - The original connection instructions and strict native completion gate stay in place. ## Linked Issues or Issue Description Refs #15489. The reduction remains draft while this separate repair is qualified. Refs #15471 for the earlier connection continuation work. ## What Changed - Add a current-turn completion reminder to compact continuation inputs. - Keep final prose insufficient for completion. Preserve permissions and retry policy. - Wait for the final successful task run's saved, attributed decline reply within the existing deadline. - Use one bounded explanation matcher for both decline checks. - Wait for a recorded tool-action rejection to dispatch its bound continuation, with strict company, task, agent and source-run checks. - Retain the exact grading input before later API refreshes. - Add failure and delay regressions and update the Runner and evaluator docs. ## Verification The fresh comparison has **15/15 original passes on each variant**: 15 unchanged pass pairs, zero new failures, and no pending pair. There are 30 case attempts and **65 actual agent runs** (baseline 33; candidate 32). All runs succeeded. All 30 cleanups passed. No model attempt was retried. | Profile | Baseline | Candidate | | --- | --- | --- | | Native Codex `gpt-5.6-sol` | 5/5 | 5/5 | | ACPX Claude `claude-sonnet-5` | 5/5 | 5/5 | | OpenCode `openrouter/deepseek/deepseek-v4-flash-0731` | 5/5 | 5/5 | Each profile covers service approval, service decline, connection decline, provider decline, and selection of the second provider. The saved replies, approved briefings, decisions, fixture observations, final task states, and native completion records were inspected. All 18 saved decline-grade snapshots match their original captured inputs and checks. Result, API snapshot, and final ledger run sets agree. - [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37711658378) · [public candidate report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711658378-1/index.html) - [Baseline campaign](https://github.com/paperclipai/paperclip/actions/runs/37711675579) · [public baseline report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711675579-1/index.html) - Candidate source: `79905343bba280d462765faad19a26e7f179259e`. Baseline source: `7c5e120158f1385a1fc5f20be41f66c58a605534`. Both use master context `fc6304dfe5f2e446e09bd052a7b45f51e930f250`. The trusted workflow source is separately frozen at `dd777f4b7343305c4e6f44c422f44a1d78e12e4f`. - Both variants use the same evaluator, fixtures, models, permissions, 720-second cell deadline, and 1,000-cent company and agent hard stops. The only production difference is the compact continuation reminder. - Suite hash: `aee30b74b4d38ada08777798db0932fbc64e368bb427d29caa8cfd86c7f59747`. Definition hash: `ea9e17f54fe0af1acbb2337ddfab8a7e92db8488bd0f59510a63095ca0229b60`. - Provider-free transport capture: startup/resume input stays at 55,726 bytes. Compact continuation input grows from 53,334 to 53,535 bytes. The 42-tool catalog stays unchanged. These are Paperclip input bytes, not complete vendor prompt tokens. - Focused evaluator tests: 31 pass. Native contract, transport delivery, and session tests: 194 pass. Evaluator support: 1,819 TypeScript tests and 128 Node tests pass, with one intentional skip. - Full build, workspace typecheck, and evaluator typecheck pass. Current-head CI passes all required gates. The current rollup has 51 successful check runs, two intentional Storybook skips, and a successful Snyk status. Review is 5/5 with zero unresolved threads. - CI attempt 1 had one initial runtime-fixture health timeout. The exact test and its full 164-test file pass locally. One CI shard retry passed. The original CI failure, its dependent verify failure, and the retry remain visible in [CI history](https://github.com/paperclipai/paperclip/actions/runs/37711235060). - The broad local `pnpm test:run` attempt was interrupted after about 49 minutes (exit 130). It recorded one failure in the unchanged Zep memory-connector disabled-setting test. That test and the full 388-test tool-access file pass in separate local checks; current-head CI also passes. The local cause is not established, and this broad local attempt is **not** claimed as passing. Three earlier local failures also pass in their isolated checks; their original logs remain retained. - The first two setup admissions were cancelled before provider jobs to include the review correction. They made no provider calls. The completed campaigns above are the first and only model attempts for these corrected variants. Cost evidence stays separate from behavior. Original result summaries report only OpenCode amounts: baseline $0.039192096 and candidate $0.063221620. Final run ledgers also retain estimates for Claude (baseline $1.232439000; candidate $1.333611200) and Codex (baseline $2.340324400; candidate $2.129335600). These estimates do not replace the original summaries. Local and GitHub runtime are unmetered here. Invoices are unknown. This is not a cheaper or faster claim. ## Risks - A single matched trial cannot prove general equivalence or causation. The reminder is an instruction change, not a new completion enforcement rule. - Candidate OpenCode service-decline finished within one continuous run; its baseline used two. That pair passed the task outcome, but it does not qualify the reminder on a resumed decline turn. No extra paid run was used to replace it. - The text matcher is bounded evidence of an explanation. It does not prove reasoning or consumption of feedback. Bounded stdout excerpts do not prove that every extra attempted tool call is absent. - Missing saved replies or continuations still fail at the original deadline. Failed native completion remains a failure even when final prose is correct. - The unresolved local-suite discrepancy above remains a validation limit. Full remote CI and both focused local reproductions pass. - The original connection instructions stay in place. These results do not qualify the reduction in #15489. Its original 11/15 versus 12/15 grades and two new failing pairs remain unchanged. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, test inspection and eval analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the focused local tests listed above and they pass; the interrupted broad local run is disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.5 |
||
|
|
5717523b9e |
fix: retain the original repository for reused task workspaces (#15528)
Apply the reviewed change for fix: retain the original repository for reused task workspaces. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.4 nightly/v2026.1008.0-nightly.0 |
||
|
|
e4b39da6f6 |
Isolate ACPX admission deadline tests from lease ports (#15527)
Apply the reviewed change for Isolate ACPX admission deadline tests from lease ports. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1d19f9b562 |
Distinguish failed Git inspection from missing worktree registration (#15525)
Apply the reviewed change for Distinguish failed Git inspection from missing worktree registration. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ed0c6958f8 |
Explain refused local Hermes gateway connections (#15520)
Apply the reviewed change for Explain refused local Hermes gateway connections. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6b171c9616 |
Keep proven local workspace configuration conflicts out of Sentry (#15521)
Apply the reviewed change for Keep proven local workspace configuration conflicts out of Sentry. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f4e6362ccc |
chore(lockfile): refresh pnpm-lock.yaml (#15512)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.1008.0-canary.2 |
||
|
|
dd777f4b73 |
feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path
> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.
## Linked Issues or Issue Description
**Subsystem affected**
Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.
**Problem or motivation**
An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.
**Proposed solution**
Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.
**Roadmap alignment**
This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.
## What Changed
- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.
## Verification
- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head
canary/v2026.1008.0-canary.1
|
||
|
|
fc6304dfe5 |
feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path > - Paperclip manages AI agents, tasks, permissions, and execution budgets. > - Paperclip Runner gives each provider the same admitted task and tool authority. > - OpenAI Dot runs outside the local process tree and needs asynchronous work delivery. > - The merged MCP gateway supplies OAuth consent and signed event delivery. > - A personal assistant grant cannot safely stand in for an assigned agent. > - This pull request adds a separate Dot agent connection and a durable Rust Runner bridge. > - The operator can assign work to Dot and inspect its accepted work, tool receipts, and result. ## Linked Issues or Issue Description **Agent or provider** OpenAI Dot, as an experimental provider of the existing Paperclip Runner adapter. **Why this adapter is useful** An operator can assign normal Paperclip tasks to an existing Dot. Dot can read its mailbox, request work on an assigned task, use admitted task tools, and submit a result. Paperclip keeps company scope, checkout, approvals, known budget limits, and activity attribution. **How the agent is invoked** A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the assignment. The Rust Runner owns the durable turn and operation receipts. The first release supports self-hosted instances with a local Runner controller. **Additional context** This extends the merged public MCP gateway from #14846 and the assistant invitation and device-consent work from #14933. This also integrates the merged assistant tool and configuration expansion in #15380. Dot retains its dedicated agent resource and cannot receive personal configuration permission. The public assistant connection remains a personal connection. ## What Changed - Add a durable Rust Dot provider and its TypeScript Runner driver. - Add closed PRP v3 external-provider operations and native execution input v6. - Add company-scoped pairing, mailbox, assignment, and operation records. - Reuse merged browser/device consent, client metadata verification, webhook admissions, refresh, secret rotation, and warm-standby gates. - Keep Dot scopes, issuer, grants, event workers, and tool access separate from personal assistant access. - Add Dot configuration, pairing, readiness, and consent UI. Keep agent grants out of the personal Connections entry. - Regenerate the Dot-only migration after master. Preserve published gateway migrations. Make the new migration safe to reapply. - Document setup, recovery, accounting limits, evidence, and remaining account qualification. - Reverify reconnect callbacks and wake outstanding work with a fresh mailbox reference; preserve the existing assignment and operation receipts. - Clean up Dot bindings and waiting runs on OAuth revoke and refresh-token replay. Old grants cannot revoke replacement bindings. - Restore the pairing reference when an unsaved agent form is reopened; document board-only pairing routes in OpenAPI. - Accept a clean Rust exit after the acknowledged shutdown receipt. Unexpected exits still require recovery. - Clear the cached binding after a successful revoke so a failed connection refresh cannot restore it. - Add production-component Storybook states and screenshots for pairing and connection review. All preview account data is synthetic. - Persist normalized completion, serialize Dot turns and durable work admission, and poll subscription readiness. - Serialize mailbox writes and cursor reads; retain paused fence acknowledgement without task authority. - Authorize admitted review runs without changing the worker assignee. Include the fenced assignment ID in production stop notices. ## Verification - This PR integrates master `4a8178e9c`. Dot migration `0317_messy_famine.sql` follows the published history and is safe to reapply. The merge preserves the reserved migration connection, batch-commit handling, private task checks, task monitors, and native accounting. - Local workspace typecheck, full build, and UI token gates pass. The server typecheck passes after the review fixes. Database and native executor regressions pass. - All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot driver tests pass. The tests cover native document writing and finalization, durable replay, queue admission, mailbox ordering, admitted reviews, stale authority, production stop references, and paused acknowledgements. - Current head `d0e7e0626` passes all 57 checks: 53 pass and four are intentionally skipped. This includes full typecheck, build, tests, Rust Runner verification, browser E2E, release verification, and Canary Dry Run. Greptile rates this exact head 5/5. All review threads are resolved. - The full local root test run is slower than the sharded CI run and has not completed. The full CI test gates pass on the current commit. Focused local regressions pass. - Real-account pairing and event delivery on this base commit remain unqualified. Live account and setup proof are recorded in the follow-up #15414. The following screenshots use synthetic preview data. They show the production pairing component and do not qualify a real account or the full agent setup journey.   ## Risks - This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP and Paperclip Runner. The separate experimental-settings follow-up in #15414 replaces this environment flag with saved operator settings. - Dot does not expose provider token usage or cost. The operator must acknowledge external billing. Known Paperclip budget gates still apply. - Cancellation fences Paperclip authority. It does not confirm that Dot stopped all external activity. - Assigned skill files and third-party MCP bindings are unsupported and reject admission. There is no mounted workspace, model selector, or provider thread identifier. - Hosted agent-broker and remote controller deployments are not qualified. - The new migration follows the merged master history. Existing prototype databases still need the normal master migration history before this Dot-only migration. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window size are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, and test inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.0 |
||
|
|
4a8178e9cf |
fix(ui): stabilize composer model selection and refine effort slider (#15444)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task composers let users select an agent, a model, and an effort level. > - The model label changes width when users move the effort slider. > - This movement makes the control harder to use. > - This pull request keeps the label fixed while the picker is open. > - A larger slider gives clearer feedback during selection. ## Linked Issues or Issue Description **What existing behavior does this improve?** The shared model and effort picker in task composers. **Current behavior** The trigger shows each effort change while the picker is open. Different label lengths move the control. The slider is small. **Proposed behavior** The trigger shows “Select model” while any picker view is open. It shows the selected model and effort after the picker closes. The slider has a thicker track and a larger thumb with hover and press feedback. **Reason and benefit** Users can change effort without moving the trigger. The larger thumb gives clearer pointer feedback. ## What Changed - Keep the trigger text fixed while the picker is open. - Add slider size and feedback tokens. - Remove the gray input outline and thumb border. Show keyboard focus on the thumb. - Keep a thumb focus outline when Windows high contrast suppresses shadows. - Keep native keyboard and pointer controls. Respect reduced motion. - Test open and closed labels on desktop and mobile. ## Verification - Picker tests: 27 pass. Composer settings and new-task tests: 86 pass. - UI typecheck and UI build passed before the final CSS-only high-contrast fix. - Token gates pass. - Full typecheck and build stop at installed Runner API type mismatches outside these files. - Revised Chromium checks pass for desktop pointer drag, no input outline during drag, stable open label, keyboard input, and mobile rendering. The earlier rendered-image check confirmed reduced-motion behavior. - The earlier full local suite did not finish. The approved visual revision passed 113 focused tests. The final high-contrast fix passed 27 picker tests and token gates. - Chromium high-contrast and reduced-motion verification shows a visible thumb focus marker and working keyboard input. - Open the model picker. Change effort. Confirm “Select model” stays visible. Close the picker. Confirm the selected model and effort appear. - All CI gates pass on final head `7a64104e14251a354b06b7916d09d44692557c4f`, including full typecheck, tests, build, and browser E2E. Greptile: 5/5 with no findings. ## Risks - Low risk: shared UI only. Native range controls stay in place. - Browser thumb rendering can differ between engines. ## Model Used - OpenAI GPT-6 Codex. Code editing, tool use, and test execution. The runtime does not expose an exact model variant or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.16 |
||
|
|
378e6d95e1 |
chore(lockfile): refresh pnpm-lock.yaml (#15488)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
a7a244ab33 |
feat: add company decision models with permission and cost controls (#15473)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Optional product features need small, typed model decisions. > - Each company needs to choose which shared API connection pays for those decisions. > - Calls must retain user and task permissions, budget limits, and cost attribution. > - This pull request adds a managed decision service, setup UI, and request history. > - Features can check availability cheaply and keep their existing behavior when decisions are unavailable. ## Linked Issues or Issue Description **Subsystem affected** Server services, shared contracts, database accounting, Company Settings, and Costs. **Problem or motivation** Paperclip has no common decision-model service. Adding provider calls within each feature would duplicate credential access, permission checks, and billing rules. **Proposed solution** Let a connection manager configure one company decision model. Support OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test and metadata-only history. Default company-sponsored background decisions to on during setup, and preserve a saved off setting. **Alternatives considered** Per-feature credentials would duplicate existing connection management. Personal overrides and provider fallback chains add permission and billing complexity; they remain deferred. **Roadmap alignment** Reviewed ROADMAP.md and searched open PRs. This extends existing connection access and budget accounting. Product features that call the service remain outside this change. No matching decision-model service PR was found. ## What Changed - Add company settings, an internal `decisionModelService`, local availability checks, and trusted human, agent/run, and system contexts. - Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and ordered-score decisions for both providers. Bound requests and time; disable paid retries. - Add durable invocation metadata and agentless decision ledger charges. Preserve fractional cents, pricing evidence, dispatch identity, and unresolved billing holds. - Share the company accounting lock and apply company, agent, and project budgets. Settle charges once, retain unknown holds, and recover interrupted calls without resubmission. - Reuse connection setup and management UI. Add a Decisions view under Costs, production-component Storybook coverage, database migration, and service documentation. ## Verification - Passed 166 current-code tests covering the decision service/provider, setup component, Costs, OpenAPI, and every failure from the earlier broad run. Coverage includes native SDK wire formats, refusals, billed malformed responses, permission and secret-rotation races, identity changes, concurrent budget admission, unresolved holds, agent/task deletion, and stale setup feedback. - Passed 170 existing connection, cost, budget, heartbeat-accounting, and profile regression tests. - Passed repository typecheck, production build, and design token gates after integrating master. Verified the generated migration on a fresh test database and upgraded the populated preview database from the branch's earlier migration without losing settings or usage. - Ran the required full `pnpm test:run`: its general phase completed with 16,354 passed and 10 failures across five files while this branch was still being updated. Every reported failure passes in the current-code rerun; the serialized phase did not run after that failure. The full GitHub CI suite passed on `b1b856a88`: general and serialized tests, browser shards, runner checks, typecheck, production build, packaging/canary, and policy gates. [CI evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360). - Passed the full-shell Storybook setup-to-history interaction test again after integrating master. Greptile rates the final revision 5/5 with zero unresolved threads. - Walked through the running app: empty setup, add each provider, save, reload, run all three sample questions, inspect fractional charges in history, switch provider while sponsorship is off, disable, and reconnect. Checked mobile settings. These tests used the actual UI, vault, server, SDKs, and database with simulated upstream responses. - Live paid setup tests remain unverified: this environment has no authorized OpenAI/OpenRouter credentials available. No mocked test is presented as live provider evidence. Reviewer journey: Company Settings → General → Decision model. Add/select a shared API connection, save, run the billed sample, open View usage, then disable decisions and verify Run test is disabled after reload. ## Risks - The SDK decision interface is experimental. Pinned versions and wire-format tests limit upgrade drift. - The migration allows agentless service charges and reservations. Existing agent cost-reporting APIs still require an agent, and decision receipts stay separate from run reconciliation. - Timeouts can have unknown provider charges. Holds remain until an audited accounting correction resolves them. - OpenAI prices use a versioned Decisions rate snapshot; OpenRouter costs use provider receipts. Unknown pricing is retained as unknown. - A configured company authorizes background spending by default. Setup explains this, and managers can turn it off. - Live provider account/model availability still needs the two credentialed acceptance checks. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser tools. The exact serving revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.15 |
||
|
|
f1c44b7b56 |
Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs. Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
cd40ebb95b |
Preserve ACP bridge terminal evidence before disconnect (#15478)
Drain terminal frames through backpressure and record bounded terminal evidence without waiting for log persistence. Guard later events and input failures against writes after socket end. Validated with 217 focused tests, full typecheck/build, independent review and green CI with Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.14 |
||
|
|
50b1f95e79 |
fix(tool-gateway): keep MCP connection healthy on oversized/malformed responses (#15462)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents call remote MCP tools through the Paperclip tool gateway > - The gateway keeps a health status for each remote MCP connection, and it hides the tools of a connection that has `healthStatus = "error"` > - One oversized, malformed, or invalid-JSON `tools/call` reply set the whole connection to `error` > - Health goes back to `ok` only after a successful call, but the hidden tools prevent that call, so the connection stayed locked until a person reconnected it > - This pull request makes these reply errors fail only the one call, and it starts a new MCP session for the caller on the next call > - The benefit is that one large or bad reply no longer disconnects a working connection for every agent ## Linked Issues or Issue Description No public issue exists. Related open PRs fix other health downgrades in the same function. They do not overlap with this change: - Refs #11910 (timeouts and JSON-RPC errors) - Refs #15325 (transport errors) **What happened?** An agent called a remote MCP tool that returned more than `MAX_REMOTE_MCP_RESPONSE_BYTES` (1 MB). The gateway returned 502 `mcp_remote_response_too_large`. It also set the connection to `healthStatus = "error"`. Then `connectedMcpConnectionFilter` hid all tools of the connection (except `per_user` connections), and `connections_search` showed the connection as `needs_user_action`. The `malformed_response` and `invalid_json` errors from `readMcpHttpResponse` did the same thing. `readMcpHttpResponse` can also cancel the reply stream before the end. The remote server can then close its MCP session. The gateway kept the cached `mcp-session-id` for up to 30 minutes and cleared it only after a 404. The next calls then failed with "Remote MCP session expired". **Expected behavior** A reply that is too large or not correct fails only that call. The connection stays healthy, and its tools stay visible. The next call uses a new MCP session. **Steps to reproduce** 1. Add a remote MCP connection with `mcpSessionRequired: true`. 2. Call a tool that returns a reply larger than 1 MB. 3. Look at the connection: `healthStatus` is `error`. 4. Start a new gateway session: the tools of the connection are not in the list. **Paperclip version or commit** `master` at `99a9de9940bf5974352d9dbfbb2f21e62e89689f` **Deployment mode** All modes. The fault is in the server tool gateway. ## What Changed - `server/src/services/tool-gateway.ts`: `too_large`, `malformed_response`, and `invalid_json` from `readMcpHttpResponse` now fail only the call. The caller gets the same 502 reason code as before. The gateway does not call `markRemoteConnectionHealth(…, "error")` for these errors. The `invalid_json` branch after `JSON.parse(body)` also does not change health now, so all `invalid_json` paths are the same. - The `mcp_remote_response_too_large` message now tells the caller to request a smaller result, for example a narrower query or a smaller page size. The error details now include `maxBytes`. - `server/src/services/mcp-http.ts`: new `forgetMcpHttpSession()`. It removes only the cached session for one scope and one credential set, and only while that entry still holds the session ID that failed. It does not touch other agents, sessions that are initializing, or a newer session that replaced the failed one. - The gateway calls `forgetMcpHttpSession()` after these reply errors, so the next call from that caller initializes a new session. The existing 404 "session expired" path now uses the same function. Before, both paths cleared all sessions and all pending initializations on the connection. That made a concurrent initialization by another agent fail with "MCP connection changed while initializing", which the gateway reported as a fetch failure and marked as a connection `error`. - `forgetMcpHttpSessions(connectionId)` is not changed. Disconnect and revocation still clear the full connection. Why `malformed_response` and `invalid_json` are also per-call errors: - Each error is about one reply body. The server was reachable, accepted the credentials, and sent HTTP 2xx. A reply can be too large or bad because of the tool and its arguments. That is not a fault of the connection. - The other malformed-reply checks in the same function (payload is not an object, or has no `result`) already throw `remote_mcp_malformed_response` and do not change health. Only the reader-level errors changed health. - Health recovers only after a successful call. An `error` status for one bad reply therefore locks the connection until a person reconnects it. - Other failures (HTTP errors, fetch failures, timeouts, JSON-RPC errors) still change health. This PR does not change them. #11910 and #15325 address some of them. ## Verification - `server/src/__tests__/tool-gateway.test.ts`: the recovery test now runs for three replies: oversized, invalid JSON, and a reply without the requested message ID. Each run uses a fake HTTP MCP server with `mcpSessionRequired: true` that gives `session-N` for each `initialize`. Each run checks that: - the call fails with the correct 502 reason code (and, for the oversized reply, the smaller-page hint and `maxBytes: 1000000`) - `healthStatus` stays `ok` - a new gateway session still lists the tool and can call it - the second `tools/call` uses `session-2`, not `session-1` - New test: agent A gets an oversized reply while agent B initializes a session on the same connection. Agent B's call completes, health stays `ok`, and the two calls use `session-1` and `session-2`. With the old connection-wide reset, this test fails: agent B gets 502 `mcp_remote_fetch_failed`. - `server/src/__tests__/remote-mcp-protocol.test.ts`: new unit test. `forgetMcpHttpSession()` keeps the session of a different identity, and a late failure from an old session does not remove the newer session. - Each regression check fails without its fix. Without the health change, the test fails on `healthStatus: 'error'`. Without the session reset, it fails with `['session-1', 'session-1']`. ``` cd server npx vitest run src/__tests__/tool-gateway.test.ts src/__tests__/tool-gateway-service.test.ts src/__tests__/remote-mcp-protocol.test.ts Test Files 3 passed (3) Tests 132 passed (132) ``` - I did not run the full server `tsc --noEmit` locally because the sandbox does not have sufficient memory. The CI typecheck covers it. ## Risks - Low risk. The change affects only the error path of remote MCP `tools/call`. - A remote server that always sends bad replies now keeps `healthStatus = "ok"`. Each call still fails with a clear 502 reason code and an audit record, so the failure stays visible. Only the connection-wide hiding of tools stops. - After one of these errors, the next call from the same caller sends one more `initialize` request. This adds one round trip. - A 404 "session expired" now clears only the session of the caller that got the 404. Before, it cleared the cached sessions of all agents on the connection. If the remote server restarts, each agent now gets its own "session expired" error one time and then initializes again. A 404 is about one session, and the old connection-wide clear also cancelled other agents' initializations. - #11910 and #15325 change the same catch block. The PR that merges last can have a small merge conflict. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, 1M-token context window. - Run as an agent in Claude Code with tool use (shell, file edit, and local test runs). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing documentation changes are necessary) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge This PR replaces #15418. It keeps the same commits on a branch name without an internal ticket ID, and adds a fix for the Greptile review on #15418. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
fd8c6b920a |
fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native tasks can pause while a human decides whether to allow a connection action. > - The next turn needs both the saved decision and complete accounting for the previous turn. > - Cancellation could discard final usage, and complete direct Claude API receipts could remain unpriced. > - A stale blocked or review report could also request approval again after the original card was declined. > - This pull request retains shutdown accounting and rejects approval waits bound to an already resolved action. > - The benefit is reliable continuation with the existing budget and approval controls. ## Linked Issues or Issue Description Refs #15420. Related: #15312 addresses requester ownership during dispatch. This change addresses receipt capture and final-response validation. ## What Changed - Retain usage events after native cancellation. Continue to reject late provider messages and work. - Drain same-turn accounting and terminal events for at most fifteen seconds after a durable governed wait. Keep incomplete accounting blocked. - Estimate complete, unpriced, direct Anthropic API receipts for the exact `claude-sonnet-5` model. Record the rate version and assumptions. Use the one-hour cache-write rate when the receipt lacks cache TTL. - Bind stale approval reports to exact interaction, action-request, or invocation IDs in the same company, task, agent, and run. Cover blocked, review, and response-wake reports. Keep independent reviews valid. - Fence checkpoint and result writes after a controller detaches for restart, including operations waiting for a database lock. Reject stale successful returns before certifying accounting. - Allow bounded subscription teardown only after retaining an actual provider terminal. - Journal the exact governed-wait trigger and disposition before provider interruption. Recover that wait independently of a later saved answer, replay retained accounting, and reject mismatched or unproven terminal evidence. - Give settling governed turns a bounded window before shutdown detaches their controller. - Keep fuzzy external app matches alongside installed capability matches instead of forcing an unrelated provider question for a generic query. - Return up to twenty exact active catalog tool names after an invalid request, after eligibility checks; still reject the request without granting access or creating an approval. - Require retained provider terminal proof before settling a governed wait, including when complete usage arrives before stream closure/error/timeout. Retain harmless numbered cancellation events so restart replay stays contiguous. - Isolate accounting-test OpenCode config from the host plugin directory. - Add regression coverage and document the accounting, restart and connection-search behavior. ## Verification Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`; later review fixes are verified separately and do not relabel those runs. - Repository typecheck and build pass. - Current runner runtime and cancellation suites: 191 tests pass. Six new regressions cover stream end/error/timeout without provider-stop proof and contiguous cancellation acknowledgement/request replay; all six failed before the fix. The existing bounded cleanup case now explicitly supplies terminal proof. All 31 adapter accounting tests pass with isolated fixture config. - New checkpoint-rebinding and approval-criterion suites: 65 tests pass; runner HTTP integration: 29 tests pass. Unchanged executor/control-plane suites: 640 tests pass; database-backed connection suites: 68 tests pass. - Current retained evidence verification covers 222 file hashes across all fifteen original result artifacts. The three-case [published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html) and all eight screenshot hashes verify. The [three-case recovery report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html) and all six screenshots also verify; the nine-result campaign did not publish. - Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node tests pass. Eval typecheck and catalog discovery pass. - Full local repository unit run was interrupted before the follow-up edits after three tool-access failures and one runner HTTP failure. Those failures pass in isolation; the 09a full run was interrupted after one rapid Slack callback-ordering failure and seven skill-service failures. All eight pass both isolated and with full-runner environment settings, and all 74 skill-service tests pass together; the subsequent full run reported two 15-second OpenCode accounting timeouts and was stopped with exit 130 to apply review fixes. The timeouts reproduce while copying this host’s 61 MB OpenCode config. All 31 tests pass after isolating config inside each fixture without increasing timeouts or changing assertions. A complete local full-suite pass is not claimed. Repository-wide CI also passes on the final review-fix head in [run 37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952). Earlier database-skipped diagnostics and the older ENFILE run are retained and are not full-suite passing evidence. - Three-case live campaign [37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525) passes all three original grades on frozen source 09a: 55/55 checks, eight succeeded run records, complete accounting receipts, matching checkpoint identities and no pending approvals. Campaign [37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353) adds nine original passes (173/173 checks, eighteen succeeded records) on the identical source. Its other three jobs failed before runner assignment or any step while GitHub could not load the paid environment; those original infrastructure failures are retained. Campaign [37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356) completes only those unstarted cells: all three original grades pass (57/57 checks, six succeeded run records). All fifteen exact cases now pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new overall failures and zero pending pairs. Total current evidence: 285/285 checks and thirty-two succeeded run records, complete accounting receipts, matching checkpoint identities, no pending approvals or retry records. Earlier campaigns retain forty-seven additional run records and two known same-run recovery attempts; actual provider-call counts and invoices remain unknown. The nine-result campaign skipped publication and its public URL returns 403; original artifacts remain retained. Previous ba9 campaign [37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840) completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline pass, while OpenCode resumes but times out searching for exact tool names and never creates the access card. That failure and incomplete cancelled-run accounting remain preserved; the new catalog error guidance targets this observed dead end. Campaign [37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312) remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker leak and unknown-criterion approval gap. The original baseline remains 7 PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS / 3 FAIL. All fifteen selected cases are qualified by their original grades in this bounded trial. These live grades belong to 09a. Its twelve governed-wait checkpoints retain matching same-turn terminal fingerprints, but passing artifacts omit detailed event journals; the later six adversarial regressions qualify the new terminal-proof and replay guards separately. Final-head repository CI passes. [Fresh Greptile review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780) is 5/5, confirms both findings are fixed, and reports no new actionable issues. All review threads are resolved and the PR has no merge conflicts. ## Risks - Governed cancellation drains accounting for up to fifteen seconds. Restart detachment gives a settling batch up to twenty seconds to finish. An incomplete receipt or unproven provider terminal still prevents successful qualification. - Claude prices are estimates, not invoices. The estimate assumes standard global API pricing and uses a conservative cache-write rate. Unsupported models, billers, and billing modes remain unpriced. - Approval identity matching must remain scoped to the current run and the requested approval. It does not authorize execution of a declined call. - Original baseline and final candidate have different merged master context. Exact-case outcomes are before/after observations, not isolated causal attribution to this repair. - No schema migration, fixture, oracle or grader change. Search-result guidance now treats fuzzy external matches as suggestions. Existing app authorization and provider-consent checks remain required. ## Model Used - OpenAI GPT-6 through Codex, with code editing, terminal tools, and test execution. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (191 current runtime tests and 31 accounting tests; interrupted full-suite history and CI coverage are disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5405b1f46f |
Synchronize chat delivery tests with completed drains (#15480)
Expose the already tracked and caught deferred-work promise to test schedulers while preserving default setImmediate behavior. Assert Slack delivery ordering while the primary lease is held, then join the drains; update the existing GitHub test to assert the resulting quiescent queue. Verified all 1,063 chat tests, full typecheck/build, negative synchronization control, independent review and exact-head CI. An unchanged browser shard retry passed and is documented. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0514e8abd5 |
Keep Cursor managed-runtime test installation offline (#15484)
Intercept the Cursor installer in its local shell fixture and guard against accidental curl downloads. Preserve real managed-home archive operations and the default timeout; change no production code. Verified 46 related tests, safe negative mutation, independent review and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
572cec0344 |
fix(ui): present missing costs as an informational notice (#15459)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Costs dashboard shows spending and gaps in cost data. > - Historical usage can have token counts without a recorded price. > - The red notice implies that the user must fix those records. > - This pull request shows the missing costs as a neutral note below the totals. > - Users can see the limit of the totals without a false call to action. ## Linked Issues or Issue Description Refs #14997. **What happened?** The Costs dashboard shows a red notice when a selected period contains usage without prices. It also shows “0 runs await accounting” when no runs are pending. Historical pricing gaps can therefore look like current errors that need cleanup. **Expected behavior** Show a neutral explanation below the totals. State that totals include only known costs. Show pending runs separately and only when the count is positive. **Steps to reproduce** 1. Open Costs for a period with unpriced usage and no pending runs. 2. Observe the red notice above the totals. 3. Apply this change. Confirm that the notice is muted and below the totals, with no zero-count pending message. **Paperclip version or commit** Reproduced on master after #14997. **Deployment mode** Board UI in both the embedded Activity page and the standalone Costs page. ## What Changed - Move the cost-data notice below the summary tiles and use the muted text token. - Explain unavailable prices with “Totals include known costs only.” - Separate pending runs from missing prices. Omit zero counts and use singular or plural copy as needed. - Update the existing UI tests for both page variants, including missing-only, pending-only, and fully accounted states after refresh. ## Verification - `pnpm --dir ui exec vitest run src/pages/Costs.test.tsx`: 21 tests passed on the final commit. Both page variants cover mixed, missing-only, pending-only, and fully accounted states. - `pnpm check:token-gates`: passed on the final commit. - `pnpm build`: passed locally. - `pnpm -r typecheck`: passed locally. - Full Linux CI passed on `0c23544c0f17e6641cc0c8de6ed498a731748bde`, including browser tests, server tests, shared-package tests, build, and typecheck. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37661883464). - Greptile Apex: 5/5 on the final commit, with no unresolved review threads. - Local full-suite limit: `pnpm test:run` first failed because the fresh checkout lacked embedded PostgreSQL library links. A runtime-skill fixture also selected an unrelated parent directory, and a load test timed out. After setup repair, the database suite (70 tests), skill suite (3 tests), email suite (39 tests), and load suite (4 tests) passed individually. The subsequent full local rerun was stopped after the complete Linux CI run passed. This is not a claim that a full local test invocation passed. ## Risks Low risk. This changes presentation only. The note remains visible and retains its accessible status role. Cost calculations, stored records, and budget enforcement do not change. The copy does not assume that all missing prices are historical. No schema migration or operator documentation change is needed. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code editing, and test execution. The exact served model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
465140596f |
Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials. Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
128c95837b |
Retain bounded orphan process-loss diagnostics (#15475)
Record bounded observer uptime, run and output ages, process-check observations and retry eligibility before orphan cleanup changes the evidence. Preserve existing recovery and reporting behavior and omit process IDs, raw paths and credentials. Verified focused helper, actual reaper and real SDK regressions, full typecheck/build, exact-head CI and independent review. An unchanged Cursor timeout passed its isolated retry; unrelated local Slack timing failure is documented separately. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
43b0aff54b |
Keep pre-dispatch secret configuration blockers out of Sentry (#15472)
Keep confirmed pre-dispatch missing-secret setup blockers out of Sentry while preserving failed runs and owner recovery actions. Runtime, provider and ambiguous failures remain reportable. Verified full exact-head CI, focused local regressions, workspace typecheck and independent review; Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ae6f95ed7a |
feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults. Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
06484b3c41 |
fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Recovery schedules bounded retries after a provider disconnects. > - A stopped run can still hold its environment lease while cleanup runs. > - Retrying before that lease is released cancels the new run before it starts and spends another retry. > - Restoring the task can also send the worker repair instructions from an already resolved recovery action. > - This pull request preserves the waiting retry and removes settled recovery instructions from later wakes. > - The task can continue after cleanup without an operator repairing the same incident again. ## Linked Issues or Issue Description **What happened?** A legacy conversation run disconnected while it was doing ordinary work. Its environment cleanup took longer than the retry delay. Two retries were cancelled before dispatch with `execution_reconciliation_required`. Those cancellations exhausted the failure budget. After an operator restored the task, the wake still told the original worker to repair the runtime and hand the task back to itself. **Expected behavior** Cleanup waits preserve the pending attempt. The same retry can continue after ownership is released, subject to all current gates. Once a recovery action is resolved or cancelled, subsequent task wakes omit its repair instructions. **Steps to reproduce** 1. Fail a legacy conversation run while its environment lease remains in `pending_cleanup`. 2. Schedule a bounded retry and run promotion before cleanup releases that lease. 3. Repeat the scheduler sweep. Before this fix, retries promote and then cancel without starting. 4. Resolve a stranded-task recovery action and build the restored task wake with that action ID. Before this fix, the wake still includes the settled repair instructions. **Paperclip version or commit** Reproduced with database regressions against `ceabc3bc880` on master. Related: #15019 restores a skipped assignment handoff after lease release. #15235 filters stale handoff evidence in the recovery sweep. This change preserves an existing scheduled conversation retry and corrects restored wake content. It does not create a new handoff wake. ## What Changed - Keep an unstarted legacy conversation retry on the same durable row while prior execution ownership remains active. Recheck after 30 seconds without increasing retry accounting. - Return a queued retry to scheduled state if it encounters that hold at the claim gate. Retain its issue claim and publish the status change. - Record one local lifecycle diagnostic per blocking run. Remove that wait marker on promotion. - Include recovery action metadata only while the referenced action is active or escalated. - Return an explicit `waiting` response and the saved schedule when Retry now meets cleanup. Show the wait inline without a false success or disabled button. - Add database, rendered-prompt, route, and UI regressions. Document the execution and run-log contracts. ## Verification - Five cleanup and restored-wake regressions fail against the original production code. The Retry now route and UI regressions also fail before their correction. - Related retry, dispatch, stale-queue, and recovery suites: 321 tests pass across seven files. All ten focused cleanup/restored-wake cases pass after rebase. The final dispatch adjustment passes all 46 adapter tests. - Retry now routes and affected UI suites: all 46 tests pass. The tests cover repeated clicks, the saved schedule, unchanged accounting, promotion after release, and no false success or error state. - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm check:token-gates` pass. - Full local `pnpm test:run` was started and then stopped after the final commit passed all GitHub CI test shards. No complete local full-suite result is claimed; CI supplies the complete test result for the final commit. - Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub checks are green or intentionally skipped. Apex review is 5/5 after two reviews, with no unresolved threads. The PR has no merge conflicts. ## Risks The wait applies only to unstarted legacy conversation retries. Native runs and non-conversation execution keep their existing recovery rules. Cleanup must actually release ownership before execution can resume. The wait does not fix a cleanup service that never finishes. Promotion and dispatch still enforce cancellation, reassignment, pause, budget, and reconciliation gates. No schema change is required. The Retry now response adds a `waiting` outcome; the shared contract and all three UI controls handle it. ## Model Used OpenAI Codex based on GPT-6, with repository analysis, tool use, and local code execution. The runtime does not expose the precise serving model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; complete final-head test coverage in CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |