mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
3dfd3da7f9901728974a3078ebc70c2f65c28358
4914
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3dfd3da7f9 |
docs(hermes): record Linux transport proof and stack replay
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3561e1edca |
ci(hermes): qualify the fixture Node interpreter
Remove group/world write bits from the dedicated CI job's Node interpreter, matching the protected paid workflow setup. Preserve the Rust launch verifier and record the first Linux fixture result and interrupted PR checks. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6cf6cef34a |
fix(hermes): bound cold native admission separately
Give Hermes native session opening a 60-second bound for verified private runtime copying and ACP initialization. Allow controller startup and per-turn restoration overhead while preserving ordinary command and stop deadlines. Add delayed-admission regressions and credential-free Linux native fixture CI. Record the failed final-head paid campaign without promoting qualification. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
42f91e33ac |
docs(hermes): record paid acceptance and qualification limits
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5c2a78fa45 |
ci(hermes): provision pinned runtime for selected paid campaigns
Select candidate assets consistently for immutable image identity and remote packs; provision local Linux assets before the protected paid step exposes credentials. Preserve default-branch dispatch and qualification gates. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
4a0b327727 |
test(hermes): add paid browser qualification campaigns
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1bd40b69e2 |
fix(hermes): release assigned skill copies after failed state save
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
19ab7e06df |
fix(hermes): pin v7 permission policy in the independent server launch
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a2643eb538 |
fix(runner): preserve Dot v6 while versioning Hermes input as v7
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
8cf6f8335d |
fix(hermes): publish and enforce native answer limits
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
fd054630f1 |
fix(runner): bound encoded Hermes attachment messages
Count serialized message, attachments, and metadata before turn admission. Keep TypeScript and Rust on a 7 MiB encoded content allowance within the existing encrypted frame bound. Validate before active-turn state changes, and prove accepted images and escaped documents traverse the secure PRP command frame. Validation: 28 attachment/permission tests, two encrypted-frame regressions, the Rust encoded-admission regression, and Runner TypeScript compilation passed. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6d68b96bc5 |
fix(hermes): preserve planning authority and native transcript results
Qualification exposed denied plan workflows, missing structured tool output, and cancellation replay in subsequent turns. Admit assigned control-plane operations through their existing semantic authority, project denied native calls by identity, and keep managed history restoration separate from transcript replay. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
eda7dd0ab7 |
docs(runner): pass resolved lock digest in image commands
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
63ec34ee79 |
fix(hermes): require the resolved image lock digest
Require the caller-verified target lock checksum instead of an incompatible checked-in default. Document lock-bot ownership and the companion browser qualification PR. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
7d89e304f0 |
fix(acpx): separate restored history from active turn events
Keep history reconstruction without publishing restored text and tools as live output. Version the changed Cursor contract as profile 16 and retain replay compatibility. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
68b1a39459 |
fix(hermes): keep native candidate independently restorable
Reject replaced learned-state parents, include authenticated run attachment and the advertised native transport fixtures in this candidate. Keep paid browser campaign definitions in the qualification PR. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c834f9dd8a | fix(hermes): release credentials and reproduce relocatable runtime assets | ||
|
|
8c8e5e45f2 |
fix(runner): include Hermes in Daytona image identity
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9e07748e91 |
feat(runner): add native Hermes ACP qualification candidate
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
75abb6a774 |
fix(runner): accept native Rust question cancellation envelopes
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d9e17db5ed |
fix(runner): settle stopped native questions without stranded recovery
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
424d9fced8 |
test(routines): retain the current master authority inventory
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c7afea8dab |
fix(routines): reject lock contention before receipt mutations
Keep owned-routine lock acquisition nonblocking after run authority locks. A scheduled firing can finish, and the same mutation may retry without a saved partial receipt. The real concurrency regression is added; local PostgreSQL bootstrap failed before assertions, so its execution still requires CI. Server typecheck passes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
79e5d38139 | fix(routines): remap annotations and verify atomic schedule edits | ||
|
|
49fdc5899e |
feat(runner): bind routine management to the live service
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5717523b9e |
fix: retain the original repository for reused task workspaces (#15528)
Apply the reviewed change for fix: retain the original repository for reused task workspaces. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.4 nightly/v2026.1008.0-nightly.0 |
||
|
|
e4b39da6f6 |
Isolate ACPX admission deadline tests from lease ports (#15527)
Apply the reviewed change for Isolate ACPX admission deadline tests from lease ports. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1d19f9b562 |
Distinguish failed Git inspection from missing worktree registration (#15525)
Apply the reviewed change for Distinguish failed Git inspection from missing worktree registration. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ed0c6958f8 |
Explain refused local Hermes gateway connections (#15520)
Apply the reviewed change for Explain refused local Hermes gateway connections. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6b171c9616 |
Keep proven local workspace configuration conflicts out of Sentry (#15521)
Apply the reviewed change for Keep proven local workspace configuration conflicts out of Sentry. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f4e6362ccc |
chore(lockfile): refresh pnpm-lock.yaml (#15512)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.1008.0-canary.2 |
||
|
|
dd777f4b73 |
feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path
> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.
## Linked Issues or Issue Description
**Subsystem affected**
Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.
**Problem or motivation**
An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.
**Proposed solution**
Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.
**Roadmap alignment**
This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.
## What Changed
- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.
## Verification
- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head
canary/v2026.1008.0-canary.1
|
||
|
|
fc6304dfe5 |
feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path > - Paperclip manages AI agents, tasks, permissions, and execution budgets. > - Paperclip Runner gives each provider the same admitted task and tool authority. > - OpenAI Dot runs outside the local process tree and needs asynchronous work delivery. > - The merged MCP gateway supplies OAuth consent and signed event delivery. > - A personal assistant grant cannot safely stand in for an assigned agent. > - This pull request adds a separate Dot agent connection and a durable Rust Runner bridge. > - The operator can assign work to Dot and inspect its accepted work, tool receipts, and result. ## Linked Issues or Issue Description **Agent or provider** OpenAI Dot, as an experimental provider of the existing Paperclip Runner adapter. **Why this adapter is useful** An operator can assign normal Paperclip tasks to an existing Dot. Dot can read its mailbox, request work on an assigned task, use admitted task tools, and submit a result. Paperclip keeps company scope, checkout, approvals, known budget limits, and activity attribution. **How the agent is invoked** A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the assignment. The Rust Runner owns the durable turn and operation receipts. The first release supports self-hosted instances with a local Runner controller. **Additional context** This extends the merged public MCP gateway from #14846 and the assistant invitation and device-consent work from #14933. This also integrates the merged assistant tool and configuration expansion in #15380. Dot retains its dedicated agent resource and cannot receive personal configuration permission. The public assistant connection remains a personal connection. ## What Changed - Add a durable Rust Dot provider and its TypeScript Runner driver. - Add closed PRP v3 external-provider operations and native execution input v6. - Add company-scoped pairing, mailbox, assignment, and operation records. - Reuse merged browser/device consent, client metadata verification, webhook admissions, refresh, secret rotation, and warm-standby gates. - Keep Dot scopes, issuer, grants, event workers, and tool access separate from personal assistant access. - Add Dot configuration, pairing, readiness, and consent UI. Keep agent grants out of the personal Connections entry. - Regenerate the Dot-only migration after master. Preserve published gateway migrations. Make the new migration safe to reapply. - Document setup, recovery, accounting limits, evidence, and remaining account qualification. - Reverify reconnect callbacks and wake outstanding work with a fresh mailbox reference; preserve the existing assignment and operation receipts. - Clean up Dot bindings and waiting runs on OAuth revoke and refresh-token replay. Old grants cannot revoke replacement bindings. - Restore the pairing reference when an unsaved agent form is reopened; document board-only pairing routes in OpenAPI. - Accept a clean Rust exit after the acknowledged shutdown receipt. Unexpected exits still require recovery. - Clear the cached binding after a successful revoke so a failed connection refresh cannot restore it. - Add production-component Storybook states and screenshots for pairing and connection review. All preview account data is synthetic. - Persist normalized completion, serialize Dot turns and durable work admission, and poll subscription readiness. - Serialize mailbox writes and cursor reads; retain paused fence acknowledgement without task authority. - Authorize admitted review runs without changing the worker assignee. Include the fenced assignment ID in production stop notices. ## Verification - This PR integrates master `4a8178e9c`. Dot migration `0317_messy_famine.sql` follows the published history and is safe to reapply. The merge preserves the reserved migration connection, batch-commit handling, private task checks, task monitors, and native accounting. - Local workspace typecheck, full build, and UI token gates pass. The server typecheck passes after the review fixes. Database and native executor regressions pass. - All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot driver tests pass. The tests cover native document writing and finalization, durable replay, queue admission, mailbox ordering, admitted reviews, stale authority, production stop references, and paused acknowledgements. - Current head `d0e7e0626` passes all 57 checks: 53 pass and four are intentionally skipped. This includes full typecheck, build, tests, Rust Runner verification, browser E2E, release verification, and Canary Dry Run. Greptile rates this exact head 5/5. All review threads are resolved. - The full local root test run is slower than the sharded CI run and has not completed. The full CI test gates pass on the current commit. Focused local regressions pass. - Real-account pairing and event delivery on this base commit remain unqualified. Live account and setup proof are recorded in the follow-up #15414. The following screenshots use synthetic preview data. They show the production pairing component and do not qualify a real account or the full agent setup journey.   ## Risks - This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP and Paperclip Runner. The separate experimental-settings follow-up in #15414 replaces this environment flag with saved operator settings. - Dot does not expose provider token usage or cost. The operator must acknowledge external billing. Known Paperclip budget gates still apply. - Cancellation fences Paperclip authority. It does not confirm that Dot stopped all external activity. - Assigned skill files and third-party MCP bindings are unsupported and reject admission. There is no mounted workspace, model selector, or provider thread identifier. - Hosted agent-broker and remote controller deployments are not qualified. - The new migration follows the merged master history. Existing prototype databases still need the normal master migration history before this Dot-only migration. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window size are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, and test inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.0 |
||
|
|
4a8178e9cf |
fix(ui): stabilize composer model selection and refine effort slider (#15444)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task composers let users select an agent, a model, and an effort level. > - The model label changes width when users move the effort slider. > - This movement makes the control harder to use. > - This pull request keeps the label fixed while the picker is open. > - A larger slider gives clearer feedback during selection. ## Linked Issues or Issue Description **What existing behavior does this improve?** The shared model and effort picker in task composers. **Current behavior** The trigger shows each effort change while the picker is open. Different label lengths move the control. The slider is small. **Proposed behavior** The trigger shows “Select model” while any picker view is open. It shows the selected model and effort after the picker closes. The slider has a thicker track and a larger thumb with hover and press feedback. **Reason and benefit** Users can change effort without moving the trigger. The larger thumb gives clearer pointer feedback. ## What Changed - Keep the trigger text fixed while the picker is open. - Add slider size and feedback tokens. - Remove the gray input outline and thumb border. Show keyboard focus on the thumb. - Keep a thumb focus outline when Windows high contrast suppresses shadows. - Keep native keyboard and pointer controls. Respect reduced motion. - Test open and closed labels on desktop and mobile. ## Verification - Picker tests: 27 pass. Composer settings and new-task tests: 86 pass. - UI typecheck and UI build passed before the final CSS-only high-contrast fix. - Token gates pass. - Full typecheck and build stop at installed Runner API type mismatches outside these files. - Revised Chromium checks pass for desktop pointer drag, no input outline during drag, stable open label, keyboard input, and mobile rendering. The earlier rendered-image check confirmed reduced-motion behavior. - The earlier full local suite did not finish. The approved visual revision passed 113 focused tests. The final high-contrast fix passed 27 picker tests and token gates. - Chromium high-contrast and reduced-motion verification shows a visible thumb focus marker and working keyboard input. - Open the model picker. Change effort. Confirm “Select model” stays visible. Close the picker. Confirm the selected model and effort appear. - All CI gates pass on final head `7a64104e14251a354b06b7916d09d44692557c4f`, including full typecheck, tests, build, and browser E2E. Greptile: 5/5 with no findings. ## Risks - Low risk: shared UI only. Native range controls stay in place. - Browser thumb rendering can differ between engines. ## Model Used - OpenAI GPT-6 Codex. Code editing, tool use, and test execution. The runtime does not expose an exact model variant or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.16 |
||
|
|
378e6d95e1 |
chore(lockfile): refresh pnpm-lock.yaml (#15488)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
a7a244ab33 |
feat: add company decision models with permission and cost controls (#15473)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Optional product features need small, typed model decisions. > - Each company needs to choose which shared API connection pays for those decisions. > - Calls must retain user and task permissions, budget limits, and cost attribution. > - This pull request adds a managed decision service, setup UI, and request history. > - Features can check availability cheaply and keep their existing behavior when decisions are unavailable. ## Linked Issues or Issue Description **Subsystem affected** Server services, shared contracts, database accounting, Company Settings, and Costs. **Problem or motivation** Paperclip has no common decision-model service. Adding provider calls within each feature would duplicate credential access, permission checks, and billing rules. **Proposed solution** Let a connection manager configure one company decision model. Support OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test and metadata-only history. Default company-sponsored background decisions to on during setup, and preserve a saved off setting. **Alternatives considered** Per-feature credentials would duplicate existing connection management. Personal overrides and provider fallback chains add permission and billing complexity; they remain deferred. **Roadmap alignment** Reviewed ROADMAP.md and searched open PRs. This extends existing connection access and budget accounting. Product features that call the service remain outside this change. No matching decision-model service PR was found. ## What Changed - Add company settings, an internal `decisionModelService`, local availability checks, and trusted human, agent/run, and system contexts. - Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and ordered-score decisions for both providers. Bound requests and time; disable paid retries. - Add durable invocation metadata and agentless decision ledger charges. Preserve fractional cents, pricing evidence, dispatch identity, and unresolved billing holds. - Share the company accounting lock and apply company, agent, and project budgets. Settle charges once, retain unknown holds, and recover interrupted calls without resubmission. - Reuse connection setup and management UI. Add a Decisions view under Costs, production-component Storybook coverage, database migration, and service documentation. ## Verification - Passed 166 current-code tests covering the decision service/provider, setup component, Costs, OpenAPI, and every failure from the earlier broad run. Coverage includes native SDK wire formats, refusals, billed malformed responses, permission and secret-rotation races, identity changes, concurrent budget admission, unresolved holds, agent/task deletion, and stale setup feedback. - Passed 170 existing connection, cost, budget, heartbeat-accounting, and profile regression tests. - Passed repository typecheck, production build, and design token gates after integrating master. Verified the generated migration on a fresh test database and upgraded the populated preview database from the branch's earlier migration without losing settings or usage. - Ran the required full `pnpm test:run`: its general phase completed with 16,354 passed and 10 failures across five files while this branch was still being updated. Every reported failure passes in the current-code rerun; the serialized phase did not run after that failure. The full GitHub CI suite passed on `b1b856a88`: general and serialized tests, browser shards, runner checks, typecheck, production build, packaging/canary, and policy gates. [CI evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360). - Passed the full-shell Storybook setup-to-history interaction test again after integrating master. Greptile rates the final revision 5/5 with zero unresolved threads. - Walked through the running app: empty setup, add each provider, save, reload, run all three sample questions, inspect fractional charges in history, switch provider while sponsorship is off, disable, and reconnect. Checked mobile settings. These tests used the actual UI, vault, server, SDKs, and database with simulated upstream responses. - Live paid setup tests remain unverified: this environment has no authorized OpenAI/OpenRouter credentials available. No mocked test is presented as live provider evidence. Reviewer journey: Company Settings → General → Decision model. Add/select a shared API connection, save, run the billed sample, open View usage, then disable decisions and verify Run test is disabled after reload. ## Risks - The SDK decision interface is experimental. Pinned versions and wire-format tests limit upgrade drift. - The migration allows agentless service charges and reservations. Existing agent cost-reporting APIs still require an agent, and decision receipts stay separate from run reconciliation. - Timeouts can have unknown provider charges. Holds remain until an audited accounting correction resolves them. - OpenAI prices use a versioned Decisions rate snapshot; OpenRouter costs use provider receipts. Unknown pricing is retained as unknown. - A configured company authorizes background spending by default. Setup explains this, and managers can turn it off. - Live provider account/model availability still needs the two credentialed acceptance checks. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser tools. The exact serving revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.15 |
||
|
|
f1c44b7b56 |
Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs. Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
cd40ebb95b |
Preserve ACP bridge terminal evidence before disconnect (#15478)
Drain terminal frames through backpressure and record bounded terminal evidence without waiting for log persistence. Guard later events and input failures against writes after socket end. Validated with 217 focused tests, full typecheck/build, independent review and green CI with Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.14 |
||
|
|
50b1f95e79 |
fix(tool-gateway): keep MCP connection healthy on oversized/malformed responses (#15462)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents call remote MCP tools through the Paperclip tool gateway > - The gateway keeps a health status for each remote MCP connection, and it hides the tools of a connection that has `healthStatus = "error"` > - One oversized, malformed, or invalid-JSON `tools/call` reply set the whole connection to `error` > - Health goes back to `ok` only after a successful call, but the hidden tools prevent that call, so the connection stayed locked until a person reconnected it > - This pull request makes these reply errors fail only the one call, and it starts a new MCP session for the caller on the next call > - The benefit is that one large or bad reply no longer disconnects a working connection for every agent ## Linked Issues or Issue Description No public issue exists. Related open PRs fix other health downgrades in the same function. They do not overlap with this change: - Refs #11910 (timeouts and JSON-RPC errors) - Refs #15325 (transport errors) **What happened?** An agent called a remote MCP tool that returned more than `MAX_REMOTE_MCP_RESPONSE_BYTES` (1 MB). The gateway returned 502 `mcp_remote_response_too_large`. It also set the connection to `healthStatus = "error"`. Then `connectedMcpConnectionFilter` hid all tools of the connection (except `per_user` connections), and `connections_search` showed the connection as `needs_user_action`. The `malformed_response` and `invalid_json` errors from `readMcpHttpResponse` did the same thing. `readMcpHttpResponse` can also cancel the reply stream before the end. The remote server can then close its MCP session. The gateway kept the cached `mcp-session-id` for up to 30 minutes and cleared it only after a 404. The next calls then failed with "Remote MCP session expired". **Expected behavior** A reply that is too large or not correct fails only that call. The connection stays healthy, and its tools stay visible. The next call uses a new MCP session. **Steps to reproduce** 1. Add a remote MCP connection with `mcpSessionRequired: true`. 2. Call a tool that returns a reply larger than 1 MB. 3. Look at the connection: `healthStatus` is `error`. 4. Start a new gateway session: the tools of the connection are not in the list. **Paperclip version or commit** `master` at `99a9de9940bf5974352d9dbfbb2f21e62e89689f` **Deployment mode** All modes. The fault is in the server tool gateway. ## What Changed - `server/src/services/tool-gateway.ts`: `too_large`, `malformed_response`, and `invalid_json` from `readMcpHttpResponse` now fail only the call. The caller gets the same 502 reason code as before. The gateway does not call `markRemoteConnectionHealth(…, "error")` for these errors. The `invalid_json` branch after `JSON.parse(body)` also does not change health now, so all `invalid_json` paths are the same. - The `mcp_remote_response_too_large` message now tells the caller to request a smaller result, for example a narrower query or a smaller page size. The error details now include `maxBytes`. - `server/src/services/mcp-http.ts`: new `forgetMcpHttpSession()`. It removes only the cached session for one scope and one credential set, and only while that entry still holds the session ID that failed. It does not touch other agents, sessions that are initializing, or a newer session that replaced the failed one. - The gateway calls `forgetMcpHttpSession()` after these reply errors, so the next call from that caller initializes a new session. The existing 404 "session expired" path now uses the same function. Before, both paths cleared all sessions and all pending initializations on the connection. That made a concurrent initialization by another agent fail with "MCP connection changed while initializing", which the gateway reported as a fetch failure and marked as a connection `error`. - `forgetMcpHttpSessions(connectionId)` is not changed. Disconnect and revocation still clear the full connection. Why `malformed_response` and `invalid_json` are also per-call errors: - Each error is about one reply body. The server was reachable, accepted the credentials, and sent HTTP 2xx. A reply can be too large or bad because of the tool and its arguments. That is not a fault of the connection. - The other malformed-reply checks in the same function (payload is not an object, or has no `result`) already throw `remote_mcp_malformed_response` and do not change health. Only the reader-level errors changed health. - Health recovers only after a successful call. An `error` status for one bad reply therefore locks the connection until a person reconnects it. - Other failures (HTTP errors, fetch failures, timeouts, JSON-RPC errors) still change health. This PR does not change them. #11910 and #15325 address some of them. ## Verification - `server/src/__tests__/tool-gateway.test.ts`: the recovery test now runs for three replies: oversized, invalid JSON, and a reply without the requested message ID. Each run uses a fake HTTP MCP server with `mcpSessionRequired: true` that gives `session-N` for each `initialize`. Each run checks that: - the call fails with the correct 502 reason code (and, for the oversized reply, the smaller-page hint and `maxBytes: 1000000`) - `healthStatus` stays `ok` - a new gateway session still lists the tool and can call it - the second `tools/call` uses `session-2`, not `session-1` - New test: agent A gets an oversized reply while agent B initializes a session on the same connection. Agent B's call completes, health stays `ok`, and the two calls use `session-1` and `session-2`. With the old connection-wide reset, this test fails: agent B gets 502 `mcp_remote_fetch_failed`. - `server/src/__tests__/remote-mcp-protocol.test.ts`: new unit test. `forgetMcpHttpSession()` keeps the session of a different identity, and a late failure from an old session does not remove the newer session. - Each regression check fails without its fix. Without the health change, the test fails on `healthStatus: 'error'`. Without the session reset, it fails with `['session-1', 'session-1']`. ``` cd server npx vitest run src/__tests__/tool-gateway.test.ts src/__tests__/tool-gateway-service.test.ts src/__tests__/remote-mcp-protocol.test.ts Test Files 3 passed (3) Tests 132 passed (132) ``` - I did not run the full server `tsc --noEmit` locally because the sandbox does not have sufficient memory. The CI typecheck covers it. ## Risks - Low risk. The change affects only the error path of remote MCP `tools/call`. - A remote server that always sends bad replies now keeps `healthStatus = "ok"`. Each call still fails with a clear 502 reason code and an audit record, so the failure stays visible. Only the connection-wide hiding of tools stops. - After one of these errors, the next call from the same caller sends one more `initialize` request. This adds one round trip. - A 404 "session expired" now clears only the session of the caller that got the 404. Before, it cleared the cached sessions of all agents on the connection. If the remote server restarts, each agent now gets its own "session expired" error one time and then initializes again. A 404 is about one session, and the old connection-wide clear also cancelled other agents' initializations. - #11910 and #15325 change the same catch block. The PR that merges last can have a small merge conflict. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, 1M-token context window. - Run as an agent in Claude Code with tool use (shell, file edit, and local test runs). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing documentation changes are necessary) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge This PR replaces #15418. It keeps the same commits on a branch name without an internal ticket ID, and adds a fix for the Greptile review on #15418. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
fd8c6b920a |
fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native tasks can pause while a human decides whether to allow a connection action. > - The next turn needs both the saved decision and complete accounting for the previous turn. > - Cancellation could discard final usage, and complete direct Claude API receipts could remain unpriced. > - A stale blocked or review report could also request approval again after the original card was declined. > - This pull request retains shutdown accounting and rejects approval waits bound to an already resolved action. > - The benefit is reliable continuation with the existing budget and approval controls. ## Linked Issues or Issue Description Refs #15420. Related: #15312 addresses requester ownership during dispatch. This change addresses receipt capture and final-response validation. ## What Changed - Retain usage events after native cancellation. Continue to reject late provider messages and work. - Drain same-turn accounting and terminal events for at most fifteen seconds after a durable governed wait. Keep incomplete accounting blocked. - Estimate complete, unpriced, direct Anthropic API receipts for the exact `claude-sonnet-5` model. Record the rate version and assumptions. Use the one-hour cache-write rate when the receipt lacks cache TTL. - Bind stale approval reports to exact interaction, action-request, or invocation IDs in the same company, task, agent, and run. Cover blocked, review, and response-wake reports. Keep independent reviews valid. - Fence checkpoint and result writes after a controller detaches for restart, including operations waiting for a database lock. Reject stale successful returns before certifying accounting. - Allow bounded subscription teardown only after retaining an actual provider terminal. - Journal the exact governed-wait trigger and disposition before provider interruption. Recover that wait independently of a later saved answer, replay retained accounting, and reject mismatched or unproven terminal evidence. - Give settling governed turns a bounded window before shutdown detaches their controller. - Keep fuzzy external app matches alongside installed capability matches instead of forcing an unrelated provider question for a generic query. - Return up to twenty exact active catalog tool names after an invalid request, after eligibility checks; still reject the request without granting access or creating an approval. - Require retained provider terminal proof before settling a governed wait, including when complete usage arrives before stream closure/error/timeout. Retain harmless numbered cancellation events so restart replay stays contiguous. - Isolate accounting-test OpenCode config from the host plugin directory. - Add regression coverage and document the accounting, restart and connection-search behavior. ## Verification Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`; later review fixes are verified separately and do not relabel those runs. - Repository typecheck and build pass. - Current runner runtime and cancellation suites: 191 tests pass. Six new regressions cover stream end/error/timeout without provider-stop proof and contiguous cancellation acknowledgement/request replay; all six failed before the fix. The existing bounded cleanup case now explicitly supplies terminal proof. All 31 adapter accounting tests pass with isolated fixture config. - New checkpoint-rebinding and approval-criterion suites: 65 tests pass; runner HTTP integration: 29 tests pass. Unchanged executor/control-plane suites: 640 tests pass; database-backed connection suites: 68 tests pass. - Current retained evidence verification covers 222 file hashes across all fifteen original result artifacts. The three-case [published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html) and all eight screenshot hashes verify. The [three-case recovery report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html) and all six screenshots also verify; the nine-result campaign did not publish. - Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node tests pass. Eval typecheck and catalog discovery pass. - Full local repository unit run was interrupted before the follow-up edits after three tool-access failures and one runner HTTP failure. Those failures pass in isolation; the 09a full run was interrupted after one rapid Slack callback-ordering failure and seven skill-service failures. All eight pass both isolated and with full-runner environment settings, and all 74 skill-service tests pass together; the subsequent full run reported two 15-second OpenCode accounting timeouts and was stopped with exit 130 to apply review fixes. The timeouts reproduce while copying this host’s 61 MB OpenCode config. All 31 tests pass after isolating config inside each fixture without increasing timeouts or changing assertions. A complete local full-suite pass is not claimed. Repository-wide CI also passes on the final review-fix head in [run 37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952). Earlier database-skipped diagnostics and the older ENFILE run are retained and are not full-suite passing evidence. - Three-case live campaign [37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525) passes all three original grades on frozen source 09a: 55/55 checks, eight succeeded run records, complete accounting receipts, matching checkpoint identities and no pending approvals. Campaign [37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353) adds nine original passes (173/173 checks, eighteen succeeded records) on the identical source. Its other three jobs failed before runner assignment or any step while GitHub could not load the paid environment; those original infrastructure failures are retained. Campaign [37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356) completes only those unstarted cells: all three original grades pass (57/57 checks, six succeeded run records). All fifteen exact cases now pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new overall failures and zero pending pairs. Total current evidence: 285/285 checks and thirty-two succeeded run records, complete accounting receipts, matching checkpoint identities, no pending approvals or retry records. Earlier campaigns retain forty-seven additional run records and two known same-run recovery attempts; actual provider-call counts and invoices remain unknown. The nine-result campaign skipped publication and its public URL returns 403; original artifacts remain retained. Previous ba9 campaign [37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840) completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline pass, while OpenCode resumes but times out searching for exact tool names and never creates the access card. That failure and incomplete cancelled-run accounting remain preserved; the new catalog error guidance targets this observed dead end. Campaign [37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312) remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker leak and unknown-criterion approval gap. The original baseline remains 7 PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS / 3 FAIL. All fifteen selected cases are qualified by their original grades in this bounded trial. These live grades belong to 09a. Its twelve governed-wait checkpoints retain matching same-turn terminal fingerprints, but passing artifacts omit detailed event journals; the later six adversarial regressions qualify the new terminal-proof and replay guards separately. Final-head repository CI passes. [Fresh Greptile review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780) is 5/5, confirms both findings are fixed, and reports no new actionable issues. All review threads are resolved and the PR has no merge conflicts. ## Risks - Governed cancellation drains accounting for up to fifteen seconds. Restart detachment gives a settling batch up to twenty seconds to finish. An incomplete receipt or unproven provider terminal still prevents successful qualification. - Claude prices are estimates, not invoices. The estimate assumes standard global API pricing and uses a conservative cache-write rate. Unsupported models, billers, and billing modes remain unpriced. - Approval identity matching must remain scoped to the current run and the requested approval. It does not authorize execution of a declined call. - Original baseline and final candidate have different merged master context. Exact-case outcomes are before/after observations, not isolated causal attribution to this repair. - No schema migration, fixture, oracle or grader change. Search-result guidance now treats fuzzy external matches as suggestions. Existing app authorization and provider-consent checks remain required. ## Model Used - OpenAI GPT-6 through Codex, with code editing, terminal tools, and test execution. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (191 current runtime tests and 31 accounting tests; interrupted full-suite history and CI coverage are disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5405b1f46f |
Synchronize chat delivery tests with completed drains (#15480)
Expose the already tracked and caught deferred-work promise to test schedulers while preserving default setImmediate behavior. Assert Slack delivery ordering while the primary lease is held, then join the drains; update the existing GitHub test to assert the resulting quiescent queue. Verified all 1,063 chat tests, full typecheck/build, negative synchronization control, independent review and exact-head CI. An unchanged browser shard retry passed and is documented. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0514e8abd5 |
Keep Cursor managed-runtime test installation offline (#15484)
Intercept the Cursor installer in its local shell fixture and guard against accidental curl downloads. Preserve real managed-home archive operations and the default timeout; change no production code. Verified 46 related tests, safe negative mutation, independent review and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
572cec0344 |
fix(ui): present missing costs as an informational notice (#15459)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Costs dashboard shows spending and gaps in cost data. > - Historical usage can have token counts without a recorded price. > - The red notice implies that the user must fix those records. > - This pull request shows the missing costs as a neutral note below the totals. > - Users can see the limit of the totals without a false call to action. ## Linked Issues or Issue Description Refs #14997. **What happened?** The Costs dashboard shows a red notice when a selected period contains usage without prices. It also shows “0 runs await accounting” when no runs are pending. Historical pricing gaps can therefore look like current errors that need cleanup. **Expected behavior** Show a neutral explanation below the totals. State that totals include only known costs. Show pending runs separately and only when the count is positive. **Steps to reproduce** 1. Open Costs for a period with unpriced usage and no pending runs. 2. Observe the red notice above the totals. 3. Apply this change. Confirm that the notice is muted and below the totals, with no zero-count pending message. **Paperclip version or commit** Reproduced on master after #14997. **Deployment mode** Board UI in both the embedded Activity page and the standalone Costs page. ## What Changed - Move the cost-data notice below the summary tiles and use the muted text token. - Explain unavailable prices with “Totals include known costs only.” - Separate pending runs from missing prices. Omit zero counts and use singular or plural copy as needed. - Update the existing UI tests for both page variants, including missing-only, pending-only, and fully accounted states after refresh. ## Verification - `pnpm --dir ui exec vitest run src/pages/Costs.test.tsx`: 21 tests passed on the final commit. Both page variants cover mixed, missing-only, pending-only, and fully accounted states. - `pnpm check:token-gates`: passed on the final commit. - `pnpm build`: passed locally. - `pnpm -r typecheck`: passed locally. - Full Linux CI passed on `0c23544c0f17e6641cc0c8de6ed498a731748bde`, including browser tests, server tests, shared-package tests, build, and typecheck. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37661883464). - Greptile Apex: 5/5 on the final commit, with no unresolved review threads. - Local full-suite limit: `pnpm test:run` first failed because the fresh checkout lacked embedded PostgreSQL library links. A runtime-skill fixture also selected an unrelated parent directory, and a load test timed out. After setup repair, the database suite (70 tests), skill suite (3 tests), email suite (39 tests), and load suite (4 tests) passed individually. The subsequent full local rerun was stopped after the complete Linux CI run passed. This is not a claim that a full local test invocation passed. ## Risks Low risk. This changes presentation only. The note remains visible and retains its accessible status role. Cost calculations, stored records, and budget enforcement do not change. The copy does not assume that all missing prices are historical. No schema migration or operator documentation change is needed. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code editing, and test execution. The exact served model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
465140596f |
Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials. Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
128c95837b |
Retain bounded orphan process-loss diagnostics (#15475)
Record bounded observer uptime, run and output ages, process-check observations and retry eligibility before orphan cleanup changes the evidence. Preserve existing recovery and reporting behavior and omit process IDs, raw paths and credentials. Verified focused helper, actual reaper and real SDK regressions, full typecheck/build, exact-head CI and independent review. An unchanged Cursor timeout passed its isolated retry; unrelated local Slack timing failure is documented separately. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
43b0aff54b |
Keep pre-dispatch secret configuration blockers out of Sentry (#15472)
Keep confirmed pre-dispatch missing-secret setup blockers out of Sentry while preserving failed runs and owner recovery actions. Runtime, provider and ambiguous failures remain reportable. Verified full exact-head CI, focused local regressions, workspace typecheck and independent review; Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ae6f95ed7a |
feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults. Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
06484b3c41 |
fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Recovery schedules bounded retries after a provider disconnects. > - A stopped run can still hold its environment lease while cleanup runs. > - Retrying before that lease is released cancels the new run before it starts and spends another retry. > - Restoring the task can also send the worker repair instructions from an already resolved recovery action. > - This pull request preserves the waiting retry and removes settled recovery instructions from later wakes. > - The task can continue after cleanup without an operator repairing the same incident again. ## Linked Issues or Issue Description **What happened?** A legacy conversation run disconnected while it was doing ordinary work. Its environment cleanup took longer than the retry delay. Two retries were cancelled before dispatch with `execution_reconciliation_required`. Those cancellations exhausted the failure budget. After an operator restored the task, the wake still told the original worker to repair the runtime and hand the task back to itself. **Expected behavior** Cleanup waits preserve the pending attempt. The same retry can continue after ownership is released, subject to all current gates. Once a recovery action is resolved or cancelled, subsequent task wakes omit its repair instructions. **Steps to reproduce** 1. Fail a legacy conversation run while its environment lease remains in `pending_cleanup`. 2. Schedule a bounded retry and run promotion before cleanup releases that lease. 3. Repeat the scheduler sweep. Before this fix, retries promote and then cancel without starting. 4. Resolve a stranded-task recovery action and build the restored task wake with that action ID. Before this fix, the wake still includes the settled repair instructions. **Paperclip version or commit** Reproduced with database regressions against `ceabc3bc880` on master. Related: #15019 restores a skipped assignment handoff after lease release. #15235 filters stale handoff evidence in the recovery sweep. This change preserves an existing scheduled conversation retry and corrects restored wake content. It does not create a new handoff wake. ## What Changed - Keep an unstarted legacy conversation retry on the same durable row while prior execution ownership remains active. Recheck after 30 seconds without increasing retry accounting. - Return a queued retry to scheduled state if it encounters that hold at the claim gate. Retain its issue claim and publish the status change. - Record one local lifecycle diagnostic per blocking run. Remove that wait marker on promotion. - Include recovery action metadata only while the referenced action is active or escalated. - Return an explicit `waiting` response and the saved schedule when Retry now meets cleanup. Show the wait inline without a false success or disabled button. - Add database, rendered-prompt, route, and UI regressions. Document the execution and run-log contracts. ## Verification - Five cleanup and restored-wake regressions fail against the original production code. The Retry now route and UI regressions also fail before their correction. - Related retry, dispatch, stale-queue, and recovery suites: 321 tests pass across seven files. All ten focused cleanup/restored-wake cases pass after rebase. The final dispatch adjustment passes all 46 adapter tests. - Retry now routes and affected UI suites: all 46 tests pass. The tests cover repeated clicks, the saved schedule, unchanged accounting, promotion after release, and no false success or error state. - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm check:token-gates` pass. - Full local `pnpm test:run` was started and then stopped after the final commit passed all GitHub CI test shards. No complete local full-suite result is claimed; CI supplies the complete test result for the final commit. - Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub checks are green or intentionally skipped. Apex review is 5/5 after two reviews, with no unresolved threads. The PR has no merge conflicts. ## Risks The wait applies only to unstarted legacy conversation retries. Native runs and non-conversation execution keep their existing recovery rules. Cleanup must actually release ownership before execution can resume. The wait does not fix a cleanup service that never finishes. Promotion and dispatch still enforce cancellation, reassignment, pause, budget, and reconciliation gates. No schema change is required. The Retry now response adds a `waiting` outcome; the shared contract and all three UI controls handle it. ## Model Used OpenAI Codex based on GPT-6, with repository analysis, tool use, and local code execution. The runtime does not expose the precise serving model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; complete final-head test coverage in CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db0bee19fa |
docs(readme): add Superagent badge and move Star History rank badge to its own line (#15460)
- Add the Superagent security badge to the README badge row, with light and dark variants via <picture>. - Move the Star History global rank badge into its own centered paragraph below the badge row, with spacing. Co-Authored-By: Amy (Opus) via Paperclip <claude@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
9fb955e9a2 |
fix(codex): gate models on the Codex CLI floor before the ChatGPT backend rejects them (#15399)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Codex adapter and the native runner run Codex in local hosts and in managed sandboxes, and the model catalog lists the models an operator can select > - The Codex backend accepts a new model with ChatGPT sign-in only from a recent enough Codex CLI. An older CLI fails every turn with `The '<model>' model is not supported when using Codex with a ChatGPT account.` > - After #14942 added `gpt-6.1-sol`, an operator selected it for an agent in a managed Daytona sandbox. The sandbox image still shipped Codex 0.156.0. The connection test and runs failed with that sentence, which reads like an account problem > - Paperclip only checked the Codex compatibility window (`>=0.149.0 <0.161.0`), so 0.156.0 passed and the failure surfaced from the backend without a cause > - This pull request records the verified Codex CLI floor for each gated model, compares the installed `codex --version` with that floor in the environment Test and in the remote runner, and names the backend rejection when it still happens > - The benefit is a precise, actionable message ("gpt-6.1-sol requires Codex CLI 0.159.0 or newer; detected 0.156.0; promote a sandbox image with Codex 0.160.0") instead of an opaque 400, and no doomed hello probe ## Linked Issues or Issue Description Refs #14942 **Bug: selecting `gpt-6.1-sol` in a managed sandbox fails with a ChatGPT account error.** **Steps to reproduce** 1. Promote a sandbox image that ships Codex CLI 0.156.0. 2. Create a Codex agent with a ChatGPT sign-in connection, select `gpt-6.1-sol`, and run the environment Test or a task in that sandbox. **Expected behavior** The Test or the run tells the operator that the Codex CLI in the sandbox is older than the model needs. **Actual behavior** The Test reports `codex_hello_probe_failed` with `{"type":"error","status":400,"error":{"type":"invalid_request_error","message":"The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account."}}`. A run fails with "The selected model is not supported by the current ChatGPT connection." **Evidence** - Codex 0.156.1 and older are rejected for `gpt-6.1-sol` with ChatGPT sign-in: https://github.com/openai/codex/issues/49396 - Codex 0.159.0 and 0.159.2 are accepted: https://github.com/openai/codex/issues/49464 and https://github.com/decolua/9router/issues/4471 (same error, fixed by raising the client version identity from 0.155.0 to 0.159.0) - Codex 0.157.0 release notes: "Add GPT-6 Sol and Luna to the model catalog": https://github.com/openai/codex/releases/tag/rust-v0.157.0 - GPT-6.1 Sol is included for Plus, Pro, Business, Enterprise, and Edu with ChatGPT sign-in: https://learn.chatgpt.com/docs/models ## What Changed - `packages/adapters/codex-local/src/index.ts`: add `minimumCodexCliVersionForModel` (`gpt-6.1-sol` → 0.159.0, `gpt-6-sol` and `gpt-6-luna` → 0.157.0, with the sources above), `parseCodexCliVersionOutput`, `codexCliVersionAtLeast`, and `CODEX_CHATGPT_MODEL_REJECTION_RE`. Models without a verified floor return `null`, so nothing probes them. The agent configuration doc describes the floors. - `packages/adapters/codex-local/src/server/cli-version.ts` (new): run `codex --version` where the run would execute it (local, SSH, or sandbox), compare it with the model floor, and build the `codex_cli_version_compatible` / `codex_cli_version_incompatible` checks. Map the backend rejection sentence to `codex_hello_probe_model_rejected` with the detected CLI version and a hint that separates a stale CLI from a plan that does not include the model. - `packages/adapters/codex-local/src/server/test.ts` (CLI lane Test): run the version check before the hello probe when the model has a floor. Skip the hello probe when the CLI is too old (`codex_hello_probe_skipped_cli_version`). When the probe still fails with the backend sentence, report `codex_hello_probe_model_rejected` instead of the generic `codex_hello_probe_failed`. The new codes do not match the auth-failure patterns, so a managed connection is not invalidated. - `packages/adapters/codex-local/src/server/acp.ts` (ACP lane Test): for remote targets, run the same version check against the shared `codex` the ACP server spawns. - `server/src/services/native-runtime/native-session-executor.ts`: after the compatibility-window check, compare the remote Codex with the configured model's floor and fail with `runner_remote_provider_artifact_incompatible: <model> requires Codex <floor> or newer with ChatGPT sign-in, received <version> from the sandbox image; promote a sandbox image with Codex 0.160.0 or configure PAPERCLIP_RUNNER_REMOTE_CODEX_NPM_SPEC=...`. When a preinstalled Codex fails this check and an npm spec is configured, the existing fallback installs the pinned release. - Tests: adapter metadata, CLI-lane Test (too old, compatible, no floor, backend rejection), ACP-lane Test (too old, compatible, no floor), and the remote runner floor (nine version/model cases). - `doc/adapter-model-audit-2026-10-02.md`: record the floors and the reason. No catalog entry, default model, saved agent configuration, pin, lockfile, or image definition changes. ## Verification Local (Node 25.9, pnpm 9.15.4 via corepack): ```sh pnpm --filter @paperclipai/adapter-codex-local typecheck pnpm --filter @paperclipai/adapter-codex-local exec vitest run # 81 passed pnpm --filter @paperclipai/server exec vitest run \ src/services/native-runtime/native-session-executor.test.ts \ src/services/native-runtime/codex-runtime-compatibility.test.ts -t Codex # 84 passed ``` Also run locally: the full `native-session-executor.test.ts` file (533 passed) and the full Codex adapter suite (81 passed). Follow-up commit `9fd4d3b25` (Greptile P2: the ACP-lane version probe ignored the agent's configured env): the probe now receives the adapter's string-valued `env` entries, the same ones `buildCodexAcpConfig` hands to remote ACP runs, so a `PATH` override selects the same `codex` for the Test as for the run. New test covers a `PATH` + `CODEX_HOME` override and a dropped non-string entry. Re-run on that head: `pnpm --filter @paperclipai/adapter-codex-local typecheck` clean; full Codex adapter suite 503 passed (31 files). Not run here: `pnpm -r typecheck` for `server` (the direct `tsc --noEmit` was killed by the sandbox memory cap; the adapter package typecheck passes and CI covers the server), `pnpm build`, browser suites, and a live sandbox probe. No UI files changed, so token gates do not apply. Manual check for a reviewer: set an agent's Codex model to `gpt-6.1-sol`, point its environment at a sandbox whose `codex --version` prints `codex-cli 0.156.0`, and run the environment Test. The result contains `codex_cli_version_incompatible` with `Detected Codex CLI 0.156.0.` and no hello probe check. With Codex 0.160.0 the Test contains `codex_cli_version_compatible` and runs the hello probe. ## Risks - A floor is enforced for all authentication modes, because the runner does not know the credential method at verification time. An OpenAI API key on Codex 0.157.0 or 0.158.0 with `gpt-6.1-sol` is now rejected before launch. Paperclip pins Codex 0.160.0 everywhere, so this only affects installs that lag the pin, and the message names the fix. - The 0.159.0 floor for `gpt-6.1-sol` is the oldest stable release verified to work; 0.157.0 and 0.158.0 were not verified either way. If OpenAI accepts an older client, the floor can be lowered in one table. - The `codex --version` probe adds one short process run to the Test only for models with a floor. - Rollback: revert this pull request. No data or configuration migrates. ## Model Used Claude Fable 5.1 (`claude-fable-5-1`), Anthropic, with extended thinking, tool use, and web research, operating as a Paperclip agent through Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Bender (Fable) <noreply@paperclip.ing> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>canary/v2026.1007.0-canary.11 |