mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
codex/hermes-native-runner
4905
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2922f65ba1 |
fix(hermes): pin v7 permission policy in the independent server launch
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b0f54fc1a2 |
fix(runner): preserve Dot v6 while versioning Hermes input as v7
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9d33a4d8a4 |
fix(hermes): publish and enforce native answer limits
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
dc76966146 |
fix(runner): bound encoded Hermes attachment messages
Count serialized message, attachments, and metadata before turn admission. Keep TypeScript and Rust on a 7 MiB encoded content allowance within the existing encrypted frame bound. Validate before active-turn state changes, and prove accepted images and escaped documents traverse the secure PRP command frame. Validation: 28 attachment/permission tests, two encrypted-frame regressions, the Rust encoded-admission regression, and Runner TypeScript compilation passed. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ef7cf8c398 |
fix(hermes): preserve planning authority and native transcript results
Qualification exposed denied plan workflows, missing structured tool output, and cancellation replay in subsequent turns. Admit assigned control-plane operations through their existing semantic authority, project denied native calls by identity, and keep managed history restoration separate from transcript replay. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
97dc4efd53 |
docs(runner): pass resolved lock digest in image commands
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2f15770a88 |
fix(hermes): require the resolved image lock digest
Require the caller-verified target lock checksum instead of an incompatible checked-in default. Document lock-bot ownership and the companion browser qualification PR. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
05d6a8f6e2 |
fix(acpx): separate restored history from active turn events
Keep history reconstruction without publishing restored text and tools as live output. Version the changed Cursor contract as profile 16 and retain replay compatibility. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0c75f68414 |
fix(hermes): keep native candidate independently restorable
Reject replaced learned-state parents, include authenticated run attachment and the advertised native transport fixtures in this candidate. Keep paid browser campaign definitions in the qualification PR. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
fdb94c95ef | fix(hermes): release credentials and reproduce relocatable runtime assets | ||
|
|
9d1512a169 |
fix(runner): include Hermes in Daytona image identity
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ba9a809ea2 |
feat(runner): add native Hermes ACP qualification candidate
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
424d9fced8 |
test(routines): retain the current master authority inventory
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c7afea8dab |
fix(routines): reject lock contention before receipt mutations
Keep owned-routine lock acquisition nonblocking after run authority locks. A scheduled firing can finish, and the same mutation may retry without a saved partial receipt. The real concurrency regression is added; local PostgreSQL bootstrap failed before assertions, so its execution still requires CI. Server typecheck passes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
79e5d38139 | fix(routines): remap annotations and verify atomic schedule edits | ||
|
|
49fdc5899e |
feat(runner): bind routine management to the live service
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5717523b9e |
fix: retain the original repository for reused task workspaces (#15528)
Apply the reviewed change for fix: retain the original repository for reused task workspaces. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.4 |
||
|
|
e4b39da6f6 |
Isolate ACPX admission deadline tests from lease ports (#15527)
Apply the reviewed change for Isolate ACPX admission deadline tests from lease ports. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1d19f9b562 |
Distinguish failed Git inspection from missing worktree registration (#15525)
Apply the reviewed change for Distinguish failed Git inspection from missing worktree registration. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ed0c6958f8 |
Explain refused local Hermes gateway connections (#15520)
Apply the reviewed change for Explain refused local Hermes gateway connections. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6b171c9616 |
Keep proven local workspace configuration conflicts out of Sentry (#15521)
Apply the reviewed change for Keep proven local workspace configuration conflicts out of Sentry. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f4e6362ccc |
chore(lockfile): refresh pnpm-lock.yaml (#15512)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.1008.0-canary.2 |
||
|
|
dd777f4b73 |
feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path
> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.
## Linked Issues or Issue Description
**Subsystem affected**
Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.
**Problem or motivation**
An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.
**Proposed solution**
Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.
**Roadmap alignment**
This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.
## What Changed
- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.
## Verification
- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head
canary/v2026.1008.0-canary.1
|
||
|
|
fc6304dfe5 |
feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path > - Paperclip manages AI agents, tasks, permissions, and execution budgets. > - Paperclip Runner gives each provider the same admitted task and tool authority. > - OpenAI Dot runs outside the local process tree and needs asynchronous work delivery. > - The merged MCP gateway supplies OAuth consent and signed event delivery. > - A personal assistant grant cannot safely stand in for an assigned agent. > - This pull request adds a separate Dot agent connection and a durable Rust Runner bridge. > - The operator can assign work to Dot and inspect its accepted work, tool receipts, and result. ## Linked Issues or Issue Description **Agent or provider** OpenAI Dot, as an experimental provider of the existing Paperclip Runner adapter. **Why this adapter is useful** An operator can assign normal Paperclip tasks to an existing Dot. Dot can read its mailbox, request work on an assigned task, use admitted task tools, and submit a result. Paperclip keeps company scope, checkout, approvals, known budget limits, and activity attribution. **How the agent is invoked** A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the assignment. The Rust Runner owns the durable turn and operation receipts. The first release supports self-hosted instances with a local Runner controller. **Additional context** This extends the merged public MCP gateway from #14846 and the assistant invitation and device-consent work from #14933. This also integrates the merged assistant tool and configuration expansion in #15380. Dot retains its dedicated agent resource and cannot receive personal configuration permission. The public assistant connection remains a personal connection. ## What Changed - Add a durable Rust Dot provider and its TypeScript Runner driver. - Add closed PRP v3 external-provider operations and native execution input v6. - Add company-scoped pairing, mailbox, assignment, and operation records. - Reuse merged browser/device consent, client metadata verification, webhook admissions, refresh, secret rotation, and warm-standby gates. - Keep Dot scopes, issuer, grants, event workers, and tool access separate from personal assistant access. - Add Dot configuration, pairing, readiness, and consent UI. Keep agent grants out of the personal Connections entry. - Regenerate the Dot-only migration after master. Preserve published gateway migrations. Make the new migration safe to reapply. - Document setup, recovery, accounting limits, evidence, and remaining account qualification. - Reverify reconnect callbacks and wake outstanding work with a fresh mailbox reference; preserve the existing assignment and operation receipts. - Clean up Dot bindings and waiting runs on OAuth revoke and refresh-token replay. Old grants cannot revoke replacement bindings. - Restore the pairing reference when an unsaved agent form is reopened; document board-only pairing routes in OpenAPI. - Accept a clean Rust exit after the acknowledged shutdown receipt. Unexpected exits still require recovery. - Clear the cached binding after a successful revoke so a failed connection refresh cannot restore it. - Add production-component Storybook states and screenshots for pairing and connection review. All preview account data is synthetic. - Persist normalized completion, serialize Dot turns and durable work admission, and poll subscription readiness. - Serialize mailbox writes and cursor reads; retain paused fence acknowledgement without task authority. - Authorize admitted review runs without changing the worker assignee. Include the fenced assignment ID in production stop notices. ## Verification - This PR integrates master `4a8178e9c`. Dot migration `0317_messy_famine.sql` follows the published history and is safe to reapply. The merge preserves the reserved migration connection, batch-commit handling, private task checks, task monitors, and native accounting. - Local workspace typecheck, full build, and UI token gates pass. The server typecheck passes after the review fixes. Database and native executor regressions pass. - All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot driver tests pass. The tests cover native document writing and finalization, durable replay, queue admission, mailbox ordering, admitted reviews, stale authority, production stop references, and paused acknowledgements. - Current head `d0e7e0626` passes all 57 checks: 53 pass and four are intentionally skipped. This includes full typecheck, build, tests, Rust Runner verification, browser E2E, release verification, and Canary Dry Run. Greptile rates this exact head 5/5. All review threads are resolved. - The full local root test run is slower than the sharded CI run and has not completed. The full CI test gates pass on the current commit. Focused local regressions pass. - Real-account pairing and event delivery on this base commit remain unqualified. Live account and setup proof are recorded in the follow-up #15414. The following screenshots use synthetic preview data. They show the production pairing component and do not qualify a real account or the full agent setup journey.   ## Risks - This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP and Paperclip Runner. The separate experimental-settings follow-up in #15414 replaces this environment flag with saved operator settings. - Dot does not expose provider token usage or cost. The operator must acknowledge external billing. Known Paperclip budget gates still apply. - Cancellation fences Paperclip authority. It does not confirm that Dot stopped all external activity. - Assigned skill files and third-party MCP bindings are unsupported and reject admission. There is no mounted workspace, model selector, or provider thread identifier. - Hosted agent-broker and remote controller deployments are not qualified. - The new migration follows the merged master history. Existing prototype databases still need the normal master migration history before this Dot-only migration. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window size are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, and test inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.0 |
||
|
|
4a8178e9cf |
fix(ui): stabilize composer model selection and refine effort slider (#15444)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task composers let users select an agent, a model, and an effort level. > - The model label changes width when users move the effort slider. > - This movement makes the control harder to use. > - This pull request keeps the label fixed while the picker is open. > - A larger slider gives clearer feedback during selection. ## Linked Issues or Issue Description **What existing behavior does this improve?** The shared model and effort picker in task composers. **Current behavior** The trigger shows each effort change while the picker is open. Different label lengths move the control. The slider is small. **Proposed behavior** The trigger shows “Select model” while any picker view is open. It shows the selected model and effort after the picker closes. The slider has a thicker track and a larger thumb with hover and press feedback. **Reason and benefit** Users can change effort without moving the trigger. The larger thumb gives clearer pointer feedback. ## What Changed - Keep the trigger text fixed while the picker is open. - Add slider size and feedback tokens. - Remove the gray input outline and thumb border. Show keyboard focus on the thumb. - Keep a thumb focus outline when Windows high contrast suppresses shadows. - Keep native keyboard and pointer controls. Respect reduced motion. - Test open and closed labels on desktop and mobile. ## Verification - Picker tests: 27 pass. Composer settings and new-task tests: 86 pass. - UI typecheck and UI build passed before the final CSS-only high-contrast fix. - Token gates pass. - Full typecheck and build stop at installed Runner API type mismatches outside these files. - Revised Chromium checks pass for desktop pointer drag, no input outline during drag, stable open label, keyboard input, and mobile rendering. The earlier rendered-image check confirmed reduced-motion behavior. - The earlier full local suite did not finish. The approved visual revision passed 113 focused tests. The final high-contrast fix passed 27 picker tests and token gates. - Chromium high-contrast and reduced-motion verification shows a visible thumb focus marker and working keyboard input. - Open the model picker. Change effort. Confirm “Select model” stays visible. Close the picker. Confirm the selected model and effort appear. - All CI gates pass on final head `7a64104e14251a354b06b7916d09d44692557c4f`, including full typecheck, tests, build, and browser E2E. Greptile: 5/5 with no findings. ## Risks - Low risk: shared UI only. Native range controls stay in place. - Browser thumb rendering can differ between engines. ## Model Used - OpenAI GPT-6 Codex. Code editing, tool use, and test execution. The runtime does not expose an exact model variant or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.16 |
||
|
|
378e6d95e1 |
chore(lockfile): refresh pnpm-lock.yaml (#15488)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
a7a244ab33 |
feat: add company decision models with permission and cost controls (#15473)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Optional product features need small, typed model decisions. > - Each company needs to choose which shared API connection pays for those decisions. > - Calls must retain user and task permissions, budget limits, and cost attribution. > - This pull request adds a managed decision service, setup UI, and request history. > - Features can check availability cheaply and keep their existing behavior when decisions are unavailable. ## Linked Issues or Issue Description **Subsystem affected** Server services, shared contracts, database accounting, Company Settings, and Costs. **Problem or motivation** Paperclip has no common decision-model service. Adding provider calls within each feature would duplicate credential access, permission checks, and billing rules. **Proposed solution** Let a connection manager configure one company decision model. Support OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test and metadata-only history. Default company-sponsored background decisions to on during setup, and preserve a saved off setting. **Alternatives considered** Per-feature credentials would duplicate existing connection management. Personal overrides and provider fallback chains add permission and billing complexity; they remain deferred. **Roadmap alignment** Reviewed ROADMAP.md and searched open PRs. This extends existing connection access and budget accounting. Product features that call the service remain outside this change. No matching decision-model service PR was found. ## What Changed - Add company settings, an internal `decisionModelService`, local availability checks, and trusted human, agent/run, and system contexts. - Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and ordered-score decisions for both providers. Bound requests and time; disable paid retries. - Add durable invocation metadata and agentless decision ledger charges. Preserve fractional cents, pricing evidence, dispatch identity, and unresolved billing holds. - Share the company accounting lock and apply company, agent, and project budgets. Settle charges once, retain unknown holds, and recover interrupted calls without resubmission. - Reuse connection setup and management UI. Add a Decisions view under Costs, production-component Storybook coverage, database migration, and service documentation. ## Verification - Passed 166 current-code tests covering the decision service/provider, setup component, Costs, OpenAPI, and every failure from the earlier broad run. Coverage includes native SDK wire formats, refusals, billed malformed responses, permission and secret-rotation races, identity changes, concurrent budget admission, unresolved holds, agent/task deletion, and stale setup feedback. - Passed 170 existing connection, cost, budget, heartbeat-accounting, and profile regression tests. - Passed repository typecheck, production build, and design token gates after integrating master. Verified the generated migration on a fresh test database and upgraded the populated preview database from the branch's earlier migration without losing settings or usage. - Ran the required full `pnpm test:run`: its general phase completed with 16,354 passed and 10 failures across five files while this branch was still being updated. Every reported failure passes in the current-code rerun; the serialized phase did not run after that failure. The full GitHub CI suite passed on `b1b856a88`: general and serialized tests, browser shards, runner checks, typecheck, production build, packaging/canary, and policy gates. [CI evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360). - Passed the full-shell Storybook setup-to-history interaction test again after integrating master. Greptile rates the final revision 5/5 with zero unresolved threads. - Walked through the running app: empty setup, add each provider, save, reload, run all three sample questions, inspect fractional charges in history, switch provider while sponsorship is off, disable, and reconnect. Checked mobile settings. These tests used the actual UI, vault, server, SDKs, and database with simulated upstream responses. - Live paid setup tests remain unverified: this environment has no authorized OpenAI/OpenRouter credentials available. No mocked test is presented as live provider evidence. Reviewer journey: Company Settings → General → Decision model. Add/select a shared API connection, save, run the billed sample, open View usage, then disable decisions and verify Run test is disabled after reload. ## Risks - The SDK decision interface is experimental. Pinned versions and wire-format tests limit upgrade drift. - The migration allows agentless service charges and reservations. Existing agent cost-reporting APIs still require an agent, and decision receipts stay separate from run reconciliation. - Timeouts can have unknown provider charges. Holds remain until an audited accounting correction resolves them. - OpenAI prices use a versioned Decisions rate snapshot; OpenRouter costs use provider receipts. Unknown pricing is retained as unknown. - A configured company authorizes background spending by default. Setup explains this, and managers can turn it off. - Live provider account/model availability still needs the two credentialed acceptance checks. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser tools. The exact serving revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.15 |
||
|
|
f1c44b7b56 |
Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs. Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
cd40ebb95b |
Preserve ACP bridge terminal evidence before disconnect (#15478)
Drain terminal frames through backpressure and record bounded terminal evidence without waiting for log persistence. Guard later events and input failures against writes after socket end. Validated with 217 focused tests, full typecheck/build, independent review and green CI with Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.14 |
||
|
|
50b1f95e79 |
fix(tool-gateway): keep MCP connection healthy on oversized/malformed responses (#15462)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents call remote MCP tools through the Paperclip tool gateway > - The gateway keeps a health status for each remote MCP connection, and it hides the tools of a connection that has `healthStatus = "error"` > - One oversized, malformed, or invalid-JSON `tools/call` reply set the whole connection to `error` > - Health goes back to `ok` only after a successful call, but the hidden tools prevent that call, so the connection stayed locked until a person reconnected it > - This pull request makes these reply errors fail only the one call, and it starts a new MCP session for the caller on the next call > - The benefit is that one large or bad reply no longer disconnects a working connection for every agent ## Linked Issues or Issue Description No public issue exists. Related open PRs fix other health downgrades in the same function. They do not overlap with this change: - Refs #11910 (timeouts and JSON-RPC errors) - Refs #15325 (transport errors) **What happened?** An agent called a remote MCP tool that returned more than `MAX_REMOTE_MCP_RESPONSE_BYTES` (1 MB). The gateway returned 502 `mcp_remote_response_too_large`. It also set the connection to `healthStatus = "error"`. Then `connectedMcpConnectionFilter` hid all tools of the connection (except `per_user` connections), and `connections_search` showed the connection as `needs_user_action`. The `malformed_response` and `invalid_json` errors from `readMcpHttpResponse` did the same thing. `readMcpHttpResponse` can also cancel the reply stream before the end. The remote server can then close its MCP session. The gateway kept the cached `mcp-session-id` for up to 30 minutes and cleared it only after a 404. The next calls then failed with "Remote MCP session expired". **Expected behavior** A reply that is too large or not correct fails only that call. The connection stays healthy, and its tools stay visible. The next call uses a new MCP session. **Steps to reproduce** 1. Add a remote MCP connection with `mcpSessionRequired: true`. 2. Call a tool that returns a reply larger than 1 MB. 3. Look at the connection: `healthStatus` is `error`. 4. Start a new gateway session: the tools of the connection are not in the list. **Paperclip version or commit** `master` at `99a9de9940bf5974352d9dbfbb2f21e62e89689f` **Deployment mode** All modes. The fault is in the server tool gateway. ## What Changed - `server/src/services/tool-gateway.ts`: `too_large`, `malformed_response`, and `invalid_json` from `readMcpHttpResponse` now fail only the call. The caller gets the same 502 reason code as before. The gateway does not call `markRemoteConnectionHealth(…, "error")` for these errors. The `invalid_json` branch after `JSON.parse(body)` also does not change health now, so all `invalid_json` paths are the same. - The `mcp_remote_response_too_large` message now tells the caller to request a smaller result, for example a narrower query or a smaller page size. The error details now include `maxBytes`. - `server/src/services/mcp-http.ts`: new `forgetMcpHttpSession()`. It removes only the cached session for one scope and one credential set, and only while that entry still holds the session ID that failed. It does not touch other agents, sessions that are initializing, or a newer session that replaced the failed one. - The gateway calls `forgetMcpHttpSession()` after these reply errors, so the next call from that caller initializes a new session. The existing 404 "session expired" path now uses the same function. Before, both paths cleared all sessions and all pending initializations on the connection. That made a concurrent initialization by another agent fail with "MCP connection changed while initializing", which the gateway reported as a fetch failure and marked as a connection `error`. - `forgetMcpHttpSessions(connectionId)` is not changed. Disconnect and revocation still clear the full connection. Why `malformed_response` and `invalid_json` are also per-call errors: - Each error is about one reply body. The server was reachable, accepted the credentials, and sent HTTP 2xx. A reply can be too large or bad because of the tool and its arguments. That is not a fault of the connection. - The other malformed-reply checks in the same function (payload is not an object, or has no `result`) already throw `remote_mcp_malformed_response` and do not change health. Only the reader-level errors changed health. - Health recovers only after a successful call. An `error` status for one bad reply therefore locks the connection until a person reconnects it. - Other failures (HTTP errors, fetch failures, timeouts, JSON-RPC errors) still change health. This PR does not change them. #11910 and #15325 address some of them. ## Verification - `server/src/__tests__/tool-gateway.test.ts`: the recovery test now runs for three replies: oversized, invalid JSON, and a reply without the requested message ID. Each run uses a fake HTTP MCP server with `mcpSessionRequired: true` that gives `session-N` for each `initialize`. Each run checks that: - the call fails with the correct 502 reason code (and, for the oversized reply, the smaller-page hint and `maxBytes: 1000000`) - `healthStatus` stays `ok` - a new gateway session still lists the tool and can call it - the second `tools/call` uses `session-2`, not `session-1` - New test: agent A gets an oversized reply while agent B initializes a session on the same connection. Agent B's call completes, health stays `ok`, and the two calls use `session-1` and `session-2`. With the old connection-wide reset, this test fails: agent B gets 502 `mcp_remote_fetch_failed`. - `server/src/__tests__/remote-mcp-protocol.test.ts`: new unit test. `forgetMcpHttpSession()` keeps the session of a different identity, and a late failure from an old session does not remove the newer session. - Each regression check fails without its fix. Without the health change, the test fails on `healthStatus: 'error'`. Without the session reset, it fails with `['session-1', 'session-1']`. ``` cd server npx vitest run src/__tests__/tool-gateway.test.ts src/__tests__/tool-gateway-service.test.ts src/__tests__/remote-mcp-protocol.test.ts Test Files 3 passed (3) Tests 132 passed (132) ``` - I did not run the full server `tsc --noEmit` locally because the sandbox does not have sufficient memory. The CI typecheck covers it. ## Risks - Low risk. The change affects only the error path of remote MCP `tools/call`. - A remote server that always sends bad replies now keeps `healthStatus = "ok"`. Each call still fails with a clear 502 reason code and an audit record, so the failure stays visible. Only the connection-wide hiding of tools stops. - After one of these errors, the next call from the same caller sends one more `initialize` request. This adds one round trip. - A 404 "session expired" now clears only the session of the caller that got the 404. Before, it cleared the cached sessions of all agents on the connection. If the remote server restarts, each agent now gets its own "session expired" error one time and then initializes again. A 404 is about one session, and the old connection-wide clear also cancelled other agents' initializations. - #11910 and #15325 change the same catch block. The PR that merges last can have a small merge conflict. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, 1M-token context window. - Run as an agent in Claude Code with tool use (shell, file edit, and local test runs). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing documentation changes are necessary) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge This PR replaces #15418. It keeps the same commits on a branch name without an internal ticket ID, and adds a fix for the Greptile review on #15418. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
fd8c6b920a |
fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native tasks can pause while a human decides whether to allow a connection action. > - The next turn needs both the saved decision and complete accounting for the previous turn. > - Cancellation could discard final usage, and complete direct Claude API receipts could remain unpriced. > - A stale blocked or review report could also request approval again after the original card was declined. > - This pull request retains shutdown accounting and rejects approval waits bound to an already resolved action. > - The benefit is reliable continuation with the existing budget and approval controls. ## Linked Issues or Issue Description Refs #15420. Related: #15312 addresses requester ownership during dispatch. This change addresses receipt capture and final-response validation. ## What Changed - Retain usage events after native cancellation. Continue to reject late provider messages and work. - Drain same-turn accounting and terminal events for at most fifteen seconds after a durable governed wait. Keep incomplete accounting blocked. - Estimate complete, unpriced, direct Anthropic API receipts for the exact `claude-sonnet-5` model. Record the rate version and assumptions. Use the one-hour cache-write rate when the receipt lacks cache TTL. - Bind stale approval reports to exact interaction, action-request, or invocation IDs in the same company, task, agent, and run. Cover blocked, review, and response-wake reports. Keep independent reviews valid. - Fence checkpoint and result writes after a controller detaches for restart, including operations waiting for a database lock. Reject stale successful returns before certifying accounting. - Allow bounded subscription teardown only after retaining an actual provider terminal. - Journal the exact governed-wait trigger and disposition before provider interruption. Recover that wait independently of a later saved answer, replay retained accounting, and reject mismatched or unproven terminal evidence. - Give settling governed turns a bounded window before shutdown detaches their controller. - Keep fuzzy external app matches alongside installed capability matches instead of forcing an unrelated provider question for a generic query. - Return up to twenty exact active catalog tool names after an invalid request, after eligibility checks; still reject the request without granting access or creating an approval. - Require retained provider terminal proof before settling a governed wait, including when complete usage arrives before stream closure/error/timeout. Retain harmless numbered cancellation events so restart replay stays contiguous. - Isolate accounting-test OpenCode config from the host plugin directory. - Add regression coverage and document the accounting, restart and connection-search behavior. ## Verification Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`; later review fixes are verified separately and do not relabel those runs. - Repository typecheck and build pass. - Current runner runtime and cancellation suites: 191 tests pass. Six new regressions cover stream end/error/timeout without provider-stop proof and contiguous cancellation acknowledgement/request replay; all six failed before the fix. The existing bounded cleanup case now explicitly supplies terminal proof. All 31 adapter accounting tests pass with isolated fixture config. - New checkpoint-rebinding and approval-criterion suites: 65 tests pass; runner HTTP integration: 29 tests pass. Unchanged executor/control-plane suites: 640 tests pass; database-backed connection suites: 68 tests pass. - Current retained evidence verification covers 222 file hashes across all fifteen original result artifacts. The three-case [published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html) and all eight screenshot hashes verify. The [three-case recovery report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html) and all six screenshots also verify; the nine-result campaign did not publish. - Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node tests pass. Eval typecheck and catalog discovery pass. - Full local repository unit run was interrupted before the follow-up edits after three tool-access failures and one runner HTTP failure. Those failures pass in isolation; the 09a full run was interrupted after one rapid Slack callback-ordering failure and seven skill-service failures. All eight pass both isolated and with full-runner environment settings, and all 74 skill-service tests pass together; the subsequent full run reported two 15-second OpenCode accounting timeouts and was stopped with exit 130 to apply review fixes. The timeouts reproduce while copying this host’s 61 MB OpenCode config. All 31 tests pass after isolating config inside each fixture without increasing timeouts or changing assertions. A complete local full-suite pass is not claimed. Repository-wide CI also passes on the final review-fix head in [run 37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952). Earlier database-skipped diagnostics and the older ENFILE run are retained and are not full-suite passing evidence. - Three-case live campaign [37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525) passes all three original grades on frozen source 09a: 55/55 checks, eight succeeded run records, complete accounting receipts, matching checkpoint identities and no pending approvals. Campaign [37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353) adds nine original passes (173/173 checks, eighteen succeeded records) on the identical source. Its other three jobs failed before runner assignment or any step while GitHub could not load the paid environment; those original infrastructure failures are retained. Campaign [37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356) completes only those unstarted cells: all three original grades pass (57/57 checks, six succeeded run records). All fifteen exact cases now pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new overall failures and zero pending pairs. Total current evidence: 285/285 checks and thirty-two succeeded run records, complete accounting receipts, matching checkpoint identities, no pending approvals or retry records. Earlier campaigns retain forty-seven additional run records and two known same-run recovery attempts; actual provider-call counts and invoices remain unknown. The nine-result campaign skipped publication and its public URL returns 403; original artifacts remain retained. Previous ba9 campaign [37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840) completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline pass, while OpenCode resumes but times out searching for exact tool names and never creates the access card. That failure and incomplete cancelled-run accounting remain preserved; the new catalog error guidance targets this observed dead end. Campaign [37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312) remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker leak and unknown-criterion approval gap. The original baseline remains 7 PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS / 3 FAIL. All fifteen selected cases are qualified by their original grades in this bounded trial. These live grades belong to 09a. Its twelve governed-wait checkpoints retain matching same-turn terminal fingerprints, but passing artifacts omit detailed event journals; the later six adversarial regressions qualify the new terminal-proof and replay guards separately. Final-head repository CI passes. [Fresh Greptile review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780) is 5/5, confirms both findings are fixed, and reports no new actionable issues. All review threads are resolved and the PR has no merge conflicts. ## Risks - Governed cancellation drains accounting for up to fifteen seconds. Restart detachment gives a settling batch up to twenty seconds to finish. An incomplete receipt or unproven provider terminal still prevents successful qualification. - Claude prices are estimates, not invoices. The estimate assumes standard global API pricing and uses a conservative cache-write rate. Unsupported models, billers, and billing modes remain unpriced. - Approval identity matching must remain scoped to the current run and the requested approval. It does not authorize execution of a declined call. - Original baseline and final candidate have different merged master context. Exact-case outcomes are before/after observations, not isolated causal attribution to this repair. - No schema migration, fixture, oracle or grader change. Search-result guidance now treats fuzzy external matches as suggestions. Existing app authorization and provider-consent checks remain required. ## Model Used - OpenAI GPT-6 through Codex, with code editing, terminal tools, and test execution. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (191 current runtime tests and 31 accounting tests; interrupted full-suite history and CI coverage are disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5405b1f46f |
Synchronize chat delivery tests with completed drains (#15480)
Expose the already tracked and caught deferred-work promise to test schedulers while preserving default setImmediate behavior. Assert Slack delivery ordering while the primary lease is held, then join the drains; update the existing GitHub test to assert the resulting quiescent queue. Verified all 1,063 chat tests, full typecheck/build, negative synchronization control, independent review and exact-head CI. An unchanged browser shard retry passed and is documented. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0514e8abd5 |
Keep Cursor managed-runtime test installation offline (#15484)
Intercept the Cursor installer in its local shell fixture and guard against accidental curl downloads. Preserve real managed-home archive operations and the default timeout; change no production code. Verified 46 related tests, safe negative mutation, independent review and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
572cec0344 |
fix(ui): present missing costs as an informational notice (#15459)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Costs dashboard shows spending and gaps in cost data. > - Historical usage can have token counts without a recorded price. > - The red notice implies that the user must fix those records. > - This pull request shows the missing costs as a neutral note below the totals. > - Users can see the limit of the totals without a false call to action. ## Linked Issues or Issue Description Refs #14997. **What happened?** The Costs dashboard shows a red notice when a selected period contains usage without prices. It also shows “0 runs await accounting” when no runs are pending. Historical pricing gaps can therefore look like current errors that need cleanup. **Expected behavior** Show a neutral explanation below the totals. State that totals include only known costs. Show pending runs separately and only when the count is positive. **Steps to reproduce** 1. Open Costs for a period with unpriced usage and no pending runs. 2. Observe the red notice above the totals. 3. Apply this change. Confirm that the notice is muted and below the totals, with no zero-count pending message. **Paperclip version or commit** Reproduced on master after #14997. **Deployment mode** Board UI in both the embedded Activity page and the standalone Costs page. ## What Changed - Move the cost-data notice below the summary tiles and use the muted text token. - Explain unavailable prices with “Totals include known costs only.” - Separate pending runs from missing prices. Omit zero counts and use singular or plural copy as needed. - Update the existing UI tests for both page variants, including missing-only, pending-only, and fully accounted states after refresh. ## Verification - `pnpm --dir ui exec vitest run src/pages/Costs.test.tsx`: 21 tests passed on the final commit. Both page variants cover mixed, missing-only, pending-only, and fully accounted states. - `pnpm check:token-gates`: passed on the final commit. - `pnpm build`: passed locally. - `pnpm -r typecheck`: passed locally. - Full Linux CI passed on `0c23544c0f17e6641cc0c8de6ed498a731748bde`, including browser tests, server tests, shared-package tests, build, and typecheck. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37661883464). - Greptile Apex: 5/5 on the final commit, with no unresolved review threads. - Local full-suite limit: `pnpm test:run` first failed because the fresh checkout lacked embedded PostgreSQL library links. A runtime-skill fixture also selected an unrelated parent directory, and a load test timed out. After setup repair, the database suite (70 tests), skill suite (3 tests), email suite (39 tests), and load suite (4 tests) passed individually. The subsequent full local rerun was stopped after the complete Linux CI run passed. This is not a claim that a full local test invocation passed. ## Risks Low risk. This changes presentation only. The note remains visible and retains its accessible status role. Cost calculations, stored records, and budget enforcement do not change. The copy does not assume that all missing prices are historical. No schema migration or operator documentation change is needed. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code editing, and test execution. The exact served model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
465140596f |
Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials. Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
128c95837b |
Retain bounded orphan process-loss diagnostics (#15475)
Record bounded observer uptime, run and output ages, process-check observations and retry eligibility before orphan cleanup changes the evidence. Preserve existing recovery and reporting behavior and omit process IDs, raw paths and credentials. Verified focused helper, actual reaper and real SDK regressions, full typecheck/build, exact-head CI and independent review. An unchanged Cursor timeout passed its isolated retry; unrelated local Slack timing failure is documented separately. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
43b0aff54b |
Keep pre-dispatch secret configuration blockers out of Sentry (#15472)
Keep confirmed pre-dispatch missing-secret setup blockers out of Sentry while preserving failed runs and owner recovery actions. Runtime, provider and ambiguous failures remain reportable. Verified full exact-head CI, focused local regressions, workspace typecheck and independent review; Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ae6f95ed7a |
feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults. Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
06484b3c41 |
fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Recovery schedules bounded retries after a provider disconnects. > - A stopped run can still hold its environment lease while cleanup runs. > - Retrying before that lease is released cancels the new run before it starts and spends another retry. > - Restoring the task can also send the worker repair instructions from an already resolved recovery action. > - This pull request preserves the waiting retry and removes settled recovery instructions from later wakes. > - The task can continue after cleanup without an operator repairing the same incident again. ## Linked Issues or Issue Description **What happened?** A legacy conversation run disconnected while it was doing ordinary work. Its environment cleanup took longer than the retry delay. Two retries were cancelled before dispatch with `execution_reconciliation_required`. Those cancellations exhausted the failure budget. After an operator restored the task, the wake still told the original worker to repair the runtime and hand the task back to itself. **Expected behavior** Cleanup waits preserve the pending attempt. The same retry can continue after ownership is released, subject to all current gates. Once a recovery action is resolved or cancelled, subsequent task wakes omit its repair instructions. **Steps to reproduce** 1. Fail a legacy conversation run while its environment lease remains in `pending_cleanup`. 2. Schedule a bounded retry and run promotion before cleanup releases that lease. 3. Repeat the scheduler sweep. Before this fix, retries promote and then cancel without starting. 4. Resolve a stranded-task recovery action and build the restored task wake with that action ID. Before this fix, the wake still includes the settled repair instructions. **Paperclip version or commit** Reproduced with database regressions against `ceabc3bc880` on master. Related: #15019 restores a skipped assignment handoff after lease release. #15235 filters stale handoff evidence in the recovery sweep. This change preserves an existing scheduled conversation retry and corrects restored wake content. It does not create a new handoff wake. ## What Changed - Keep an unstarted legacy conversation retry on the same durable row while prior execution ownership remains active. Recheck after 30 seconds without increasing retry accounting. - Return a queued retry to scheduled state if it encounters that hold at the claim gate. Retain its issue claim and publish the status change. - Record one local lifecycle diagnostic per blocking run. Remove that wait marker on promotion. - Include recovery action metadata only while the referenced action is active or escalated. - Return an explicit `waiting` response and the saved schedule when Retry now meets cleanup. Show the wait inline without a false success or disabled button. - Add database, rendered-prompt, route, and UI regressions. Document the execution and run-log contracts. ## Verification - Five cleanup and restored-wake regressions fail against the original production code. The Retry now route and UI regressions also fail before their correction. - Related retry, dispatch, stale-queue, and recovery suites: 321 tests pass across seven files. All ten focused cleanup/restored-wake cases pass after rebase. The final dispatch adjustment passes all 46 adapter tests. - Retry now routes and affected UI suites: all 46 tests pass. The tests cover repeated clicks, the saved schedule, unchanged accounting, promotion after release, and no false success or error state. - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm check:token-gates` pass. - Full local `pnpm test:run` was started and then stopped after the final commit passed all GitHub CI test shards. No complete local full-suite result is claimed; CI supplies the complete test result for the final commit. - Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub checks are green or intentionally skipped. Apex review is 5/5 after two reviews, with no unresolved threads. The PR has no merge conflicts. ## Risks The wait applies only to unstarted legacy conversation retries. Native runs and non-conversation execution keep their existing recovery rules. Cleanup must actually release ownership before execution can resume. The wait does not fix a cleanup service that never finishes. Promotion and dispatch still enforce cancellation, reassignment, pause, budget, and reconciliation gates. No schema change is required. The Retry now response adds a `waiting` outcome; the shared contract and all three UI controls handle it. ## Model Used OpenAI Codex based on GPT-6, with repository analysis, tool use, and local code execution. The runtime does not expose the precise serving model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; complete final-head test coverage in CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db0bee19fa |
docs(readme): add Superagent badge and move Star History rank badge to its own line (#15460)
- Add the Superagent security badge to the README badge row, with light and dark variants via <picture>. - Move the Star History global rank badge into its own centered paragraph below the badge row, with spacing. Co-Authored-By: Amy (Opus) via Paperclip <claude@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
9fb955e9a2 |
fix(codex): gate models on the Codex CLI floor before the ChatGPT backend rejects them (#15399)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Codex adapter and the native runner run Codex in local hosts and in managed sandboxes, and the model catalog lists the models an operator can select > - The Codex backend accepts a new model with ChatGPT sign-in only from a recent enough Codex CLI. An older CLI fails every turn with `The '<model>' model is not supported when using Codex with a ChatGPT account.` > - After #14942 added `gpt-6.1-sol`, an operator selected it for an agent in a managed Daytona sandbox. The sandbox image still shipped Codex 0.156.0. The connection test and runs failed with that sentence, which reads like an account problem > - Paperclip only checked the Codex compatibility window (`>=0.149.0 <0.161.0`), so 0.156.0 passed and the failure surfaced from the backend without a cause > - This pull request records the verified Codex CLI floor for each gated model, compares the installed `codex --version` with that floor in the environment Test and in the remote runner, and names the backend rejection when it still happens > - The benefit is a precise, actionable message ("gpt-6.1-sol requires Codex CLI 0.159.0 or newer; detected 0.156.0; promote a sandbox image with Codex 0.160.0") instead of an opaque 400, and no doomed hello probe ## Linked Issues or Issue Description Refs #14942 **Bug: selecting `gpt-6.1-sol` in a managed sandbox fails with a ChatGPT account error.** **Steps to reproduce** 1. Promote a sandbox image that ships Codex CLI 0.156.0. 2. Create a Codex agent with a ChatGPT sign-in connection, select `gpt-6.1-sol`, and run the environment Test or a task in that sandbox. **Expected behavior** The Test or the run tells the operator that the Codex CLI in the sandbox is older than the model needs. **Actual behavior** The Test reports `codex_hello_probe_failed` with `{"type":"error","status":400,"error":{"type":"invalid_request_error","message":"The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account."}}`. A run fails with "The selected model is not supported by the current ChatGPT connection." **Evidence** - Codex 0.156.1 and older are rejected for `gpt-6.1-sol` with ChatGPT sign-in: https://github.com/openai/codex/issues/49396 - Codex 0.159.0 and 0.159.2 are accepted: https://github.com/openai/codex/issues/49464 and https://github.com/decolua/9router/issues/4471 (same error, fixed by raising the client version identity from 0.155.0 to 0.159.0) - Codex 0.157.0 release notes: "Add GPT-6 Sol and Luna to the model catalog": https://github.com/openai/codex/releases/tag/rust-v0.157.0 - GPT-6.1 Sol is included for Plus, Pro, Business, Enterprise, and Edu with ChatGPT sign-in: https://learn.chatgpt.com/docs/models ## What Changed - `packages/adapters/codex-local/src/index.ts`: add `minimumCodexCliVersionForModel` (`gpt-6.1-sol` → 0.159.0, `gpt-6-sol` and `gpt-6-luna` → 0.157.0, with the sources above), `parseCodexCliVersionOutput`, `codexCliVersionAtLeast`, and `CODEX_CHATGPT_MODEL_REJECTION_RE`. Models without a verified floor return `null`, so nothing probes them. The agent configuration doc describes the floors. - `packages/adapters/codex-local/src/server/cli-version.ts` (new): run `codex --version` where the run would execute it (local, SSH, or sandbox), compare it with the model floor, and build the `codex_cli_version_compatible` / `codex_cli_version_incompatible` checks. Map the backend rejection sentence to `codex_hello_probe_model_rejected` with the detected CLI version and a hint that separates a stale CLI from a plan that does not include the model. - `packages/adapters/codex-local/src/server/test.ts` (CLI lane Test): run the version check before the hello probe when the model has a floor. Skip the hello probe when the CLI is too old (`codex_hello_probe_skipped_cli_version`). When the probe still fails with the backend sentence, report `codex_hello_probe_model_rejected` instead of the generic `codex_hello_probe_failed`. The new codes do not match the auth-failure patterns, so a managed connection is not invalidated. - `packages/adapters/codex-local/src/server/acp.ts` (ACP lane Test): for remote targets, run the same version check against the shared `codex` the ACP server spawns. - `server/src/services/native-runtime/native-session-executor.ts`: after the compatibility-window check, compare the remote Codex with the configured model's floor and fail with `runner_remote_provider_artifact_incompatible: <model> requires Codex <floor> or newer with ChatGPT sign-in, received <version> from the sandbox image; promote a sandbox image with Codex 0.160.0 or configure PAPERCLIP_RUNNER_REMOTE_CODEX_NPM_SPEC=...`. When a preinstalled Codex fails this check and an npm spec is configured, the existing fallback installs the pinned release. - Tests: adapter metadata, CLI-lane Test (too old, compatible, no floor, backend rejection), ACP-lane Test (too old, compatible, no floor), and the remote runner floor (nine version/model cases). - `doc/adapter-model-audit-2026-10-02.md`: record the floors and the reason. No catalog entry, default model, saved agent configuration, pin, lockfile, or image definition changes. ## Verification Local (Node 25.9, pnpm 9.15.4 via corepack): ```sh pnpm --filter @paperclipai/adapter-codex-local typecheck pnpm --filter @paperclipai/adapter-codex-local exec vitest run # 81 passed pnpm --filter @paperclipai/server exec vitest run \ src/services/native-runtime/native-session-executor.test.ts \ src/services/native-runtime/codex-runtime-compatibility.test.ts -t Codex # 84 passed ``` Also run locally: the full `native-session-executor.test.ts` file (533 passed) and the full Codex adapter suite (81 passed). Follow-up commit `9fd4d3b25` (Greptile P2: the ACP-lane version probe ignored the agent's configured env): the probe now receives the adapter's string-valued `env` entries, the same ones `buildCodexAcpConfig` hands to remote ACP runs, so a `PATH` override selects the same `codex` for the Test as for the run. New test covers a `PATH` + `CODEX_HOME` override and a dropped non-string entry. Re-run on that head: `pnpm --filter @paperclipai/adapter-codex-local typecheck` clean; full Codex adapter suite 503 passed (31 files). Not run here: `pnpm -r typecheck` for `server` (the direct `tsc --noEmit` was killed by the sandbox memory cap; the adapter package typecheck passes and CI covers the server), `pnpm build`, browser suites, and a live sandbox probe. No UI files changed, so token gates do not apply. Manual check for a reviewer: set an agent's Codex model to `gpt-6.1-sol`, point its environment at a sandbox whose `codex --version` prints `codex-cli 0.156.0`, and run the environment Test. The result contains `codex_cli_version_incompatible` with `Detected Codex CLI 0.156.0.` and no hello probe check. With Codex 0.160.0 the Test contains `codex_cli_version_compatible` and runs the hello probe. ## Risks - A floor is enforced for all authentication modes, because the runner does not know the credential method at verification time. An OpenAI API key on Codex 0.157.0 or 0.158.0 with `gpt-6.1-sol` is now rejected before launch. Paperclip pins Codex 0.160.0 everywhere, so this only affects installs that lag the pin, and the message names the fix. - The 0.159.0 floor for `gpt-6.1-sol` is the oldest stable release verified to work; 0.157.0 and 0.158.0 were not verified either way. If OpenAI accepts an older client, the floor can be lowered in one table. - The `codex --version` probe adds one short process run to the Test only for models with a floor. - Rollback: revert this pull request. No data or configuration migrates. ## Model Used Claude Fable 5.1 (`claude-fable-5-1`), Anthropic, with extended thinking, tool use, and web research, operating as a Paperclip agent through Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Bender (Fable) <noreply@paperclip.ing> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>canary/v2026.1007.0-canary.11 |
||
|
|
ceabc3bc88 |
fix(e2e): wait on the server's own deferral signal in the signoff wakeup helper (#15345)
## Thinking Path > - Paperclip coordinates work for AI agents through server-managed tasks and heartbeats > - Signoff policy tests exercise stage transitions that wake the next participant > - A stage transition can defer a wake while the previous run still holds the issue execution lock > - The test helper used a fixed wait and could fail before the server released that lock > - This pull request waits on the server deferral signal and verifies the returned run identity > - The benefit is a stable test with clear failure details and no repeated wake request ## Linked Issues or Issue Description **What happened?** The signoff policy end-to-end spec failed intermittently with `No issue-bound heartbeat run became available for agent <id>`. The server had deferred the wake while another run held the issue execution lock. **Expected behavior** The test helper waits for the named blocking run to finish, then reads and verifies the new run for the requested agent and issue. **Steps to reproduce** 1. Run `npx playwright test --config tests/e2e/playwright.config.ts tests/e2e/signoff-policy.spec.ts`. 2. Exercise a signoff stage transition while the previous stage run still holds the issue execution lock. 3. Confirm that the helper waits on the server deferral signal and returns only a matching run. This change relates to [PR #11299](https://github.com/paperclipai/paperclip/pull/11299), which also touches the signoff policy end-to-end spec. ## What Changed - Call the wakeup endpoint that reports the server deferral reason. - Wait for the named blocking run instead of using a fixed delay. - Verify that each returned run belongs to the requested agent and issue. - Bound read and wait loops and report the last server signal and lock fields on failure. - Keep one wakeup request per helper call. ## Verification - `npx playwright test --config tests/e2e/playwright.config.ts tests/e2e/signoff-policy.spec.ts` passed locally. - All five signoff policy cases passed locally. - The spec keeps all 60 assertions. - The diff contains only `tests/e2e/signoff-policy.spec.ts`. - The full CI end-to-end shards must pass before merge. ## Risks - The helper now depends on the wakeup endpoint's deferral fields. - A server change to those fields can fail the test with a named diagnostic. - The change affects end-to-end test behavior only. ## Model Used OpenAI GPT-5 Codex, exact runtime model `gpt-5-codex`, with tool use and code review assistance. The runtime does not expose a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.10 |
||
|
|
f669194298 |
fix(runner): propagate configured environment to future turns (#15451)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runners start provider processes and enforce a separate tool environment policy. > - The server resolves task environment bindings at each run boundary. > - Fixed launch allowlists dropped custom variables after resolution. > - Warm process reuse and a blanket fingerprint exclusion also hid configuration changes. > - This pull request carries a bounded, server-selected list of task variable names through each process boundary. > - The benefit is that later turns and tool commands receive configured values while ambient host secrets stay excluded. ## Linked Issues or Issue Description **What happened?** A native Codex agent could not use configured task credentials. Adding the variables in Settings did not fix the next turn. The provider sanitizer, runner launch, Rust provider launch, and shell policy each used fixed allowlists. Effective config fingerprints also excluded all variables with the `PAPERCLIP_` prefix. **Expected behavior** Explicitly bound task variables must reach provider processes and tool commands. Added, changed, and removed values must take effect at the next run boundary. A running turn keeps its original configuration. **Steps to reproduce** 1. Start a native Codex task without a custom environment binding. 2. Add a fake `PAPERCLIP_PAGE_BUCKET` value and a fake Pages credential binding in Settings. 3. Continue the task and inspect the tool environment. 4. Before this fix, those values are absent even from a fresh runner launch. **Paperclip version or commit** Reproduced at `b31558064`. The fix is rebased on current master. **Deployment mode** Self-hosted server with a native process runner. Related changes: [the legacy Codex MCP environment fix](https://github.com/paperclipai/paperclip/pull/13321) and [ambient server-secret exclusion](https://github.com/paperclipai/paperclip/pull/12870). These affect different launch paths. This change preserves their credential boundaries. ## What Changed - Capture scoped task bindings after resolution, before managed provider credential injection. Mint and validate the names-only projection at native dispatch before host inheritance. Legacy adapters retain their previous environment limits. - Strip user-supplied projection markers from agent, environment, project, and routine config. - Carry selected values through the Codex, ACPX, OpenCode, runnerd, and Rust subprocess launch boundaries. - Add selected names to native and ACPX Codex shell include lists. Keep selected values out of command arguments, including selected bootstrap values. - Reject malformed projections, reserved authority and loader names, missing values, null bytes, and oversized input. - Replace a retained native process when projected values change. Keep unchanged processes reusable. - Fingerprint custom namespaced variables while excluding known generated runtime variables. - Document next-run behavior and add regression coverage. ## Verification - Red: the permanent reproduction failed at four launch/tool boundaries and the namespaced fingerprint check. Two control checks passed. - Green: the initial regression plus existing Codex environment and shell tests passed (39 tests). - Server config resolution, fingerprints, and native-session suites passed (653 tests), including addition, rotation, removal, and unchanged warm-session reuse. - Rust regression tests passed. A real shell child received added and rotated values, then lost them after removal. Unselected host variables stayed absent. - `pnpm -r typecheck` and `pnpm build` passed after rebase. The full `pnpm test:run` was attempted but could not complete: fresh embedded PostgreSQL databases fail during bootstrap on this macOS host. An isolated suite and a disposable native `initdb` probe reproduced the failure before test execution. `shmget` reports `No space left on device` because the host has exhausted shared-memory IDs. This is not disk exhaustion. The run was stopped after confirming the external setup failure. CI results will be recorded separately. - Runner boundary suites: 361 tests passed. Two process-launch errors during concurrent binary staging passed on isolated rerun. - Rust Codex provider and process supervisor integration suites: 97 passed, 2 intentionally ignored. - ACPX shell and selected-bootstrap argv regressions: 3 failed before the fix, then all 55 relevant tests passed. - Review compatibility regression: 129-variable and large-value legacy configurations failed before the correction and passed after moving native-only validation to dispatch. - Final review head `3f11e8d25`: repository typechecks and production build passed. CI completed its implementation checks; one general-server shard hit SQL `40P01` in `heartbeat-runtime-skills.test.ts` during its `beforeEach` table truncate (1,199 tests passed in that shard). The shard passed on its single rerun. All CI checks for this head are green. Greptile reviewed this head at 5/5 with no new actionable findings; the compatibility thread is resolved. - No live provider credentials or model calls are required by these tests. ## Risks - Configured task credentials now reach the tools they were configured for. The controller selects names only after existing scope and secret-binding authorization. - TypeScript and Rust validate the same bounded projection. Their reserved-name rules must stay aligned. - A changed projection replaces an idle provider process. Unchanged values preserve reuse. Active turns retain their original environment. - No database migration or API schema change. > This fixes existing runner configuration behavior. It does not add a new roadmap capability. ## Model Used - OpenAI GPT-6 (Codex), with reasoning, local code execution, and repository tools. The session does not expose a more specific API model identifier or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0830737030 |
feat(ui): add CSV table previews to task file tabs (#15445)
## Thinking Path > - Paperclip helps people manage AI agents and inspect their work. > - Task file tabs show the files that agents produce. > - Markdown has a readable view, but CSV files show only source text. > - People need to scan CSV reports without a separate app. > - This pull request adds a table view to attachment and workspace file tabs. > - View buttons let people inspect the original source beside download. ## Linked Issues or Issue Description **Subsystem affected** ui/ — task file tabs. **Problem or motivation** CSV exports are difficult to scan as source text. Related work: https://github.com/paperclipai/paperclip/pull/14297 added text attachment tabs. **Proposed solution** Render CSV as a table by default. Add row numbers, sticky headers, and record counts. Keep rendered and raw icons beside download in task tabs. **Roadmap alignment** This is a small improvement to existing task file previews. It does not add a new core subsystem. ## What Changed - Add a shared CSV parser and table preview. - Handle quoted commas, escaped quotes, UTF-8 BOM, CRLF, and multiline cells. - Show at most 500 data rows and 100 columns. Show a notice for larger files. Keep the complete download. - Add CSV attachment and workspace previews. Put workspace tab view controls in the header. - Add parser and interaction tests, Storybook examples, and product documentation. ## Verification - 28 focused tests pass across the parser, attachment tabs, and file viewer. - UI typecheck, UI build, and token gates pass. - Playwright checked desktop and mobile rendered → raw → rendered flows using the real attachment preview component in Storybook. Screenshots are recorded on the driving task. - Repository-wide typecheck and build were attempted. Both require cargo, which is absent in this environment. - The earlier repository-wide local test run stopped after a server test race failure. The prior commit passed all GitHub CI jobs. - Review fixes add source-line navigation coverage and CSV boundary tests. All 22 file-viewer tests pass. UI typecheck and token gates pass. - All 54 active GitHub checks pass on the final commit. Two Storybook checks are skipped. Greptile gives 5/5 with no actionable findings. - The final version preserves HTML previews from master and uses the shared toolbar control. All 35 focused tests, UI typecheck, UI build, and token gates pass. ## Risks - The first CSV record supplies column headers. Files without headers use that same rule. - The table preview is bounded. Download and raw view retain the complete file. - No API, database, or dependency changes. ## Model Used - OpenAI Codex, GPT-6. The runtime does not expose a more precise model ID or context window. Used code editing, shell tools, and browser checks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1640c5b6ab |
feat(ui): render HTML artifacts in a secure sandbox (#15447)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents deliver reports as task attachments and workspace files. > - The board already renders Markdown, but it shows HTML reports as source or rejects their preview. > - Reports can need inline scripts to build charts and tables. > - Artifact scripts must not read board cookies, storage, or the parent page. > - This pull request adds an opaque-origin HTML preview and view controls beside Download. > - Users can explore a report, inspect its source, and download the original file. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task attachment panel and workspace file viewers. **Subsystem affected** ui/ and the server workspace file preview service. **Current behavior** HTML attachments show source text. Workspace HTML files cannot be previewed. **Proposed behavior** Render self-contained HTML reports in a sandboxed iframe. Place eye and code controls beside Download. Preserve the original file for raw view and download. **Reason and benefit** Users can read and filter agent reports in Paperclip without giving artifact scripts access to the board session. **Breaking changes** HTML previews now open rendered. Scripts and styles must be embedded in the report. Remote resources, API requests, forms, popups, and host navigation are blocked. Workspace HTML content uses the existing bounded UTF-8 JSON response. **Additional context** Related closed proposals: #3293 and #4857. This change uses the current attachment and workspace viewers and adds no report-serving endpoint. The Artifacts and Work Products roadmap item is complete; this improves its existing preview behavior. ## What Changed - Add a shared HTML iframe renderer with `sandbox="allow-scripts"` and no `allow-same-origin`. - Install a restrictive CSP before artifact markup. Keep inline report scripts and styles. - Add shared rendered/raw icon controls beside Download in the attachment panel, workspace panel, and file sheet. - Return workspace HTML as bounded text inside JSON. Keep download and path-access protections. - Add Storybooks for reports, security probes, workspace viewers, a task journey, and mobile layouts. - Document the security boundary and browser test command. ## Verification - `pnpm -r typecheck` and `pnpm build` pass. - Token gates and the static Storybook build pass. - Four browser tests pass. They cover report filters, raw mode, task entry, workspace viewers, cookies, storage, host DOM, resource requests, forms, popups, and host navigation. - The focused UI suite passes 26 tests. The file-resource server suite passes 36 tests. - Local full-stack acceptance passes with an actual 52 KB HTML report. Its chart and filter render. Raw mode preserves the source. Download bytes match the original file. - The complete local UI suite passes: 698 files and 7,754 tests. - `pnpm test:run` was attempted. Its server group recorded two unrelated timeouts and eight connector failures. All failed cases pass in isolated reruns, including the complete 388-test connector suite. The remaining local run was stopped after the full CI test coverage passed, to release test database resources. The full local command did not complete cleanly. - The separate local CLI group passes 511 tests. Three worktree database cases fail to start embedded PostgreSQL because this Mac has exhausted its shared-memory allocation. One isolated rerun fails at the same database startup step. No CLI code was changed. The corresponding CI test jobs pass. - All 56 PR checks are clean on `c771255f1a2e86bd8825f22eafb5d14a5e342cce`. Greptile gives 5/5 with no actionable findings. The security scan passes. The branch has no merge conflict. - Review the **HTML artifacts** section in Storybook. Run `pnpm exec playwright test --config tests/html-preview/playwright.config.ts` for the browser checks. ## Risks - Never add `allow-same-origin` to this iframe. Its opaque origin is the cookie, storage, and host-DOM boundary. - Reports that depend on CDN assets need to embed their dependencies. - The frame can navigate itself. Sandbox restrictions remain after navigation. This is not full network isolation. - Inline scripts can consume browser resources. This change does not isolate CPU or memory use. - No database migrations or API shape changes are required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository editing, code execution, and browser tool use. The runtime does not expose a more specific model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — feature and UI checks pass; the full local run has the infrastructure limits recorded above. All CI test jobs pass. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cfedd7df8 |
fix(runner): restore task monitors and durable timed waits (#15446)
## Thinking Path > - Paperclip manages AI agents and their task execution. > - Agents need a durable way to return to work after a delayed check. > - The issue monitor scheduler already provides a one-shot wake for an assignee. > - Native runners reject generic execution-policy writes and had no bound monitor tool. > - Scheduling alone is insufficient because native completion also needs to accept a timed wait. > - This pull request adds an authorized monitor tool and connects it to completion and the existing scheduler. > - An agent can now schedule its next check, end the run, and resume on the same task. ## Linked Issues or Issue Description **What happened?** A native runner could not set its own task monitor. `call_api` correctly rejected execution-policy writes, while `schedule_wake` had no production binding. `paperclip_finish` also rejected monitor waits. **Expected behavior** A standard native run can set a one-shot monitor on its current task or another accessible task assigned to the same agent. After a confirmed schedule on the current task, it can yield. The scheduler later delivers `issue_monitor_due`. **Steps to reproduce** 1. Start a standard native task. 2. Ask the agent to check the task again later and end its current run. 3. Inspect available tools and try the generic issue execution-policy update. 4. Observe the missing native tool and the lifecycle-write denial. Related PRs: #14680 concerns monitor notes in the shared wake prompt. #11919 changes attempt-limit scope. This PR adds native scheduling and completion authority and retains the existing cumulative attempt bounds. It does not depend on either PR. ## What Changed - Add provider-neutral `set_task_monitor` with a default current-task target, future timestamp, required notes, existing bounds, and explicit clearing. - Check company, task visibility, ownership, runtime permissions, work mode, and active-run authority. Preserve review-only restrictions. Reject the reserved server-owned quota-recovery name before saving or accepting a native wait. - Commit the monitor, audit event, and retry receipt together. Retry receipts survive a successor run without re-arming cleared or consumed timers. - Permit `paperclip_finish` to yield to a persisted monitor. Recheck ownership and the schedule when committing final disposition. Release execution without an immediate continuation. - Preserve due monitors during native execution. Fence wake admission and consumption against replacement, clearing, reassignment, and completion. Preserve unrelated review policy. - Expose scheduled and consumed monitor instructions in task context. Update provider schemas, Rust validation, generated contracts, and execution documentation. - Add an opt-in live Codex smoke script with isolated data and explicit run/session/runner/process evidence. ## Verification - Repository `pnpm -r typecheck` and `pnpm build` passed after rebase. Server typecheck passed again after review fixes. All CI test shards pass on `3def77b1b`, including runner TypeScript/Rust, server, serialized server, workspace, and browser tests. All CI gates are green, including the canary dry run. Greptile is 5/5 on the same commit with zero unresolved threads. - The local monolithic `pnpm test:run`, started before the rebase, was interrupted after current-head CI test coverage passed. It is not counted as a standalone full-suite pass; the focused local regression suites passed. - Targeted server tests cover scheduling, replacement, clearing, policy preservation, cumulative bounds, cross-run retries, permissions, provider-neutral discovery, review restrictions, completion authority, and scheduler/finalizer races. - Runner contract/catalog/semantic tests and Rust terminal-tool tests cover the new operation and monitor completion. - Live Codex test passed twice (latest live run on `e0bcd63e6`) in a temporary database and workspace, with a 300,000 ms warm window. First run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at `2026-10-07T13:15:03.274Z`. Second run `97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`, received `issue_monitor_due`, and completed the same task. Exactly one monitor wake was recorded. - Both live runs used native session `7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner `e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session `01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same process start time. This proves warm reuse for that local Codex test, not only successful scheduling. - Reproduce the paid live test with `node --import ./server/node_modules/tsx/dist/loader.mjs server/scripts/smoke-native-task-monitor.ts --run`, with the installed Codex binary on `PATH` and a valid local login. ## Risks - The scheduler now defers monitor dispatch while the task has an active native run. A stuck run still depends on the existing recovery lifecycle. - Idempotency uses the existing run ledger; no table or migration is added. - Other providers share the tested tool and completion contracts. Only Codex received a live model test. - Existing `call_api` lifecycle restrictions remain enforced. Monitor waits do not bypass task blockers, reviews, or approvals. ## Model Used OpenAI Codex, GPT-6 family, with tool use, code execution, and TypeScript/Rust editing. The session does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.9 |
||
|
|
b315580640 |
feat: add private task creation and sharing controls (#14718)
Add private task creation and sharing controls, effective access explanations, private project management, and locked task references. Constrain the mobile composer plus menu to the viewport and explain named inherited privacy on hover, focus, or tap. Include 125 production-component Storybook stories and thirteen interactive user journeys. Final head passes all CI checks and Greptile Apex 5/5 with no findings. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.8 |
||
|
|
abbd88007f |
fix(hermes): keep managed instructions out of resumed user turns (#15439)
Deliver managed instructions through Hermes's native system overlay while keeping current wake and runtime identity in user turns. Add fresh/resumed regression coverage and include Hermes in the default CI test roster. Fixes #15385 Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3a726e676f |
feat: enforce private task permissions across execution and data (#10633)
Enforce private task and project access across direct reads, search, execution, files, plugins, live delivery, and sharing mutations. Preserve downward-only sharing, current responsible-user authorization, and audited emergency access. Bind historical draft assets with migration 0314. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b67db12d90 |
feat: add durable storage for private tasks (#14717)
Add private task and project ownership, downward access grants, and immutable run/workspace provenance. Apply migration 0313 with bounded batch commits and concurrent indexes. Keep the enforcement and sharing changes in their dependent PRs. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1007.0-canary.7 |