mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 21:03:51 +02:00
codex/plugin-task-execution
231
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
941a3fa991 |
Pin Copilot native dependencies with the maintained lockfile refresh (#15572)
Pin the three optional GitHub Copilot 1.0.88 native packages for Runner and server using the maintained lockfile workflow. Synchronize the package contract and bound initial render readiness in the deliberately throttled browser fixture. Current-head CI and focused checks pass. Co-Authored-By: Dotta <cryppadotta@users.noreply.github.com> Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d0f69670db |
fix(runner): recover saved execution prompts after upgrades (#15518)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs save an immutable execution context for restart recovery. > - The context includes the prompt text, revision, and content hashes. > - The parser required that saved prompt to match the current release. > - A server upgrade could reject a valid saved run before provider recovery. > - This pull request validates and preserves the saved prompt snapshot. > - Routine prompt changes no longer need a catalog of past strings. ## Linked Issues or Issue Description **What happened?** A hot restart selected a dead native runner for same-run recovery. Reading its saved v5 execution input failed with `input.runtimeContext.prompt must match the fixed Paperclip prompt revision`. The new controller accepted only v6. **Expected behavior** Recovery uses the saved prompt and validates its content hashes. It preserves the same run and provider session without starting a duplicate turn. **Steps to reproduce** 1. Start a native run and save its execution input and provider checkpoint. 2. Change the fixed execution prompt in the server release. 3. Stop the runner and recover the saved run with the new controller. 4. Observe that the old parser rejects the saved prompt before provider recovery. **Paperclip version or commit** The v5-to-v6 prompt change was introduced in #15446. The defect also reproduces on current master before this fix. **Deployment mode** Source-built server with the native runner. Related work: #15446 added task-monitor guidance. The held prompt-size experiment in #15489 changes prompt wording but does not add recovery compatibility. ## What Changed - Read the prompt text and revision from the saved execution snapshot. - Treat the revision as non-empty metadata and preserve the exact saved bytes. - Validate the prompt SHA-256 and the aggregate context digest. - Keep fresh-run builders on the current prompt constants. - Test arbitrary saved prompts, malformed fields, altered text, stale hashes, and aggregate drift. - Test recovery parsing for Codex input versions v3-v5 and OpenCode, ACPX Pi, and Dot v6 inputs. - Extend the real-process restart suite with both the incident's v5 wire fixture and a prompt unknown to this release. - Run the restart recovery suite in the existing Rust-equipped PR lane, where its runner and fake-provider binaries are built. Verify complete, non-overlapping test coverage for PR, release, and local callers. - Document recovery from saved snapshots without a historical prompt catalog. ## Verification - Red: the new contract regressions fail against the catalog-based parser with the original prompt-validation error. - Green: 53 focused contract and materialization tests pass. - Red: the real-process unknown-prompt regression fails with master's original parser at the saved-input recovery read after process loss. - Green: all 15 real-process restart tests pass locally on the final branch. The saved-prompt cases keep the run and provider session, replace the PID, and record one `turn/start`. - Local repository `pnpm -r typecheck` and `pnpm build` passed after rebase on `89f09dad723766e5351953f0731b9aa5daada28d`. The 53 focused tests also passed on that head. - Red: the new test-roster checks fail against the old CI placement. - Green: all 26 test-scheduling checks pass after moving the restart suite. - CI ran all 15 restart recovery tests with no skips on final head `6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`. [Runner test job](https://github.com/paperclipai/paperclip/actions/runs/37764782334/job/113271464966). - Greptile scored 5/5 on that exact head with no actionable findings. - The complete CI matrix passed on final head `6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`: general/workspace tests, serialized server suites, both runner Vitest lanes, Rust and static checks, all browser shards, typecheck, build, and the canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37764782334). - There are no unresolved review threads or merge conflicts. - Reproduce focused tests with `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/contracts/runtime-context.test.ts src/contracts/native-execution.test.ts src/drivers/runtime-context-materializer.test.ts`. - Reproduce restart tests with `pnpm --filter @paperclipai/paperclip-runner build:rust` followed by `pnpm exec vitest run server/src/services/native-runtime/native-runner-restart-recovery.integration.test.ts`. They use temporary PostgreSQL, real runner processes, and a fake Codex provider. They do not use paid inference. ## Risks - The parser now accepts internally consistent saved prompt text that is absent from the current source. Inputs must come from trusted server persistence. Content hashes verify consistency; they do not authenticate authorship. - Existing execution-schema, ownership, checkpoint, provider, permission, and session-compatibility checks still apply. - This change validates the saved base prompt. It does not make all additional code-generated instruction strings versioned. - The process-level recovery proof uses Codex. Other provider coverage verifies the shared input parser and retained provider configuration. ## Model Used OpenAI Codex, GPT-6 family, with repository inspection, code editing, and test tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b5342febe5 |
fix(runner): require current-turn completion after connection continuations (#15514)
## Thinking Path > - Paperclip manages agents, tasks, permissions, and execution budgets. > - Native tasks can resume after a connection decision in the same provider conversation. > - Each turn still needs an accepted completion report. > - Compact continuation messages did not explain that reports from earlier turns cannot finish the new turn. > - The connection evaluator could also grade before the final reply was stored or reject valid unavailable-access wording. > - This PR clarifies the current-turn report requirement and fixes those observation boundaries. > - The original connection instructions and strict native completion gate stay in place. ## Linked Issues or Issue Description Refs #15489. The reduction remains draft while this separate repair is qualified. Refs #15471 for the earlier connection continuation work. ## What Changed - Add a current-turn completion reminder to compact continuation inputs. - Keep final prose insufficient for completion. Preserve permissions and retry policy. - Wait for the final successful task run's saved, attributed decline reply within the existing deadline. - Use one bounded explanation matcher for both decline checks. - Wait for a recorded tool-action rejection to dispatch its bound continuation, with strict company, task, agent and source-run checks. - Retain the exact grading input before later API refreshes. - Add failure and delay regressions and update the Runner and evaluator docs. ## Verification The fresh comparison has **15/15 original passes on each variant**: 15 unchanged pass pairs, zero new failures, and no pending pair. There are 30 case attempts and **65 actual agent runs** (baseline 33; candidate 32). All runs succeeded. All 30 cleanups passed. No model attempt was retried. | Profile | Baseline | Candidate | | --- | --- | --- | | Native Codex `gpt-5.6-sol` | 5/5 | 5/5 | | ACPX Claude `claude-sonnet-5` | 5/5 | 5/5 | | OpenCode `openrouter/deepseek/deepseek-v4-flash-0731` | 5/5 | 5/5 | Each profile covers service approval, service decline, connection decline, provider decline, and selection of the second provider. The saved replies, approved briefings, decisions, fixture observations, final task states, and native completion records were inspected. All 18 saved decline-grade snapshots match their original captured inputs and checks. Result, API snapshot, and final ledger run sets agree. - [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37711658378) · [public candidate report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711658378-1/index.html) - [Baseline campaign](https://github.com/paperclipai/paperclip/actions/runs/37711675579) · [public baseline report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711675579-1/index.html) - Candidate source: `79905343bba280d462765faad19a26e7f179259e`. Baseline source: `7c5e120158f1385a1fc5f20be41f66c58a605534`. Both use master context `fc6304dfe5f2e446e09bd052a7b45f51e930f250`. The trusted workflow source is separately frozen at `dd777f4b7343305c4e6f44c422f44a1d78e12e4f`. - Both variants use the same evaluator, fixtures, models, permissions, 720-second cell deadline, and 1,000-cent company and agent hard stops. The only production difference is the compact continuation reminder. - Suite hash: `aee30b74b4d38ada08777798db0932fbc64e368bb427d29caa8cfd86c7f59747`. Definition hash: `ea9e17f54fe0af1acbb2337ddfab8a7e92db8488bd0f59510a63095ca0229b60`. - Provider-free transport capture: startup/resume input stays at 55,726 bytes. Compact continuation input grows from 53,334 to 53,535 bytes. The 42-tool catalog stays unchanged. These are Paperclip input bytes, not complete vendor prompt tokens. - Focused evaluator tests: 31 pass. Native contract, transport delivery, and session tests: 194 pass. Evaluator support: 1,819 TypeScript tests and 128 Node tests pass, with one intentional skip. - Full build, workspace typecheck, and evaluator typecheck pass. Current-head CI passes all required gates. The current rollup has 51 successful check runs, two intentional Storybook skips, and a successful Snyk status. Review is 5/5 with zero unresolved threads. - CI attempt 1 had one initial runtime-fixture health timeout. The exact test and its full 164-test file pass locally. One CI shard retry passed. The original CI failure, its dependent verify failure, and the retry remain visible in [CI history](https://github.com/paperclipai/paperclip/actions/runs/37711235060). - The broad local `pnpm test:run` attempt was interrupted after about 49 minutes (exit 130). It recorded one failure in the unchanged Zep memory-connector disabled-setting test. That test and the full 388-test tool-access file pass in separate local checks; current-head CI also passes. The local cause is not established, and this broad local attempt is **not** claimed as passing. Three earlier local failures also pass in their isolated checks; their original logs remain retained. - The first two setup admissions were cancelled before provider jobs to include the review correction. They made no provider calls. The completed campaigns above are the first and only model attempts for these corrected variants. Cost evidence stays separate from behavior. Original result summaries report only OpenCode amounts: baseline $0.039192096 and candidate $0.063221620. Final run ledgers also retain estimates for Claude (baseline $1.232439000; candidate $1.333611200) and Codex (baseline $2.340324400; candidate $2.129335600). These estimates do not replace the original summaries. Local and GitHub runtime are unmetered here. Invoices are unknown. This is not a cheaper or faster claim. ## Risks - A single matched trial cannot prove general equivalence or causation. The reminder is an instruction change, not a new completion enforcement rule. - Candidate OpenCode service-decline finished within one continuous run; its baseline used two. That pair passed the task outcome, but it does not qualify the reminder on a resumed decline turn. No extra paid run was used to replace it. - The text matcher is bounded evidence of an explanation. It does not prove reasoning or consumption of feedback. Bounded stdout excerpts do not prove that every extra attempted tool call is absent. - Missing saved replies or continuations still fail at the original deadline. Failed native completion remains a failure even when final prose is correct. - The unresolved local-suite discrepancy above remains a validation limit. Full remote CI and both focused local reproductions pass. - The original connection instructions stay in place. These results do not qualify the reduction in #15489. Its original 11/15 versus 12/15 grades and two new failing pairs remain unchanged. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, test inspection and eval analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the focused local tests listed above and they pass; the interrupted broad local run is disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e4b39da6f6 |
Isolate ACPX admission deadline tests from lease ports (#15527)
Apply the reviewed change for Isolate ACPX admission deadline tests from lease ports. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
dd777f4b73 |
feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path
> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.
## Linked Issues or Issue Description
**Subsystem affected**
Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.
**Problem or motivation**
An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.
**Proposed solution**
Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.
**Roadmap alignment**
This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.
## What Changed
- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.
## Verification
- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head
|
||
|
|
fc6304dfe5 |
feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path > - Paperclip manages AI agents, tasks, permissions, and execution budgets. > - Paperclip Runner gives each provider the same admitted task and tool authority. > - OpenAI Dot runs outside the local process tree and needs asynchronous work delivery. > - The merged MCP gateway supplies OAuth consent and signed event delivery. > - A personal assistant grant cannot safely stand in for an assigned agent. > - This pull request adds a separate Dot agent connection and a durable Rust Runner bridge. > - The operator can assign work to Dot and inspect its accepted work, tool receipts, and result. ## Linked Issues or Issue Description **Agent or provider** OpenAI Dot, as an experimental provider of the existing Paperclip Runner adapter. **Why this adapter is useful** An operator can assign normal Paperclip tasks to an existing Dot. Dot can read its mailbox, request work on an assigned task, use admitted task tools, and submit a result. Paperclip keeps company scope, checkout, approvals, known budget limits, and activity attribution. **How the agent is invoked** A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the assignment. The Rust Runner owns the durable turn and operation receipts. The first release supports self-hosted instances with a local Runner controller. **Additional context** This extends the merged public MCP gateway from #14846 and the assistant invitation and device-consent work from #14933. This also integrates the merged assistant tool and configuration expansion in #15380. Dot retains its dedicated agent resource and cannot receive personal configuration permission. The public assistant connection remains a personal connection. ## What Changed - Add a durable Rust Dot provider and its TypeScript Runner driver. - Add closed PRP v3 external-provider operations and native execution input v6. - Add company-scoped pairing, mailbox, assignment, and operation records. - Reuse merged browser/device consent, client metadata verification, webhook admissions, refresh, secret rotation, and warm-standby gates. - Keep Dot scopes, issuer, grants, event workers, and tool access separate from personal assistant access. - Add Dot configuration, pairing, readiness, and consent UI. Keep agent grants out of the personal Connections entry. - Regenerate the Dot-only migration after master. Preserve published gateway migrations. Make the new migration safe to reapply. - Document setup, recovery, accounting limits, evidence, and remaining account qualification. - Reverify reconnect callbacks and wake outstanding work with a fresh mailbox reference; preserve the existing assignment and operation receipts. - Clean up Dot bindings and waiting runs on OAuth revoke and refresh-token replay. Old grants cannot revoke replacement bindings. - Restore the pairing reference when an unsaved agent form is reopened; document board-only pairing routes in OpenAPI. - Accept a clean Rust exit after the acknowledged shutdown receipt. Unexpected exits still require recovery. - Clear the cached binding after a successful revoke so a failed connection refresh cannot restore it. - Add production-component Storybook states and screenshots for pairing and connection review. All preview account data is synthetic. - Persist normalized completion, serialize Dot turns and durable work admission, and poll subscription readiness. - Serialize mailbox writes and cursor reads; retain paused fence acknowledgement without task authority. - Authorize admitted review runs without changing the worker assignee. Include the fenced assignment ID in production stop notices. ## Verification - This PR integrates master `4a8178e9c`. Dot migration `0317_messy_famine.sql` follows the published history and is safe to reapply. The merge preserves the reserved migration connection, batch-commit handling, private task checks, task monitors, and native accounting. - Local workspace typecheck, full build, and UI token gates pass. The server typecheck passes after the review fixes. Database and native executor regressions pass. - All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot driver tests pass. The tests cover native document writing and finalization, durable replay, queue admission, mailbox ordering, admitted reviews, stale authority, production stop references, and paused acknowledgements. - Current head `d0e7e0626` passes all 57 checks: 53 pass and four are intentionally skipped. This includes full typecheck, build, tests, Rust Runner verification, browser E2E, release verification, and Canary Dry Run. Greptile rates this exact head 5/5. All review threads are resolved. - The full local root test run is slower than the sharded CI run and has not completed. The full CI test gates pass on the current commit. Focused local regressions pass. - Real-account pairing and event delivery on this base commit remain unqualified. Live account and setup proof are recorded in the follow-up #15414. The following screenshots use synthetic preview data. They show the production pairing component and do not qualify a real account or the full agent setup journey.   ## Risks - This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP and Paperclip Runner. The separate experimental-settings follow-up in #15414 replaces this environment flag with saved operator settings. - Dot does not expose provider token usage or cost. The operator must acknowledge external billing. Known Paperclip budget gates still apply. - Cancellation fences Paperclip authority. It does not confirm that Dot stopped all external activity. - Assigned skill files and third-party MCP bindings are unsupported and reject admission. There is no mounted workspace, model selector, or provider thread identifier. - Hosted agent-broker and remote controller deployments are not qualified. - The new migration follows the merged master history. Existing prototype databases still need the normal master migration history before this Dot-only migration. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window size are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, and test inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fd8c6b920a |
fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native tasks can pause while a human decides whether to allow a connection action. > - The next turn needs both the saved decision and complete accounting for the previous turn. > - Cancellation could discard final usage, and complete direct Claude API receipts could remain unpriced. > - A stale blocked or review report could also request approval again after the original card was declined. > - This pull request retains shutdown accounting and rejects approval waits bound to an already resolved action. > - The benefit is reliable continuation with the existing budget and approval controls. ## Linked Issues or Issue Description Refs #15420. Related: #15312 addresses requester ownership during dispatch. This change addresses receipt capture and final-response validation. ## What Changed - Retain usage events after native cancellation. Continue to reject late provider messages and work. - Drain same-turn accounting and terminal events for at most fifteen seconds after a durable governed wait. Keep incomplete accounting blocked. - Estimate complete, unpriced, direct Anthropic API receipts for the exact `claude-sonnet-5` model. Record the rate version and assumptions. Use the one-hour cache-write rate when the receipt lacks cache TTL. - Bind stale approval reports to exact interaction, action-request, or invocation IDs in the same company, task, agent, and run. Cover blocked, review, and response-wake reports. Keep independent reviews valid. - Fence checkpoint and result writes after a controller detaches for restart, including operations waiting for a database lock. Reject stale successful returns before certifying accounting. - Allow bounded subscription teardown only after retaining an actual provider terminal. - Journal the exact governed-wait trigger and disposition before provider interruption. Recover that wait independently of a later saved answer, replay retained accounting, and reject mismatched or unproven terminal evidence. - Give settling governed turns a bounded window before shutdown detaches their controller. - Keep fuzzy external app matches alongside installed capability matches instead of forcing an unrelated provider question for a generic query. - Return up to twenty exact active catalog tool names after an invalid request, after eligibility checks; still reject the request without granting access or creating an approval. - Require retained provider terminal proof before settling a governed wait, including when complete usage arrives before stream closure/error/timeout. Retain harmless numbered cancellation events so restart replay stays contiguous. - Isolate accounting-test OpenCode config from the host plugin directory. - Add regression coverage and document the accounting, restart and connection-search behavior. ## Verification Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`; later review fixes are verified separately and do not relabel those runs. - Repository typecheck and build pass. - Current runner runtime and cancellation suites: 191 tests pass. Six new regressions cover stream end/error/timeout without provider-stop proof and contiguous cancellation acknowledgement/request replay; all six failed before the fix. The existing bounded cleanup case now explicitly supplies terminal proof. All 31 adapter accounting tests pass with isolated fixture config. - New checkpoint-rebinding and approval-criterion suites: 65 tests pass; runner HTTP integration: 29 tests pass. Unchanged executor/control-plane suites: 640 tests pass; database-backed connection suites: 68 tests pass. - Current retained evidence verification covers 222 file hashes across all fifteen original result artifacts. The three-case [published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html) and all eight screenshot hashes verify. The [three-case recovery report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html) and all six screenshots also verify; the nine-result campaign did not publish. - Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node tests pass. Eval typecheck and catalog discovery pass. - Full local repository unit run was interrupted before the follow-up edits after three tool-access failures and one runner HTTP failure. Those failures pass in isolation; the 09a full run was interrupted after one rapid Slack callback-ordering failure and seven skill-service failures. All eight pass both isolated and with full-runner environment settings, and all 74 skill-service tests pass together; the subsequent full run reported two 15-second OpenCode accounting timeouts and was stopped with exit 130 to apply review fixes. The timeouts reproduce while copying this host’s 61 MB OpenCode config. All 31 tests pass after isolating config inside each fixture without increasing timeouts or changing assertions. A complete local full-suite pass is not claimed. Repository-wide CI also passes on the final review-fix head in [run 37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952). Earlier database-skipped diagnostics and the older ENFILE run are retained and are not full-suite passing evidence. - Three-case live campaign [37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525) passes all three original grades on frozen source 09a: 55/55 checks, eight succeeded run records, complete accounting receipts, matching checkpoint identities and no pending approvals. Campaign [37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353) adds nine original passes (173/173 checks, eighteen succeeded records) on the identical source. Its other three jobs failed before runner assignment or any step while GitHub could not load the paid environment; those original infrastructure failures are retained. Campaign [37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356) completes only those unstarted cells: all three original grades pass (57/57 checks, six succeeded run records). All fifteen exact cases now pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new overall failures and zero pending pairs. Total current evidence: 285/285 checks and thirty-two succeeded run records, complete accounting receipts, matching checkpoint identities, no pending approvals or retry records. Earlier campaigns retain forty-seven additional run records and two known same-run recovery attempts; actual provider-call counts and invoices remain unknown. The nine-result campaign skipped publication and its public URL returns 403; original artifacts remain retained. Previous ba9 campaign [37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840) completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline pass, while OpenCode resumes but times out searching for exact tool names and never creates the access card. That failure and incomplete cancelled-run accounting remain preserved; the new catalog error guidance targets this observed dead end. Campaign [37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312) remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker leak and unknown-criterion approval gap. The original baseline remains 7 PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS / 3 FAIL. All fifteen selected cases are qualified by their original grades in this bounded trial. These live grades belong to 09a. Its twelve governed-wait checkpoints retain matching same-turn terminal fingerprints, but passing artifacts omit detailed event journals; the later six adversarial regressions qualify the new terminal-proof and replay guards separately. Final-head repository CI passes. [Fresh Greptile review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780) is 5/5, confirms both findings are fixed, and reports no new actionable issues. All review threads are resolved and the PR has no merge conflicts. ## Risks - Governed cancellation drains accounting for up to fifteen seconds. Restart detachment gives a settling batch up to twenty seconds to finish. An incomplete receipt or unproven provider terminal still prevents successful qualification. - Claude prices are estimates, not invoices. The estimate assumes standard global API pricing and uses a conservative cache-write rate. Unsupported models, billers, and billing modes remain unpriced. - Approval identity matching must remain scoped to the current run and the requested approval. It does not authorize execution of a declined call. - Original baseline and final candidate have different merged master context. Exact-case outcomes are before/after observations, not isolated causal attribution to this repair. - No schema migration, fixture, oracle or grader change. Search-result guidance now treats fuzzy external matches as suggestions. Existing app authorization and provider-consent checks remain required. ## Model Used - OpenAI GPT-6 through Codex, with code editing, terminal tools, and test execution. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (191 current runtime tests and 31 accounting tests; interrupted full-suite history and CI coverage are disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f669194298 |
fix(runner): propagate configured environment to future turns (#15451)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runners start provider processes and enforce a separate tool environment policy. > - The server resolves task environment bindings at each run boundary. > - Fixed launch allowlists dropped custom variables after resolution. > - Warm process reuse and a blanket fingerprint exclusion also hid configuration changes. > - This pull request carries a bounded, server-selected list of task variable names through each process boundary. > - The benefit is that later turns and tool commands receive configured values while ambient host secrets stay excluded. ## Linked Issues or Issue Description **What happened?** A native Codex agent could not use configured task credentials. Adding the variables in Settings did not fix the next turn. The provider sanitizer, runner launch, Rust provider launch, and shell policy each used fixed allowlists. Effective config fingerprints also excluded all variables with the `PAPERCLIP_` prefix. **Expected behavior** Explicitly bound task variables must reach provider processes and tool commands. Added, changed, and removed values must take effect at the next run boundary. A running turn keeps its original configuration. **Steps to reproduce** 1. Start a native Codex task without a custom environment binding. 2. Add a fake `PAPERCLIP_PAGE_BUCKET` value and a fake Pages credential binding in Settings. 3. Continue the task and inspect the tool environment. 4. Before this fix, those values are absent even from a fresh runner launch. **Paperclip version or commit** Reproduced at `b31558064`. The fix is rebased on current master. **Deployment mode** Self-hosted server with a native process runner. Related changes: [the legacy Codex MCP environment fix](https://github.com/paperclipai/paperclip/pull/13321) and [ambient server-secret exclusion](https://github.com/paperclipai/paperclip/pull/12870). These affect different launch paths. This change preserves their credential boundaries. ## What Changed - Capture scoped task bindings after resolution, before managed provider credential injection. Mint and validate the names-only projection at native dispatch before host inheritance. Legacy adapters retain their previous environment limits. - Strip user-supplied projection markers from agent, environment, project, and routine config. - Carry selected values through the Codex, ACPX, OpenCode, runnerd, and Rust subprocess launch boundaries. - Add selected names to native and ACPX Codex shell include lists. Keep selected values out of command arguments, including selected bootstrap values. - Reject malformed projections, reserved authority and loader names, missing values, null bytes, and oversized input. - Replace a retained native process when projected values change. Keep unchanged processes reusable. - Fingerprint custom namespaced variables while excluding known generated runtime variables. - Document next-run behavior and add regression coverage. ## Verification - Red: the permanent reproduction failed at four launch/tool boundaries and the namespaced fingerprint check. Two control checks passed. - Green: the initial regression plus existing Codex environment and shell tests passed (39 tests). - Server config resolution, fingerprints, and native-session suites passed (653 tests), including addition, rotation, removal, and unchanged warm-session reuse. - Rust regression tests passed. A real shell child received added and rotated values, then lost them after removal. Unselected host variables stayed absent. - `pnpm -r typecheck` and `pnpm build` passed after rebase. The full `pnpm test:run` was attempted but could not complete: fresh embedded PostgreSQL databases fail during bootstrap on this macOS host. An isolated suite and a disposable native `initdb` probe reproduced the failure before test execution. `shmget` reports `No space left on device` because the host has exhausted shared-memory IDs. This is not disk exhaustion. The run was stopped after confirming the external setup failure. CI results will be recorded separately. - Runner boundary suites: 361 tests passed. Two process-launch errors during concurrent binary staging passed on isolated rerun. - Rust Codex provider and process supervisor integration suites: 97 passed, 2 intentionally ignored. - ACPX shell and selected-bootstrap argv regressions: 3 failed before the fix, then all 55 relevant tests passed. - Review compatibility regression: 129-variable and large-value legacy configurations failed before the correction and passed after moving native-only validation to dispatch. - Final review head `3f11e8d25`: repository typechecks and production build passed. CI completed its implementation checks; one general-server shard hit SQL `40P01` in `heartbeat-runtime-skills.test.ts` during its `beforeEach` table truncate (1,199 tests passed in that shard). The shard passed on its single rerun. All CI checks for this head are green. Greptile reviewed this head at 5/5 with no new actionable findings; the compatibility thread is resolved. - No live provider credentials or model calls are required by these tests. ## Risks - Configured task credentials now reach the tools they were configured for. The controller selects names only after existing scope and secret-binding authorization. - TypeScript and Rust validate the same bounded projection. Their reserved-name rules must stay aligned. - A changed projection replaces an idle provider process. Unchanged values preserve reuse. Active turns retain their original environment. - No database migration or API schema change. > This fixes existing runner configuration behavior. It does not add a new roadmap capability. ## Model Used - OpenAI GPT-6 (Codex), with reasoning, local code execution, and repository tools. The session does not expose a more specific API model identifier or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cfedd7df8 |
fix(runner): restore task monitors and durable timed waits (#15446)
## Thinking Path > - Paperclip manages AI agents and their task execution. > - Agents need a durable way to return to work after a delayed check. > - The issue monitor scheduler already provides a one-shot wake for an assignee. > - Native runners reject generic execution-policy writes and had no bound monitor tool. > - Scheduling alone is insufficient because native completion also needs to accept a timed wait. > - This pull request adds an authorized monitor tool and connects it to completion and the existing scheduler. > - An agent can now schedule its next check, end the run, and resume on the same task. ## Linked Issues or Issue Description **What happened?** A native runner could not set its own task monitor. `call_api` correctly rejected execution-policy writes, while `schedule_wake` had no production binding. `paperclip_finish` also rejected monitor waits. **Expected behavior** A standard native run can set a one-shot monitor on its current task or another accessible task assigned to the same agent. After a confirmed schedule on the current task, it can yield. The scheduler later delivers `issue_monitor_due`. **Steps to reproduce** 1. Start a standard native task. 2. Ask the agent to check the task again later and end its current run. 3. Inspect available tools and try the generic issue execution-policy update. 4. Observe the missing native tool and the lifecycle-write denial. Related PRs: #14680 concerns monitor notes in the shared wake prompt. #11919 changes attempt-limit scope. This PR adds native scheduling and completion authority and retains the existing cumulative attempt bounds. It does not depend on either PR. ## What Changed - Add provider-neutral `set_task_monitor` with a default current-task target, future timestamp, required notes, existing bounds, and explicit clearing. - Check company, task visibility, ownership, runtime permissions, work mode, and active-run authority. Preserve review-only restrictions. Reject the reserved server-owned quota-recovery name before saving or accepting a native wait. - Commit the monitor, audit event, and retry receipt together. Retry receipts survive a successor run without re-arming cleared or consumed timers. - Permit `paperclip_finish` to yield to a persisted monitor. Recheck ownership and the schedule when committing final disposition. Release execution without an immediate continuation. - Preserve due monitors during native execution. Fence wake admission and consumption against replacement, clearing, reassignment, and completion. Preserve unrelated review policy. - Expose scheduled and consumed monitor instructions in task context. Update provider schemas, Rust validation, generated contracts, and execution documentation. - Add an opt-in live Codex smoke script with isolated data and explicit run/session/runner/process evidence. ## Verification - Repository `pnpm -r typecheck` and `pnpm build` passed after rebase. Server typecheck passed again after review fixes. All CI test shards pass on `3def77b1b`, including runner TypeScript/Rust, server, serialized server, workspace, and browser tests. All CI gates are green, including the canary dry run. Greptile is 5/5 on the same commit with zero unresolved threads. - The local monolithic `pnpm test:run`, started before the rebase, was interrupted after current-head CI test coverage passed. It is not counted as a standalone full-suite pass; the focused local regression suites passed. - Targeted server tests cover scheduling, replacement, clearing, policy preservation, cumulative bounds, cross-run retries, permissions, provider-neutral discovery, review restrictions, completion authority, and scheduler/finalizer races. - Runner contract/catalog/semantic tests and Rust terminal-tool tests cover the new operation and monitor completion. - Live Codex test passed twice (latest live run on `e0bcd63e6`) in a temporary database and workspace, with a 300,000 ms warm window. First run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at `2026-10-07T13:15:03.274Z`. Second run `97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`, received `issue_monitor_due`, and completed the same task. Exactly one monitor wake was recorded. - Both live runs used native session `7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner `e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session `01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same process start time. This proves warm reuse for that local Codex test, not only successful scheduling. - Reproduce the paid live test with `node --import ./server/node_modules/tsx/dist/loader.mjs server/scripts/smoke-native-task-monitor.ts --run`, with the installed Codex binary on `PATH` and a valid local login. ## Risks - The scheduler now defers monitor dispatch while the task has an active native run. A stuck run still depends on the existing recovery lifecycle. - Idempotency uses the existing run ledger; no table or migration is added. - Other providers share the tested tool and completion contracts. Only Codex received a live model test. - Existing `call_api` lifecycle restrictions remain enforced. Monitor waits do not bypass task blockers, reviews, or approvals. ## Model Used OpenAI Codex, GPT-6 family, with tool use, code execution, and TypeScript/Rust editing. The session does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d0db8820db |
fix(connections): repair native baseline and approval continuations (#15420)
## Thinking Path > - Paperclip manages AI agents and the tools they may use. > - Connection setup separates provider preference from permission to use a tool. > - The first native connection baseline could not exercise its intended decisions. > - The browser used mutable task titles, and the provider fixture already granted access. > - Native provider-choice instructions also disagreed with the preferred question format. Schema rejection gave no field guidance. > - This pull request repairs those test preconditions and native guidance, then fixes restart/approval defects exposed by the corrected baseline. It also restores missing OpenCode tool-error evidence. > - The benefit is an inspectable baseline before any further instruction reduction. ## Linked Issues or Issue Description Refs #15407. The original 15-cell baseline remains 0 PASS / 15 FAIL. Ten cells stopped on stale titles, two Codex cells had schema denials, two OpenCode cells used already-granted tools, and one Claude cell returned no native result. No intended user decisions were submitted. The exact invalid Codex field and underlying Claude failure cause remain unknown. [Original campaign](https://github.com/paperclipai/paperclip/actions/runs/37562577199) · [Original report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37562577199-1/index.html) ## What Changed - Match the browser's task route and visible identifier instead of a title the agent can change. - Start native provider-choice fixtures with no agent tool access. Verify the public effective-access records. - In the positive case, select Arcade, then grant its exact HubSpot tool through the real access card. Require both saved decisions and exactly one observed call. - Return a canonical `providerQuestionSet` for native input and retain the equivalent legacy `providerQuestion`. - Keep invalid input rejected. Return bounded schema locations and required field names without submitted values. - Preserve Claude's exact session/content identity while allowing authenticated registered instruction-copy paths to rotate on a new run. - Reject duplicate approval reports for an existing exact tool-action card before they create another human review. - Wait for a recorded service-approval continuation within the existing deadline; retain missing or failed continuation grades. - Forward OpenCode tool activity through the runner facade, preserving bounded errors and execution-part identity without inventing host-call joins or exposing arguments. - Preserve original grades, costs, scope limits and diagnoses in the dated repair report. ## Verification - Eval typecheck passes. Support suite: 1,801 PASS, one intentional skip; Node checks: 128 PASS. - Connection/schema tests: 51 PASS. Real-server public fixture setup: one PASS with zero providers. - Browser support regression: five PASS, including renamed and wrong tasks. - Focused Rust safe-feedback test: one PASS. - Repository typecheck and build pass before the latest master replay. Post-replay connection/shared/real-server fixture checks: 52 PASS; eval typecheck passes. The browser review fix additionally passes all five browser checks and seven suite checks. - The full local repository run was interrupted incomplete after about 45 minutes, with five integration failures retained. All five pass in a separate targeted invocation (1,250 unrelated tests skipped). No full local-suite pass or root cause for the initial local failures is claimed. - Corrected frozen source `162cc90fdabe7f505b88ae095044531b82784c92`: **10 PASS / 5 FAIL** across the [passing Codex canary](https://github.com/paperclipai/paperclip/actions/runs/37575158761) and [remaining 14 cells](https://github.com/paperclipai/paperclip/actions/runs/37576261807). The canary passes all 17 checks. Claude's two provider-choice continuations fail on restart, Claude service approval exposes an early evaluator rejection, Codex service approval creates a duplicate approval, and OpenCode provider-second times out after both decisions with no HubSpot call. No original result is regraded. - Final ledgers count 31 actual runs: 27 succeeded, two failed, two cancelled during cleanup. All 15 cleanup/budget checks pass. The late Claude continuation is absent from its earlier workflow snapshot; it remains in the result/API/final ledger. Original evidence retains 279 hashes. Recorded LLM subtotal $0.04553787 is incomplete billing, not actual total cost; local runtime is unmetered. - New repair regressions reproduce the Claude attach failure, duplicate approval acceptance and dropped OpenCode tool events before their respective fixes. Nine Rust attachment checks, 127 ACPX host/adapter tests, 33 completion/control-plane checks, nine eval deadline tests, 59 OpenCode proxy/driver tests, one Rust tool-error/redaction check, and TypeScript/Rust composer parity pass. Eval typecheck, repository typecheck and build pass. Existing support coverage is 1,802 PASS plus 128 Node PASS, one intentional support skip; two additional deadline tests also pass. - New-source full CI/review and live canaries are pending. The next bounded selection is Claude provider-decline, Codex service-approve and one OpenCode provider-second diagnostic with repaired event evidence. No broader campaign or instruction-reduction qualification is claimed. - Initial corrected campaign [37574251834](https://github.com/paperclipai/paperclip/actions/runs/37574251834) was cancelled during shared build after review found the breadcrumb whitespace assumption. Its matrix job has zero steps and no provider execution. The real adjacent-span browser regression now reproduces the old failure and passes after the fix. ## Risks - The corrected baseline remains 10/15. The new restart/approval fixes require live qualification; OpenCode evidence forwarding does not itself establish or fix its prior behavioral failure. - The positive provider case now expects three runs, including separate access approval. Its new results are distinct from the original invalid fixture. - The old Claude missing-result cause and rejected Codex field are unknown. These repairs do not retroactively explain or erase either failure. - Path rotation must preserve prompt, custom instruction, skill/content identity and protected provider settings; regression checks reject stale or changed content. No connection authorization, JSON schema, budget, cleanup, or final-result requirement is relaxed. Historical Everyday prompts and gateway setup remain unchanged. ## Model Used OpenAI Codex, GPT-6. The exact deployment variant and context window are not exposed in this session. Used repository inspection, code editing, test execution and retained-evidence analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
799e4d556f |
fix: make accounting durable and synchronize cost reporting (#14997)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a6306ba606 |
feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
faa8e452c7 |
fix(tasks): stop repeated reminders for historical questions (#15392)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task questions remain saved so a person can answer them later. > - The composer moves an old question into history after a newer human message. > - Native completion still treated every pending question as a required response. > - This caused agents to demand an old answer after the person moved work forward. > - This pull request shares the historical-question rule across context and task execution. > - Agents can finish verified work while the original question remains answerable. ## Linked Issues or Issue Description Refs #15229, Refs #14613, Refs #13130. **What happened?** An agent repeatedly asked a person to answer a question that had moved into feed history. Completion feedback explicitly told the agent to request a response. The saved pending row also blocked task completion. **Expected behavior** A question before newer human direction remains answerable in the feed. Its pending state alone must not require another reminder or stop completed work. A new input blocker, an approval, or a configured review stage must keep its gate. **Steps to reproduce** 1. Let a task agent create an ordinary question. 2. Dismiss the question and send a newer task message. 3. Let the agent finish the requested work and submit its completion report. 4. Observe a demand to answer the old question and a retained completion gate. ## What Changed - Add one company-scoped predicate for historical questions. Only later human comments count. Exclude agent attribution, run attribution, system notices, and untrusted source data. - Apply the predicate to completion feedback, native waits, finalization, commit validation, retry validation, blocked routing, and successful-run handoff. - Include question classification and guidance in heartbeat context and both native task-context tools. Add the guidance to fresh and resumed task prompts. - Replace automatic reminders for current ordinary questions with instructions to assess the real blocker, continue independent work, and withdraw obsolete questions through the existing API. - Preserve historical question rows during completion while cancelling their live native source runs through the existing post-commit and recovery paths. Let an authorized human answer them after completion without reopening work or creating a response wake. Preserve cancellation, current-input, approval, permission, credential, connection, and review gates. Add no dismissal storage or migration. - Document the rule in the execution contract and agent skill. Refresh generated capability source anchors. Add database-backed status, context, attribution, and governance regression tests. ## Verification - `pnpm build` passed. The server rebuild also passed after the lifecycle fix. Generated capability contract and inventory checks passed after the agent documentation update. - `pnpm -r typecheck` passed. Final `pnpm --filter @paperclipai/server exec tsc --noEmit` also passed after the last test additions. - Lifecycle and interaction regressions passed: 224 tests in 3 suites. Context and prompt tests also passed. Final historical-question cases passed (33 tests), native cancellation/recovery cases passed (6 tests), and the existing interaction/confirmation suites passed (72 tests). The full local `pnpm test:run` was attempted and stopped after more than two hours with unrelated fixture/hook timeout failures; it did not pass. All 52 successful GitHub checks are green on the latest commit, including the complete test matrix; no checks are pending or failing. Greptile is 5/5 and both review threads are resolved. - Regression cases cover the old-question/new-human-message sequence, final task status, answering after completion with no wake, live-run cancellation and crash recovery, current input blockers, both task-context tools, API context, timestamp precision, attribution boundaries, and protected gates. ## Risks - A later human task message makes an earlier ordinary question historical even if its input is still missing. The agent must identify the current blocker and ask only for information that still prevents work. - Browser dismissal remains a local preference. Dismissal without a later human message is not recorded by this change. - No schema change or data migration. Completion retains ordinary historical questions; cancellation still expires them. A completed task accepts historical answers only from an authorized human and creates no response-delivery outbox row. Governed requests retain their gates. ## Model Used OpenAI Codex, GPT-6, with repository editing, code execution, and browser diagnostics. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eab93fd4a0 |
fix: checkpoint adapter usage and preserve unknown prices (#14991)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b508a05c43 |
feat: add internal agent complaints and suggestions (#15367)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use legacy skills or native runner tools to work on tasks. > - Those agents can encounter friction that does not belong in the task thread. > - A complaint should preserve the raw reaction. A suggestion should describe an improvement. > - This pull request adds attributed local storage and both submission paths. > - Agents can submit feedback once and continue their primary work. ## Linked Issues or Issue Description **Subsystem affected** Server, database, shared contracts, runtime skills, and native runner tools. **Problem or motivation** Agents have no default internal channel for incidental complaints and suggestions. Sending this feedback through task comments adds noise and can alter task workflows. **Proposed solution** Store free-form feedback in the current instance database. Derive agent, run, company, and task attribution from active authority. Provide default legacy skills and provider-neutral native actions. Keep the instructions close to Warp's MIT-licensed originals. **Alternatives considered** Task comments and external Slack delivery add unwanted side effects. Mandatory suggestion fields and short editorial limits would discard useful feedback. This release has no listing API, UI, read tool, automatic triage, or external forwarding. **Roadmap alignment** This is a maintainer-requested addition to the existing runtime skills and runner tool paths. It does not duplicate a listed roadmap milestone. Searches for complaint tooling, suggestion-box, and agent commentary found no overlapping public PR or issue. ## What Changed - Add the company-scoped `agent_commentary` table, shared validation, and idempotent migration `0310`. - Add one transactional service and the agent-only POST route. Validate active authority before writes or replay. Redact known credentials. Commit a content-free audit with each new record. - Add `submit_complaint` and `submit_suggestion` to standard, ask, and planning modes. Keep review, revocation, and completion restrictions. Store replay identity on the commentary row. - Mount `complain` and `suggestion-box` by default for legacy agents. Bundle a dependency-free Node.js stdin helper in the operational skill and allow its POST through the sandbox bridge. - Preserve Warp's complaint voice and suggestion guidance, with attribution and local transport adaptations. Keep source attribution and MIT notices in each skill's LICENSE, outside runtime instructions. - Document custom-runtime HTTP use and database inspection. Add real-database tests and a repeatable live Codex smoke for local and Daytona execution. - Pin the lagging-source migration fixture before the identity-repair migration so later migrations preserve its regression coverage. ## Verification - Personally ran real Codex submissions in all four environments on 2026-10-06. Local runs passed at 20:35 UTC. Daytona native passed at 20:31 UTC; Daytona legacy passed at 20:33 UTC. Each stored exactly two rows with company, agent, run, and task attribution, wrote the continuation marker, exited zero, created no task comments, and left task status unchanged. Each recorded two content-free activity entries. - Daytona used production provider hooks, real remote execution and file transfer, the legacy queue callback bridge, and native private WebSocket ingress. The current Linux runner was built from `abf47b595`, staged, and verified against controller contracts. Both sandboxes were confirmed deleted. This is a focused feedback transport smoke; it does not claim full Runner E2E catalog or browser qualification. - The immutable base image and Linux binary digest are recorded in [the verification documentation](https://github.com/paperclipai/paperclip/blob/codex/agent-commentary/doc/agent-commentary.md#verification). The smoke script can save content-free JSON evidence. No credentials or feedback bodies are in these reports. | Environment | Runner | Complaint row | Suggestion row | | --- | --- | --- | --- | | local | legacy Codex | `59413a00-1de2-4bb1-bcc6-9c4b54c64aa6` | `3db2364d-3e15-4f47-846f-875d3902999d` | | local | native Codex | `5da22b5f-41df-4de5-8ba0-d9345ab01267` | `2d5abe17-dd41-403c-a5ee-4729f2d58921` | | daytona | legacy Codex | `27c9d0aa-8477-409f-9da0-e8ffa48dee50` | `209681c9-d1e9-4ce1-999e-48fa07692389` | | daytona | native Codex | `6eb001bb-4bcf-43f7-8717-f662f53dc7c3` | `77c383d8-a997-49e5-a33e-25c70e15c0b2` | - Run the local check with `node cli/node_modules/tsx/dist/cli.mjs server/scripts/verify-agent-commentary-live.ts`. The documentation gives the Daytona invocation. Both use disposable instance databases and normal Codex provider usage. - Repository `pnpm -r typecheck` and `pnpm build` passed after the test extension. The build includes runner generation, contracts, and replay checks. The smoke scripts also passed a separate TypeScript check. The lagging-source migration regression passed. All equivalent current-head Vitest CI shards passed. The local monolithic `pnpm test:run` invocation was stopped after CI supplied that coverage; it did not complete locally. - Focused tests cover company isolation, spoofing, revoked credentials, stale ownership, post-finish rejection, concurrent replay, conflicting keys, atomic rollback, and deletion through existing services. Boundary tests cover empty text, Unicode, text beyond 8,000 characters, and the 524,288-character ceiling without truncation. Mounting tests cover Codex, Claude, and sandbox staging. Helper tests cover standalone Node execution, stdin, invalid UTF-8, redirects, HTTP failure, and its deadline. Privacy and bridge tests cover successful and rejected requests. - Instructions were compared with Warp's originals. MIT notices and source credits live only in LICENSE files. Native tools preserve truthful disclosure when asked, without routine announcements. - [Full CI](https://github.com/paperclipai/paperclip/actions/runs/37508559190) and Greptile 5/5 passed on the earlier feature commit `5209c3501`. The later head found the migration-fixture assumption fixed in this update. On `8a4965164`, all 55 check contexts passed after one browser shard rerun. Its initial reviewer signoff failure also passed an isolated local browser run (1 test). Greptile scored that head 5/5 and identified one smoke cleanup gap. `6ecbafb0b` fixes failed-acquisition cleanup with four passing tests and a passing smoke-script typecheck. Fresh CI is pending for this final test-only fix. No commentary production code changed during verification. ## Risks - Feedback is internally attributed. It is not anonymous. Existing redaction removes known credentials, but agents must still omit sensitive content. Normal provider transcripts can include their submitted arguments. - Default skill availability changes for existing legacy agents. Runtime policy filtering still applies. The helper uses the existing Node.js runtime with no extra dependencies; custom runtimes can call the HTTP endpoint. - Feedback is removed with its run, agent, or company. Task deletion clears only the issue pointer. Normal database backups include the table. - The migration is additive and has no backfill. Writes serialize on the active run for replay consistency. No server suggestion quota is imposed. ## Model Used OpenAI `gpt-6-astra` through Codex, with `xhigh` reasoning effort and a reported 258,400-token context window. Capabilities used: repository inspection, code execution, and live runtime verification. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
854af7df19 |
build(deps-dev): bump vitest from 4.1.11 to 5.0.3 (#12969)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.11 to 5.0.3. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v5.0.3</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Isolate <code>result.status</code> between <code>repeats</code> runs - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11218">vitest-dev/vitest#11218</a> <a href="https://github.com/vitest-dev/vitest/commit/5dbebe9e3"><!-- raw HTML omitted -->(5dbeb)<!-- raw HTML omitted --></a></li> <li>Don't print an interceptor warning in browser mode - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11377">vitest-dev/vitest#11377</a> <a href="https://github.com/vitest-dev/vitest/commit/15cc006aa"><!-- raw HTML omitted -->(15cc0)<!-- raw HTML omitted --></a></li> <li>Don't retry when <code>test.fails</code> expectedly failed - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11219">vitest-dev/vitest#11219</a> <a href="https://github.com/vitest-dev/vitest/commit/b24585f08"><!-- raw HTML omitted -->(b2458)<!-- raw HTML omitted --></a></li> <li>Scope cache key generators to projects - by <a href="https://github.com/ecoyoung"><code>@ecoyoung</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11281">vitest-dev/vitest#11281</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11301">vitest-dev/vitest#11301</a> <a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1d"><!-- raw HTML omitted -->(92ba7)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Delay server <code>listen</code> until tests start running - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11366">vitest-dev/vitest#11366</a> <a href="https://github.com/vitest-dev/vitest/commit/7d8ed3e9b"><!-- raw HTML omitted -->(7d8ed)<!-- raw HTML omitted --></a></li> <li>Check mock path boundaries - by <a href="https://github.com/saryn17"><code>@saryn17</code></a>, <strong>Ryosei Sato</strong> and <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11361">vitest-dev/vitest#11361</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11362">vitest-dev/vitest#11362</a> <a href="https://github.com/vitest-dev/vitest/commit/1c3888bce"><!-- raw HTML omitted -->(1c388)<!-- raw HTML omitted --></a></li> <li>Keep config of browser-consumed environments - by <a href="https://github.com/kasperpeulen"><code>@kasperpeulen</code></a> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11378">vitest-dev/vitest#11378</a> <a href="https://github.com/vitest-dev/vitest/commit/aafc0996f"><!-- raw HTML omitted -->(aafc0)<!-- raw HTML omitted --></a></li> <li>Ignore page crash while cancelling - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11386">vitest-dev/vitest#11386</a> <a href="https://github.com/vitest-dev/vitest/commit/7c36748fa"><!-- raw HTML omitted -->(7c367)<!-- raw HTML omitted --></a></li> <li><code>toMatchScreenshot</code> uses wrong reference on retried tests - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11393">vitest-dev/vitest#11393</a> <a href="https://github.com/vitest-dev/vitest/commit/c22aba992"><!-- raw HTML omitted -->(c22ab)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>cache</strong>: <ul> <li>Revalidate imports of cached modules - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11381">vitest-dev/vitest#11381</a> <a href="https://github.com/vitest-dev/vitest/commit/38f98855f"><!-- raw HTML omitted -->(38f98)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>deps</strong>: <ul> <li>Pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into <code>ERR_PNPM_TRUST_DOWNGRADE</code> - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11403">vitest-dev/vitest#11403</a> <a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977"><!-- raw HTML omitted -->(f6c9a)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Pass current equality testers to <code>expect.extend</code> asymmetric matchers - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11401">vitest-dev/vitest#11401</a> <a href="https://github.com/vitest-dev/vitest/commit/3e794a96b"><!-- raw HTML omitted -->(3e794)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Support Blob on jsdom 30.1 - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11379">vitest-dev/vitest#11379</a> <a href="https://github.com/vitest-dev/vitest/commit/6c49b7197"><!-- raw HTML omitted -->(6c49b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>pool</strong>: <ul> <li>Preserve unique pool ids when <code>groupOrder</code> is set - by <a href="https://github.com/mtorp"><code>@mtorp</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11392">vitest-dev/vitest#11392</a> <a href="https://github.com/vitest-dev/vitest/commit/50312ebb4"><!-- raw HTML omitted -->(50312)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>ui</strong>: <ul> <li>Split-pane handle overlapping iframe - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11221">vitest-dev/vitest#11221</a> <a href="https://github.com/vitest-dev/vitest/commit/f91db0dfd"><!-- raw HTML omitted -->(f91db)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vitest</strong>: <ul> <li>Remove root temp dir on close - by <a href="https://github.com/abhinav-phi"><code>@abhinav-phi</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11248">vitest-dev/vitest#11248</a> <a href="https://github.com/vitest-dev/vitest/commit/7c7119cf7"><!-- raw HTML omitted -->(7c711)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vm</strong>: <ul> <li>Do not optimize deps from index.html - by <a href="https://github.com/ezefernandezyf"><code>@ezefernandezyf</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11329">vitest-dev/vitest#11329</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11360">vitest-dev/vitest#11360</a> <a href="https://github.com/vitest-dev/vitest/commit/caf2887de"><!-- raw HTML omitted -->(caf28)<!-- raw HTML omitted --></a></li> <li>Don't reuse scripts across vite environments - by <a href="https://github.com/MO2k4"><code>@MO2k4</code></a>, <strong>Martin Oehlert</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11395">vitest-dev/vitest#11395</a> <a href="https://github.com/vitest-dev/vitest/commit/346d3896b"><!-- raw HTML omitted -->(346d3)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v5.0.2...v5.0.3">View changes on GitHub</a></h5> <h2>v5.0.2</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Bind <code>process</code> in case global is overwritten - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11343">vitest-dev/vitest#11343</a> <a href="https://github.com/vitest-dev/vitest/commit/0b79231ad"><!-- raw HTML omitted -->(0b792)<!-- raw HTML omitted --></a></li> <li><strong>detect-async-leaks</strong>: <ul> <li>Ignore <code>process.stdio</code> handles - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11333">vitest-dev/vitest#11333</a> <a href="https://github.com/vitest-dev/vitest/commit/0fd6b9790"><!-- raw HTML omitted -->(0fd6b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Fix <code>toMatchObject</code> with asymmetric matchers - by <a href="https://github.com/ShreeBohara"><code>@ShreeBohara</code></a>, <strong>Claude Opus 5</strong>, <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-5)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11100">vitest-dev/vitest#11100</a> <a href="https://github.com/vitest-dev/vitest/commit/42523289e"><!-- raw HTML omitted -->(42523)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Fix <code>Request</code> with <code>Blob</code> body on jsdom 28+ - by <a href="https://github.com/harshit-d3v"><code>@harshit-d3v</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11295">vitest-dev/vitest#11295</a> <a href="https://github.com/vitest-dev/vitest/commit/d1c3ecc93"><!-- raw HTML omitted -->(d1c3e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporter</strong>: <ul> <li><code>agent</code> to respect <code>--silent</code> - by <a href="https://github.com/Raj4478"><code>@Raj4478</code></a> and <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11271">vitest-dev/vitest#11271</a> <a href="https://github.com/vitest-dev/vitest/commit/5b95efb6d"><!-- raw HTML omitted -->(5b95e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporters</strong>: <ul> <li>Handle concurrent <code>createReport</code> calls - by <a href="https://github.com/7rulnik"><code>@7rulnik</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11278">vitest-dev/vitest#11278</a> <a href="https://github.com/vitest-dev/vitest/commit/e8e556ff7"><!-- raw HTML omitted -->(e8e55)<!-- raw HTML omitted --></a></li> <li><code>hanging-process</code> to use ESM entrypoint - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11316">vitest-dev/vitest#11316</a> <a href="https://github.com/vitest-dev/vitest/commit/4e91e5668"><!-- raw HTML omitted -->(4e91e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>spy</strong>: <ul> <li>Fix stack overflow when spying <code>Set.prototype.add</code> - by <a href="https://github.com/fengmk2"><code>@fengmk2</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11299">vitest-dev/vitest#11299</a> <a href="https://github.com/vitest-dev/vitest/commit/a0a939653"><!-- raw HTML omitted -->(a0a93)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/33cadea62e8763c455c7fca38d9ab1dda87c5f75"><code>33cadea</code></a> chore: release v5.0.3 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11409">#11409</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/346d3896b65c3c907174447a035807342799f346"><code>346d389</code></a> fix(vm): don't reuse scripts across vite environments (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11395">#11395</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977ad3363f796a737834572e54c6ad5c18"><code>f6c9a49</code></a> fix(deps): pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into `...</li> <li><a href="https://github.com/vitest-dev/vitest/commit/062c75d8b63519211d951d8293ea81b5a9e3c124"><code>062c75d</code></a> chore: fix standalone docs build, update exports maps (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11394">#11394</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/caf2887dee8987a60118d53933f6e9cabd6b3e2a"><code>caf2887</code></a> fix(vm): do not optimize deps from index.html (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11329">#11329</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11360">#11360</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/50312ebb4eca6a98f6d0b2b61d5d9d38cbbabcef"><code>50312eb</code></a> fix(pool): preserve unique pool ids when <code>groupOrder</code> is set (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11392">#11392</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/7c36748fad1eae9687312f2f7ceadce6ec88b5df"><code>7c36748</code></a> fix(browser): ignore page crash while cancelling (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11386">#11386</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1df16a4fa5bbee3f198c582fbd56689d8"><code>92ba7fc</code></a> fix: scope cache key generators to projects (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11281">#11281</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11301">#11301</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/38f98855fa9cd7fd376afb84094eba0fda256a74"><code>38f9885</code></a> fix(cache): revalidate imports of cached modules (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11381">#11381</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/b24585f08f2ea267746a2d6ca0e43edcbb29726f"><code>b24585f</code></a> fix: don't retry when <code>test.fails</code> expectedly failed (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11219">#11219</a>)</li> <li>Additional commits viewable in <a href="https://github.com/vitest-dev/vitest/commits/v5.0.3/packages/vitest">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a590ab769d |
Give managed agents persistent cryptographic identities (#15352)
Give agents persistent Ed25519 identities encrypted with the existing instance master key. Create keys transactionally for new agents and lazily before supported managed runs, expose public identities in the API and agent UI, and protect private material during runtime delivery and output persistence. Preserve identities in recovery backups while giving imported and development-cloned agents fresh keys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9f7057e122 |
feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use a harness, a model, and a credential to run tasks. > - Connections already store credentials and control who can use them. > - Custom providers also need an endpoint and a supported API format. > - A per-agent endpoint would duplicate credentials and access rules. > - This pull request stores routing on the connection and projects it into the harness. > - Isolated credentials and protocol checks keep the selected connection authoritative. ## Linked Issues or Issue Description Refs #37, #13083, #14104, #14565, #12692. #14016 is a reference only. This PR has its own schema, vault persistence, routing validation, runtime projection, and tests. None of the five commits in #14016 is an ancestor of this branch. We do not depend on or plan to merge it. #14967 addresses task-pinned account pools. #14422 addresses another provider integration. This is the first of two linked PRs. Merge this connection change before #15341, which refines agent setup and adds the qualification harness. The split keeps each review below 100 changed files. Provider catalog entries have usable setup forms in this PR. Local browser subscription sign-in is included. ## What Changed - Store non-secret routing metadata on AI connections. Vault provider API keys, including Bedrock bearer API keys. Reject general AWS access keys. - Enforce company, owner, human audience, agent access, connection status, and protocol checks before resolving credentials. Keep reconnect destinations immutable and retain connection identity during key rotation. - Project OpenRouter and compatible custom endpoints into Codex, Claude, OpenCode, and local Hermes. Carry these settings through both legacy and native runner transports. Clear conflicting host credentials and redact keys from diagnostics. - Preserve older OpenRouter accounts and native personal defaults. Add Google API-key accounts and migration `0306` for the two provider-default constraints. - Run local Claude and Codex subscription sign-in behind the existing browser sign-in card. Use private attempt homes and owner-bound completion instead of a copied terminal command. - Seed isolated Gemini authentication and preserve OpenCode workspace permissions. Keep the selected connection authoritative. The independent Gemini and Grok workflow fixes are in #15341. - Keep native OpenCode custom gateway keys in a runner-owned selected-model proxy; the harness config contains only a session-scoped capability. Honor runtime outgoing proxy and certificate settings. Preserve streamed responses and revoke the proxy on close or startup failure. - Allow ordinary members to connect native personal accounts before an agent exists. - Repair routed accounts from task cards using the saved provider destination, protocol, model aliases, and connection identity. - Add provider catalog definitions, model discovery, pinned logos, and complete native and routed setup forms. Allow a personal routed connection before a new agent exists. Keep endpoint authentication keys out of Hermes terminal children. - Recover cancelled or restarted browser sign-in with a clear restart action. Support no-auth endpoints without a vault credential. Add isolation and recovery regressions and runtime documentation. ## Verification - Updated with `origin/master` at `22a3ea341`. Migration `0306` follows the new master migration and passes migration and snapshot checks. - The integrated connection regressions passed 152 tests and 50 native OpenCode driver tests, including key-free child-shell configuration reads, authenticated/no-auth forwarding, streaming, model/path restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass. Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34 executable also completed a turn through the proxy against a local synthetic provider; the reusable key was absent from its config. A second real-executable smoke passed with an HTTPS CONNECT proxy and runtime-specific synthetic certificate trust. Certificate-file and certificate-directory regressions pass. - Task-card repair passed 48 tests, including OpenRouter, Bedrock, and custom gateway reconnect cases. UI typecheck and token gates passed. - The prior core regression set passed 133 tests across new-agent setup, provider forms, browser sign-in, routing projection, and connection authorization. Token gates and UI typecheck passed. - Full workspace typecheck and production build passed again after the latest integration and credential-proxy fix. The merged deterministic runner E2E suite passed 1,400 Vitest tests and 128 Node tests. - Full workspace typecheck passed on the prior linked combined implementation. Production build, Storybook build, 1,316 browser-harness Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the latest master integration and regenerated migration. This exact head passed 54 remote checks with four expected skips and Greptile 5/5; no review threads remain open. An unchanged server fixture had a random six-character issue-prefix collision on its first attempt. All 245 tests passed locally and the single CI retry passed. - A provider-free terminal check used the cited supported Hermes source and dummy keys. Gateway and OpenRouter terminal children could not read the selected key. - The broad local Vitest attempt passed 15,442 tests but was not green. It had an embedded-Postgres startup failure, an HTTP logger timeout, an origin socket error, and a browser cancellation wait timeout. The cancellation wait was corrected. The relevant connection tests and the full origin test file passed separately. Latest-head CI must pass before merge. - Prior credential-backed acceptance exercised task creation, tool use, artifact delivery, completion, and context-dependent follow-up. Claude legacy and native runners passed Bedrock with `us-east-1` and `us.anthropic.claude-sonnet-4-6`. - Historical local qualification retained 43 passing API/gateway cells out of 46. Those attempts span earlier builds. They do not qualify this exact commit or staging. All subscription combinations and staging remain unqualified. - Verify native subscription and API-key setup. Connect a regular provider catalog row. Verify an incompatible harness and a changed reconnect URL are rejected. Use #15341 for the complete browser campaign. ## Risks - Migration `0306` changes two check constraints. It preserves rows and is safe to reapply. It takes normal constraint-change locks. - Credential projection touches several harnesses. CLI upgrades can change provider configuration and session behavior. - The native OpenCode proxy adds a loopback hop, pins requests to the selected model, limits request bodies to 16 MiB, rejects redirects, and expires at session close. It prevents reusable keys in the child configuration; it is not an OS isolation boundary against a process debugger running as the same user. - Custom endpoints must be reachable from the agent environment. Saving a connection does not prove connectivity. Bedrock keys require rotation before expiry. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Provider overloads and an unresolved follow-up timeout also affect live Gemini qualification. We have not patched the installed CLI or marked those cases as passing. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS identity, arbitrary auth headers, and custom routing for other harnesses are excluded. - These PRs do not establish production or staging qualification for every provider and login method. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0fe47882cf |
Allow configurable Runner listening ports (#15353)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner supports authenticated provider ingress. > - Each listening Runner currently binds port 43127. > - Concurrent Runner processes on one computer need distinct listening ports. > - This pull request permits an explicit listening port and retains the default. > - Providers can route each run without changing the Runner protocol. ## Linked Issues or Issue Description **Subsystem affected** Paperclip Runner launch and durable transport. **Problem or motivation** Two listening Runner processes in one network namespace cannot bind the same fixed port. The CLI accepts a port flag but rejects every value except 43127. **Proposed solution** Accept `--listen-port` values from 1 through 65535. Default to 43127 when the flag is omitted. Preserve the wildcard bind address, exact run path, PRP authentication, and secure frames. A warm attachment retains its existing listening port. **Alternatives considered** Separate network namespaces or a shared Runner daemon need more changes. Configurable launch ports preserve the existing process model. **Roadmap alignment** This extends existing Cloud / Sandbox agent support. The duplicate search found no matching Runner listener-port change. ## What Changed - Default an omitted listener port to 43127 and reject invalid values. - Validate configurable ports in the durable transport. - Reuse the selected port during warm attachment and reject port changes. - Cover default and explicit ports, invalid input, concurrent listeners, and warm attachment. - Update transport documentation. Daytona still uses its existing default port. ## Verification - Native `cargo test --locked --workspace` passed (two existing tests ignored). - Targeted listener and CLI tests passed, including executable launches on two concurrent ports and warm listener retention. - `pnpm -r typecheck` and `pnpm build` passed. - Complete GitHub CI is green, including general and serialized test suites, all browser shards, Runner Rust/Vitest lanes, builds, and release checks. - The additional full local `pnpm test:run` is still running; no final local result is claimed. - `git diff --check` passes. ## Risks An explicit port can already be occupied. Runner fails its bind without choosing a different port. Port allocation and ingress authorization remain provider responsibilities. No schema or PRP wire format changes. Existing explicit port 43127 callers continue to work. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository inspection, and code execution. The session does not expose the exact model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e34abee670 |
feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path > - Paperclip gives teams durable tasks, agent execution, budgets, and approvals. > - People also use assistants in Codex, Claude, and other MCP clients. > - Those assistants need a scoped connection that preserves the person’s permissions and attribution. > - Delegating a task must not turn the assistant into the assigned agent. > - This PR adds opt-in user OAuth, ten first-party tools, browser consent, and workflow packages. > - Paid product evals verify the resulting tasks, documents, attribution, retries, and access boundaries. > - The team keeps working after the assistant conversation ends. ## Linked Issues or Issue Description **Problem or motivation** A person cannot connect an external assistant to an existing team through browser consent and safely delegate durable work as themselves. **Proposed solution** Expose an opt-in `/mcp/paperclip` endpoint with individually described first-party operations. Bind every connection to a person, client, company, resource, and scopes. Reuse domain authorization and scheduling. Package shared team-review, delegation, and follow-up workflows for OpenAI/Codex and Claude. **Alternatives considered** Related PRs #9393 and #12549 cover earlier remote MCP and board-operator approaches. This change uses user OAuth and a bounded public catalog. It does not expose a generic executor, operator administration, static shared board credentials, or external agent execution. Registry listing work in #9851 is a separate distribution step. **Roadmap alignment** This maintainer-requested implementation extends the governed MCP gateway, activity attribution, durable work products, and hosted deployment direction in `ROADMAP.md`. It implements the first release of the saved design plan; external agent participation and granted third-party tools remain later releases. ## What Changed - Add MCP 2.0 discovery and task status/comment/document Events on the same authenticated endpoint. Persist subscriptions and delivery receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt callback material, recheck permissions/Cloud membership, and bound retries/expiry. Older MCP clients keep their existing tools. - Add discovery, dynamic client registration, S256 PKCE, resource validation, rotating refresh tokens, revocation, and company consent. Store credentials as hashes and recheck membership at execution. - Add tools for connection identity, agents/projects, task search/read/create, human comments, documents/deliverables, and pending-approval links. Preserve current domain permissions and scheduling. - Add durable mutation receipts across reconnects. Matching retries replay results; uncertain outcomes keep the same request ID and require inspection. - Add consent and connection-management pages, OAuth log redaction, shared plugin workflows, and separate OpenAI/Codex and Claude package outputs. - Add eight paid Product E2E cases across three models, independent durable-state grading, usage evidence, cleanup, and report integration. Add task-document guidance and regenerate the runner capability inventories. - Add migrations 0301 and 0302, the dated implementation plan, result notes, and direct-client setup instructions in `doc/public-mcp.md`. ## Verification - Merge integration `e180b1948`: resolved conflicts with current master, preserved both eval registries, regenerated capability catalogs, and regenerated migrations as 0301/0302 while keeping the original replay-safe SQL byte-identical. Local migration safety/snapshot tests (26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval catalog/grading tests (198) pass. Token and capability gates pass. Full recursive typecheck passed. Fresh Greptile review is 5/5 with no unresolved findings. CI is green on this exact head (55 successes, two intentional skips, one neutral result): one unchanged Cursor sandbox test timed out at 10 seconds, then passed locally in 856 ms. A single retry of that failed shard and the aggregate workflow passed. Merge remains blocked on the repository code-owner approval rule. Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55 successes, two intentional skips and one neutral result. [The earlier CI run](https://github.com/paperclipai/paperclip/actions/runs/36901592350) includes all test shards, browser tests, typecheck, build and canary dry run. Greptile was 5/5 on that commit with no unresolved review threads. GitHub still requires code-owner review under the repository merge rules; passing checks do not bypass that approval. Paid source fingerprints remain separate below and in the dated result note. - Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku 4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback, signature verification and report retrieval in a fresh conversation. A final Mini regression passes after the quota/status fixes. All evidence validates. Bounded tunnel startup retries occur before provider calls and remain visible; failed earlier attempts retain their original grades. - The earlier complete seven-case matrix passes **21/21**, with a separate **3/3** delegation regression. Two preceding matrices also passed 21/21 each. A complete 24-cell matrix including Events has not been run. [The dated results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain exact source fingerprints, failures, model IDs and partial costs. - Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass after merging master. Server typecheck passes after the final quota/status changes. Eval typecheck and all 892 eval-support tests pass. - All 33 real MCP/OAuth tests pass. The preceding combined MCP, redaction, private-address and DNS-rebinding run passed 129 tests; two later MCP regressions cover quota reuse and unchanged-status suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI then found a null checkout result in the existing concurrent-workspace path; logging now uses optional status access. All 12 closed-workspace tests and all 33 MCP tests pass after that correction. The exact-start event calibration exposed a timestamp gap; scanning now includes the subscription start, with all 33 MCP tests and server typecheck passing. These two narrow corrections follow the paid regression. - A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0 subscription/delivery, current membership loss, unsubscribe, legacy SDK tools, refresh and revocation. Its callback transport is a fixture with independent HMAC verification. The paid Events campaigns separately prove public HTTPS delivery. - Earlier component qualification passed UI 7,117 tests, CLI 502, shared 832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module boundaries, migration order and plugin regeneration passed. CI covers general/serialized suites, eight browser shards, runner checks, typecheck, build and canary dry run. - **Local full-suite limitation:** the earlier monolithic run was not clean. It encountered overlapping schema rebuilding, Mac database shared-memory limits and isolated CLI/fixture failures. Targeted reruns passed. The existing >32 MiB Git filename stress test still hit its 300-second Mac timeout. The additional serialized sweep stopped after 62 passing suites once CI passed. Original failures and partial logs remain; this PR does not claim a wholly green local monolithic run. - Local Codex CLI and Claude Code OAuth login and MCP SDK interoperability were verified. Public-store installation, actual ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted newcomer provisioning remain release gates. Enablement is moving to **Settings → Experimental → Assistant connections (MCP)** in the stacked follow-up [#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge both for the intended setup experience. This foundation branch alone still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set `PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and connect to `/mcp/paperclip`. Select a team and allow writes in browser consent. Configure an available agent and budget, then delegate and retrieve results later. For Events, rescan the deployed plugin catalog in ChatGPT Work Cloud; the host supplies its webhook credentials when the user asks to watch a task. See [the setup runbook](doc/public-mcp.md). ## Risks - Events are at-least-once and may arrive out of order. No replay cursor is advertised. Clients must refresh finite subscriptions, read current state and avoid comment feedback loops. Callback material uses the instance secrets master key; hosted subscriptions require the updated Cloud broker and are bounded to five minutes/the access proof expiry. - ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging subscription remain deployment gates. Local signed-webhook and paid model evidence does not claim those surfaces have been exercised. - Disabled by default. Merging adds schema and opt-in code; it does not deploy a public endpoint, publish a store listing, create a team, or start paid agents. - Migrations 0301 and 0302 are additive and idempotent. Their SQL is unchanged from the earlier preview numbers, so hash-aware upgrade reconciliation preserves prior staging applications. Normal instance upgrades must apply it before enabling MCP. - Task creation and comments can schedule paid agent work. Consent and tool descriptions disclose that effect. Revocation blocks future calls but does not undo delegated work. - Public deployments need edge rate limits and credential-safe logging. Internal dispatch is restricted to the closed catalog and carries a request-local verified actor. - Hosted onboarding requires the companion Cloud broker, encryption-key configuration, and tenant rollout. Self-hosted direct connections can use this PR alone. - Store acceptance and agent-mode participation are not claimed. Checked-in plugin endpoints are development defaults; rebuild packages for a real deployment before installation. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A more specific serving version and context-window size were not exposed by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`, `claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted/component checks; full local-run limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
16b7db35ff |
Shorten planning skills and measure task decomposition (#15296)
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1477d1ecea |
test: remove the no-op sequential describe modifier (#15286)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs the server and runner test suites with Vitest. > - Vitest 5 removes the deprecated `describe.sequential` property, so the pending Vitest 5 upgrade fails the type-check and test jobs. > - `describe.sequential` only changes behaviour inside a `describe.concurrent` suite, or when `sequence.concurrent` is on. > - This repository has neither, so the modifier changed nothing at run time. > - The benefit is that the Vitest 5 upgrade can land, and the test files lose a modifier that did no work. ## Linked Issues or Issue Description Refs: #12969 ## What Changed - Replace every `describe.sequential` use with a plain `describe` call. - Drop the `{ concurrent: false }` suite option from the two runner test files. - Add a comment to `server/vitest.config.ts` that records why these suites must run one test at a time. - Leave the package manifests and the lockfile unchanged. ## Why the modifier did nothing The Vitest documentation states that `describe.sequential` is useful to run tests in sequence inside a `describe.concurrent` suite, or with the `--sequence.concurrent` option. `sequence.concurrent` defaults to `false`. This repository satisfies neither condition: - No test file uses `describe.concurrent`, `it.concurrent`, or `test.concurrent`. - `server/vitest.config.ts` sets `sequence.concurrent: false`, with `maxWorkers: 1`, `maxConcurrency: 1`, and `isolate: true`. - `packages/paperclip-runner/vitest.config.ts` sets no `sequence` block, so the `false` default applies. `packages/db` and `cli` already run the same embedded-Postgres suites with a plain `describe`, and those jobs are green. The server package was the only outlier. The modifier did carry one real piece of knowledge: these suites need their tests to run one at a time. The new comment in `server/vitest.config.ts` records that reason next to the setting that enforces it. ## Verification - `git grep` for `describe.sequential` returns nothing outside `node_modules`. - The author ran the changed server test files under the installed Vitest 4, and the results match the results without this change. - Two very large embedded-Postgres test files exceeded the author's local memory limit, so the CI test jobs cover those two. - The two changed runner test files have pre-existing local failures caused by a missing Rust toolchain and a missing global `pnpm` binary. The failures are identical with and without this change. - The author type-checked the changed files and found no new error. - CI must pass the typecheck, build, server test, and runner verify jobs. ## Risks - Low risk. Suite execution stays serial, because the Vitest config enforces it. - The change adds no dependency and changes no package manifest or lockfile. - A future change that turns `sequence.concurrent` on would break these suites. The new config comment warns against it. ## Model Used - Claude Sonnet 5 — code edits and local verification. - OpenAI Codex, GPT-5 — the earlier revision of this branch. ## Test plan - [x] Every CI check reaches a terminal green state. A pending or queued check is not a pass. - [x] The `Typecheck + Release Registry` job passes. This change must not introduce a type error. - [x] The `Build` job passes. - [x] The server test jobs and the runner verify jobs pass. - [x] Greptile re-reviews this commit set and posts a passing verdict. The dependabot waiver does not apply to this pull request. - [x] `mergeable` reads `MERGEABLE` as a terminal value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the related public issue with `Refs: #12969` - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Priya Raman <priya.raman@paperclip.ing> --------- Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com> |
||
|
|
63f3aa2dbf |
fix(runner): continue restart-interrupted Codex turns (#15297)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs retain their provider conversation across server restarts. > - A dead runner can restore a Codex conversation after its active turn is lost. > - The old recovery path synthesized a failed task result from that interruption. > - The task then required operator action even though its conversation and workspace were available. > - This pull request preserves the interruption cause and uses the admitted restart attempt for one continuation in the same conversation. > - The agent can reconcile unfinished actions and complete the current request without resending the original task. ## Linked Issues or Issue Description Related: #12845 added native restart recovery. #15042 covers admission during shutdown. #14796 covers legacy shutdown recovery. This change covers a lost native Codex turn after successful conversation restoration. **What happened?** After a server restart killed the local runner, Paperclip restored the saved Codex thread. Runnerd found that the old turn was no longer active. It synthesized a failed terminal and a needs-review result from the last progress message. The transport also discarded the terminal error when it reconstructed thread history. The task failed instead of continuing. **Expected behavior** After proving that the old process stopped and admitting a bounded recovery attempt, resume the current request in the same conversation. Preserve the workspace. Inspect unfinished actions before proceeding. Keep real provider failures, accepted results, intentional stops, unknown unreconciled effects, and exhausted attempts subject to their existing rules. **Steps to reproduce** 1. Start a local native Codex run and leave its turn active. 2. Kill the isolated runner and provider processes, as can happen during a server restart. 3. Restore the same provider thread with no active turn. 4. Observe the synthetic task failure. The new real-process regression reproduces this boundary with a scripted provider. ## What Changed - Record an explicit recoverable process-loss cause without inventing a task result. - Preserve terminal errors and prior turns in reconstructed provider history. Recover the authoritative saved result when adopting an accepted continuation. - Send one continuation in the same conversation for an admitted dead-runner recovery. Require reconciliation of unfinished commands and external actions. - Persist the interrupted terminal before submission and retain the existing recovery marker across another controller loss. - Keep provider attempt limits, terminal failures, and intentional cancellation behavior. - Add red/green regressions, real process-kill coverage, restart checkpoint coverage, and retry-budget coverage. Document the behavior and run-log evidence. ## Verification - Red: the new native runtime regression rejected with `NativeProviderTerminalFailure` on the original code; the Rust restore regression found a missing recovery cause. - Red/green: if restoring the conversation fails and replacement is allowed, the replacement receives the full task and fresh-session handoff. Both prepared and legacy execution inputs are covered. - Red: a second controller crash after the provider accepted the continuation caused an extra `turn/start`. The regression now proves there are exactly two submissions total: the original and its continuation. - Green: focused runtime, backend, driver recovery, and real-process restart suites (207 tests). After the final history/result changes, driver recovery and real-process restart suites passed again (36 tests). - Green: complete Codex transport suite (186 tests), server restart classification/database integration suites (34 tests), and Rust Codex provider suite (92 passed, 2 ignored). - Full `pnpm -r typecheck` and `pnpm build` passed on `60141e649`. The subsequent replacement-prompt guard passed the Runner TypeScript check and the complete runtime plus process-restart suites (148 tests). - Local full-suite attempt: `pnpm test:run` reported two failures in the untouched chat integration suite. Both passed individually, and the complete chat suite passed on rerun (1,063 tests). After all remote test shards passed, the duplicate serial local run was stopped with SIGINT; it is not claimed as a full local-suite pass. - Latest-head CI (`b1297dcd4`): 55 successful checks and 4 intentionally skipped checks, including all test shards, typecheck, build, native Runner verification, end-to-end tests, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37405346303). - Greptile: 5/5 on the latest head, with no unresolved review threads. The PR is mergeable. - The process tests use the real runner binary and a scripted Codex provider. They do not call a live model service. ## Risks - This changes local Codex recovery after process loss. A continuation can execute more work in the retained conversation. Its prompt requires state inspection before repeating an uncertain action; the system does not replay tool calls. - Recovery shares the existing three-attempt budget and one-shot continuation marker. Real failures and older unmarked failed checkpoints are not reopened. - No database migration or API change is required. ## Model Used - OpenAI Codex, based on GPT-6. The exact model ID and context-window size are not exposed in this session. - Capabilities: reasoning, source inspection, tool use, code editing, code execution, and test analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7eadc714d2 |
Verify native semantic input against its raw wire digest (#15301)
## Thinking Path > - Paperclip manages AI agents and their work. > - Native runners send authenticated semantic tool inputs to the control plane. > - The runner hashes the complete input, but the receiver used a redacted receipt hash. > - Protected fields and credential-like document text can therefore fail integrity validation even when the input is unchanged. > - This pull request verifies the complete input with the existing canonical hash function. > - Receipt redaction and permanent rejection of altered input remain in place. ## Linked Issues or Issue Description **What happened?** The Rust runner preserves tool arguments and hashes their canonical JSON. The TypeScript receiver instead redacts the input before hashing. A valid input such as a synthetic document containing `Bearer fixture_token_123456` fails with `native_event_replay_conflict`. A digest of redacted input can also pass without proving the original protected values. **Expected behavior** Verify the complete transmitted input after authentication and exact scope checks. Reject any incorrect digest before durable commit, dispatch, or ACK. **Steps to reproduce** Run the new authenticated controller regressions against the prior receiver. The protected-field and credential-like document cases fail, and the redacted-digest rejection case receives an ACK. The same tests pass with this change. **Paperclip version or commit** Reproduced from source at `858094ba8123c7edb56623597cd96f0391f7e2d4` with synthetic fixtures. Applies to native runner deployments. Related: #14937 preserves semantic input bytes for execution. GitHub issue and PR searches found no duplicate fix; #14591 touches a separate question-draft contract. ## What Changed - Use the existing bounded raw canonical digest for incoming semantic and MCP tool inputs. - Keep receipt and diagnostic redaction unchanged. - Share four digest fixtures between Rust and the authenticated TypeScript controller, including protected fields, document text, Unicode keys, and number boundaries. - Test raw acceptance, altered protected values, forged and redacted digests, canonicalization limits, permanent reconnect fences, and authentication/scope rejection. - Document the separate wire and receipt contracts and the unchanged recovery fence. ## Verification - Before the production fix, the new selected regressions produced three expected failures and eight passes. - Focused controller, receipt, and semantic-tool suites: 138 tests passed. - Rust shared digest fixture test: 1 test passed. - Full workspace typecheck and build passed; the final receiver delta also passed its TypeScript typecheck. Canonical Linux PR CI passed the full aggregate test gate, runner checks, build, and canary dry run at `efcd38e2d8a9fa6340db4f4e863b6febbe8ca32a`. - Independent review passed, including 18 independently run authenticated regressions and the Rust golden test, plus both matching-digest size-limit cases after the test-only follow-up. Greptile is 5/5 on the current head with no unresolved threads. Diff whitespace and local secret/PII scan passed; fixtures are synthetic. ## Risks The verifier remains strict: there is no redacted-digest fallback. Existing authentication, scope, sequence, replay, settlement, and authorization checks remain in force. Receipt storage and redaction behavior are unchanged. This patch does not clear failed-run fences or replay prior work. Synthetic tests prove the protocol mismatch; they do not identify the contents of any historical rejected input. ## Model Used OpenAI Codex (GPT-6), with reasoning, code execution, and independent agent review. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4857799a88 |
feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes. Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6c36c07a4f |
feat(adapters): add GPT-6.1 Sol and refresh shared coding harness pins (#14942)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents run through coding-agent adapters and the native runner. Both use the same installed provider CLIs, model catalogs, and reasoning controls. > - OpenAI released GPT-6.1 Sol (`gpt-6.1-sol`) in Codex. Anthropic released Claude Sonnet 5.5. The static Codex, Bedrock, and OpenCode catalogs do not list these IDs. > - The shared provider pack pins Codex 0.156.0 and OpenCode 1.18.32. The evaluation image pins older Grok, Gemini, Kimi, Cursor, and GitHub CLI releases. Codex 0.156.0 has no bundled metadata for GPT-6.1 Sol. > - A model entry without a current harness, or a harness pin without its runner integrity checks, fails at run time. > - This pull request adds the verified model IDs and moves the harness pins, executable digests, controller checks, and image pins together. > - The benefit is that operators can select the current models, and the native and local adapters share one current CLI installation. ## Linked Issues or Issue Description Refs #13829 and #13838 (the September 22, 2026 model and harness refresh). Related pull requests: #14993 (merged October 5, 2026, superseding #14816) added the direct Claude Sonnet 5.5 entry and refreshed the Claude runtime to Agent SDK 0.3.286 / Claude Code 2.1.286. This pull request does not change the Claude runtime or the direct Claude model list; it keeps the #14993 pins and adds only the Bedrock Sonnet 5.5 ID. After #14993 merged, this branch was rebased onto `master` (October 5, 2026). The six overlapping pin regions (`docker/daytona-runner/Dockerfile`, `docker/daytona-runner/README.md`, `package.json`, `pnpm-workspace.yaml`, `packages/adapters/claude-local/src/index.test.ts`, `packages/paperclip-runner/src/backends/native-backend-factory.test.ts`) were resolved by keeping this pull request's Codex 0.160.0 and OpenCode 1.18.34 pins next to #14993's Claude 0.3.286 / 2.1.286 pins, taking the union of the Sonnet 5.5 model IDs in the Claude test, and merging both README paragraphs. The Sonnet 5.5 effort and CLI-gate lines in the Claude adapter were identical in both pull requests and merged without a diff. #14917 and #14918 reordered the Claude and Codex model lists earlier; the new entries sit where those ordering rules put them. Sources checked on 2026-10-02: - [OpenAI Codex models](https://learn.chatgpt.com/docs/models): GPT-6.1 Sol uses `gpt-6.1-sol`, supports reasoning efforts from Light to Ultra, and has Standard and Fast modes at launch. The page also records that `gpt-5.4` and `gpt-5.4-mini` retired from Codex with ChatGPT sign-in on August 31, 2026, and that `gpt-5.5` retires on October 14, 2026. Neither retirement applies to the OpenAI API. - [Codex CLI releases](https://github.com/openai/codex/releases) 0.157.0 through 0.160.0. The bundled model metadata in the 0.160.0 Linux binary contains `gpt-6.1-sol`. - [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview): Bedrock ID `anthropic.claude-sonnet-5-5`, released September 28, 2026. - [OpenCode releases](https://github.com/anomalyco/opencode/releases) 1.18.33 and 1.18.34 (fixes only). The OpenCode model registry lists both added provider-qualified IDs. - npm `latest` tags for `@xai-official/grok` 1.0.46, `@google/gemini-cli` 0.62.0, and `@moonshot-ai/kimi-code` 2.1.1. [xAI](https://docs.x.ai/docs/models), [Google](https://ai.google.dev/gemini-api/docs/models), and [Kimi](https://www.kimi.com/code/docs/en/kimi-code/models.html) list no newer coding models. - Cursor CLI 2026.10.01-e373342 is the version the official installer resolves. The pinned digest is the SHA-256 of the versioned Linux x64 archive. - [GitHub CLI 2.102.0](https://github.com/cli/cli/releases/tag/v2.102.0) (security fixes). The pinned digest matches the release `checksums.txt`. ## What Changed - Codex adapter: add `gpt-6.1-sol` to the model list, the Fast mode list, and the Ultra effort set. It is the first entry: #14918 orders the list newest version first, and its description notes the ChatGPT app lists GPT-6.1 Sol first. Update the adapter documentation text. - Claude adapter: add `us.anthropic.claude-sonnet-5-5` (Bedrock Sonnet 5.5) to the Bedrock catalog in the newest-Sonnet slot after Opus 5.5; `us.anthropic.claude-sonnet-5` moves into the older-Sonnet group, matching what `sortClaudeModels` from #14917 produces at runtime. Any Sonnet 5.5 ID (direct or Bedrock-qualified) now gets the documented `xhigh` and `max` efforts and requires Claude Code 2.1.284 or later on the CLI lane (the Claude Code changelog entry for 2.1.284 adds `claude-sonnet-5-5`). These two lines are identical to the ones #14993 merged, so the branch carries no diff for them. - OpenCode adapter: add `openai/gpt-6.1-sol` and `anthropic/claude-sonnet-5-5` to the static fallback catalog. - Codex runtime pin 0.156.0 → 0.160.0 in the root and workspace overrides, the runner package, the Codex ACP package patch, the qualified ACPX profiles, the Linux x64 executable digest, the Rust provider backend and its tests, the provider-pack manifest pins, the remote controller pins, the sandbox npm install spec, and the opt-in qualification scripts. - Remote Codex compatibility window: upper bound 0.157.0 → 0.161.0. The minimum stays at 0.149.0. - OpenCode runtime pin 1.18.32 → 1.18.34 in the runner package, the materialization script, the server and Rust qualified versions, the eval and live-session labels, fixtures, and the configuration label. - Evaluation image (`docker/daytona-runner/Dockerfile`): Grok CLI 1.0.46, Gemini CLI 0.62.0, Kimi Code 2.1.1, Cursor CLI 2026.10.01-e373342 with its digest, GitHub CLI 2.102.0 with its digest, Codex and OpenCode version probes, and the refreshed lockfile digest. The Claude Code 2.1.286 probe comes from #14993 and is unchanged here. - `pnpm-lock.yaml` is not part of this pull request. The repository's pull request gate rejects lockfile edits, and the refresh bot regenerates the lockfile on master (the same flow #13838 used). The Dockerfile `PAPERCLIP_RUNNER_LOCK_SHA256` default is the digest of the lockfile that `pnpm install --resolution-only --ignore-scripts --no-frozen-lockfile` (the refresh workflow's command) produces for the combined pins on the rebased branch (`e1856797…`); that lockfile differs from master only in the `@openai/codex` 0.160.0 platform packages, the `@anthropic-ai/claude-agent-sdk` 0.3.286 override that #14993 introduced (the open refresh-bot pull request #14872 carries that part), `opencode-ai` 1.18.34 with its Linux x64 baseline, and the `codex-acp` patch hash. - Documentation: runner README, runner compatibility doc, environment variable example, and a new `doc/adapter-model-audit-2026-10-02.md` with sources and deferred items. - Tests: Codex adapter catalog, server adapter models, Codex compatibility window, native session executor pins, runner package contract, OpenCode materialization, and UI effort options. Unchanged on purpose: Claude Agent SDK 0.3.286 / Claude Code 2.1.286 (already on `master` from #14993), ACP bridges (`acpx` 0.13.1, `claude-agent-acp` 0.73.0, `codex-acp` 1.6.2; newer upstream releases need a separate qualification), the native Grok runtime 1.0.13, Pi 0.84.2 / 0.87.1 (the Pi 1.0 runner stack covers it), and Hermes 0.19.0 (current). `gpt-5.4` and `gpt-5.4-mini` stay in the picker because the OpenAI API still serves them. ## Verification Run on Linux x64 with Node 25.9.0 and pnpm 9.15.4 after `pnpm install --no-frozen-lockfile` (the refreshed lockfile stays local; see above). The results below were re-run on the rebased head (October 5, 2026) for the suites the conflict resolution touches; the other rows are from the original run and are covered by CI on every push: - Rebased head: `packages/adapters/codex-local` 482 passed; `packages/adapters/claude-local` 340 passed, 4 failed (`execute.remote`, `test.probe`, `execute.acp-fallback`, `acp` spawn/env-hardening cases that fail identically on unchanged `master` in this host environment); `server` adapter-models + codex-runtime-compatibility + native-session-executor + adapter-registry 607 passed, 1 failed (the same adapter-registry override-pause case as before, also failing on `master` here); `packages/paperclip-runner` native-backend-factory + qualified-profiles 36 passed; `ui` codex-reasoning-effort + config-fields + model-utils 19 passed. Rust, full typecheck, build, and the Docker image are left to CI as before. - `vitest run` in `packages/adapters/codex-local`: 13 passed. `vitest run` in `packages/adapters/claude-local` (whole package, including the new Sonnet 5.5 gate and effort tests): see the latest CI run and the comment below. `vitest run` in `packages/adapters/opencode-local`: 48 passed, 1 failed (`runtime-config.test.ts` reads the host `PAPERCLIP_OPENCODE_PROVIDERS` variable; it fails the same way on the unchanged base). - `vitest run src/__tests__/adapter-models.test.ts src/services/native-runtime/codex-runtime-compatibility.test.ts src/__tests__/adapter-registry.test.ts` in `server`: 84 passed, 1 failed (`adapter-registry.test.ts` override pause test; it fails the same way on the unchanged base). - `vitest run` in `ui` for `codex-reasoning-effort`, `agent-setup-fields`, `config-fields`, and `ComposerRunSettingsPicker`: 25 passed. - `node --test test/acpx-codex-package-contract.test.mjs scripts/materialize-opencode-binary.test.mjs scripts/runner-protocol-eval-campaign.test.mjs` in `packages/paperclip-runner`: 23 passed. The package contract test verifies the installed Codex ACP executable digest and the 0.160.0 patch pin. - `vitest run src/drivers/acpx src/backends src/drivers/opencode src/live/live-session.test.ts` in `packages/paperclip-runner`: 626 passed, 5 failed, 1 skipped. The 5 failures (`installation-integrity.test.ts` `/proc/self/fd` module loading and one OpenCode answer-selection test) also fail on the unchanged base under Node 25; Linux CI runs Node 24. - `pnpm run test:opencode:qualification` in `packages/paperclip-runner` against the installed OpenCode 1.18.34 executable: passed. - `codex --version` from the installed pack prints `codex-cli 0.160.0`. The Linux x64 executable digest `12eb3e81…652aad` was computed from the `@openai/codex@0.160.0-linux-x64` archive after checking its registry `dist.integrity`. - `pnpm check:token-gates`: all gates clean. - `pnpm run typecheck:typescript` in `packages/paperclip-runner`: passed. Package typechecks ran one at a time; see the comment below for the server and UI results. Not run here, and needed from CI: - Rust tests and `pnpm -r typecheck` / `pnpm build` for the server (no `cargo` in this environment; the server typecheck prepares the runner vendor build). - The Docker evaluation image build and the real-binary Codex startup and session-resume probes (no Docker; the probes need the compiled `paperclip-runnerd`). The trusted CI runner workflow covers them. - Authenticated inference with any new model. This change is metadata and startup validation only. ## Risks - Codex 0.160.0 changes the bundled model catalog and app-server behaviour (authoritative provider catalogs, incremental running-turn tracking). The patched `codex-acp` 1.6.2 bridge is unchanged and declares `^0.148.0`; it worked with 0.156.0 under the same override. If CI probes show a protocol change, the pin can return to 0.156.0 by reverting this pull request. - The compatibility window upper bound moves to `<0.161.0`. Remote images with Codex 0.157 to 0.160 become accepted. Older images stay accepted down to 0.149.0. - Until the refresh bot lands the regenerated lockfile on master, the Dockerfile lockfile digest default does not match the committed lockfile. The trusted CI workflow computes the digest from its own resolution at build time, so this affects only a local build that passes no digest. - Existing saved model selections and effort settings are not changed. Agents on `gpt-5.4` or `gpt-5.5` with ChatGPT sign-in need a model change before the OpenAI retirement dates; that is documented, not enforced. - Rollout order: deploy the controller and runner from this change before promoting a sandbox image that carries these pins. Older controllers reject the new provider-pack pins. ## Model Used - Claude Fable 5.1 (Anthropic, model ID `claude-fable-5-1`), 1M context window, adaptive thinking, tool use. The model ran as a Paperclip agent through the Claude Code harness, performed the web research, edited the code, and ran the tests listed above. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Bender (Fable) <noreply@paperclip.ing> Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
cab4263dc9 |
feat(claude-local): add Sonnet 5.5 and refresh the qualified Claude runtime (#14993)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Claude local adapter lists models for users with a Claude subscription. > - The list needs Claude Sonnet 5.5 and its supported effort levels. > - Sonnet 5.5 needs Claude Code 2.1.284 or later on both execution paths. > - The qualified ACP runtime previously used Claude Code 2.1.280. > - This change adds Sonnet 5.5 and pins Agent SDK 0.3.286, which includes Claude Code 2.1.286. > - Users can select the model and run it with a qualified runtime. ## Linked Issues or Issue Description Refs #3936. This replaces #14816 because the maintainer integration cannot write to the contributor fork. Thank you to @SkilLab-Tech for the model support, runtime refresh, tests, and platform digest verification. This branch preserves both original commits: `8d2f3261af61a2ac1120e51e8a8618732ace543b` and `e68d1d002a3ed745f016fb11c50ac5a3c5a9ff8d`. Related work: - #14917 added Claude model ordering. This branch includes that merged change and resolves its conflicts with #14816. - #14942 updates the other models and harnesses. It remains separate. Its matching Sonnet effort and CLI-gate changes are identical. Both PRs merge with master. The second PR will need a rebase after the first merges because adjacent runtime-pin and test edits conflict. - #14954 is another Sonnet 5.5 change. It overlaps with the model additions but does not include the qualified runtime refresh. - #14039 makes the per-task effort picker model-aware. #3937 is also related to effort selection. The original author checked the [Claude Code changelog](https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md) and [effort documentation](https://platform.claude.com/docs/en/build-with-claude/effort) on 2026-10-01. ## What Changed - Add the direct `claude-sonnet-5-5` model and Low, Medium, High, X-High, and Max effort levels. - Require Claude Code 2.1.284 or later for that model on the CLI path. - Put Sonnet 5.5 after Opus 5.5 in the current-model group. Keep Sonnet 5 in the older-model group. - Retain the Sonnet 5.5 assertions and the model-order assertions in the server tests. - Pin Agent SDK 0.3.286 and Claude Code 2.1.286 across overrides, integrity digests, qualified profiles, Rust provider pins, and the Daytona version check. - Update the related adapter and runtime documentation. ## Verification Local verification uses the resolved source tree and pnpm 9.15.4. Model tests passed on Node 25.9.0. Runtime integrity tests use CI's Node 24.21.0. - Five focused Claude test files pass: 62 tests. They cover model defaults, model ordering, CLI gates, and remote execution probes. - Server model-list and UI setup tests pass: 30 tests. - The Claude adapter typecheck passes. - Runner integrity and qualification tests pass on Node 24.21.0: 78 tests. Four descriptor-loader tests fail on Node 25.9.0; all four pass on the CI version. - The runner package contract passes: 10 tests. - Full local typecheck stopped with exit 137 in the database package under the container's 4 GB memory limit. The production build reached the runner Rust build, then stopped because `cargo` is absent. - The full stable local Vitest run was stopped after all current-head CI test shards passed. It did not complete locally. The 180 focused tests listed above passed. - `git diff --check` passes. The branch changes 20 files against master. It has no lockfile or workflow changes. - The original author verified all three platform digests against registry integrity and ran the Linux executable. Its version was `2.1.286 (Claude Code)`. See #14816 for that evidence. - Greptile reviewed head `6db3d3f1` and gave 5/5 with zero comments. Both Superagent scans and Commitperclip pass. All current-head CI jobs pass, including build, typecheck, Rust, test shards, browser tests, and the canary dry run. ## Risks - The controller and provider pack must use matching runtime pins. Deploy them together. - CI owns `pnpm-lock.yaml`. The master lockfile refresh must resolve the SDK override. Refresh the Daytona lock digest with that lockfile. - Images built with Claude Code older than 2.1.284 need a rebuild before the CLI path can use Sonnet 5.5. - The runtime remains at SDK 0.3.286. This PR does not take the later 0.3.287 patch. - A live Sonnet 5.5 session and a Daytona image build are not part of the local verification. ## Model Used - Original work: Anthropic Claude Code, `claude-sonnet-5-5`. Review: `claude-opus-5-5`. The author reported `xhigh` effort, tool use, and code execution. The original context window was not reported. - Merge repair and PR preparation: OpenAI Codex, based on GPT-6, with tool use and code execution. The runtime does not expose the exact model identifier or context window in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Squash Attribution Keep these trailers in the squash commit to preserve the original author and AI attribution: ```text Co-Authored-By: Claude Code (Ivan) <SkilLab-Tech@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> ``` --------- Co-authored-by: Claude Code (Ivan) <ivan@skillab.com.br> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a65ca09508 |
fix(runner): settle accepted results after shutdown failures (#15217)
## Thinking Path > - Paperclip manages AI agents and the tasks they perform. > - The native runner saves tool results and completion reports before it releases a session. > - Large project discovery responses can exceed the durable command limit. > - A shutdown failure can leave a saved answer waiting for workspace repair. > - Recovery reused the old assessment for a different status decision, which violated a database constraint. > - This pull request bounds discovery responses and lets recovery commit the saved result after workspace repair. > - The benefit is a task that reaches its correct final status without another provider turn. ## Linked Issues or Issue Description **What happened?** A native run can save its final answer, fail during shutdown, and leave the task In Progress after workspace repair succeeds. Reconciliation tries to reuse the failed-workspace assessment for a new decision. The one-decision-per-assessment constraint rejects the write. Replaying the old decision can also retain a fresh coordinator lease. Separately, retained session cleanup only recognizes the old `adapter_failed` label. The project-list tool returns full project records, including large descriptions and workspace configuration. A large response exceeds the runner's durable command limit. The settlement diagnostic previously recorded only a failure flag. **Expected behavior** Project discovery stays within the command limit. Recovery finishes the saved result after workspace repair, releases its lease, and preserves the original error for inspection. It does not repeat provider work or relax session ownership checks. **Steps to reproduce** 1. Return large project records from `list_projects` and observe an oversized semantic result. 2. Persist an accepted native completion result, then record a shutdown failure. 3. Finalize with a failed workspace, repeat that attempt, then record successful workspace repair. 4. Reconcile the run. Before this fix, the issue stays In Progress. **Paperclip version or commit** Reproduced on `a386a599983519eb1d399f8b770bfccdb2a74762`. **Deployment mode** Self-hosted server with Paperclip Runner. Related transport work: #12208 drains queued events; #12241 resumes interrupted semantic calls. This change addresses bounded project discovery and accepted-result finalization. ## What Changed - Read bounded project summary projections from the database and return at most 50 authorized summaries with a continuation cursor and explicit description truncation. Agent and run trust boundaries narrow the database candidates; project-specific policies still receive full authorization. Only visible projects determine continuations. The default project-list API remains unchanged. - Record bounded, content-free settlement failure causes for command limits, storage errors, and rejected dispatch. - Include workspace state in assessment identity. Preserve the initial assessment for interrupted finalization, and commit replacement assessment and decision references together. - Release the coordinator lease when an existing decision is replayed, without repeating its effects. - Clear stale errors when recovery succeeds and retain them in `recoveredExecutionFailure`. - Accept both current and legacy close-failure labels in the existing exact-state cleanup path. - Add regression tests and update the tool contracts and recovery documentation. ## Verification - Red: the new project paging, settlement diagnostic, current cleanup label, and repaired-workspace regressions failed on the original implementation. - Green: protocol/catalog/tool checks (117 tests), the full project-tool and finalizer suites (47 tests), cleanup ownership cases (70 tests), and cleanup sweep cases (4 tests) pass. Database-backed pagination covers large descriptions/configuration, complete enumeration, uppercase cursors, agent/run restrictions, project-policy scope contributions, and identical results/cursors when hidden projects are added. - The initial CI failures in interrupted Board waits, contended endpoint proof, and semantic schema validation were reproduced and fixed; all affected cases pass locally. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — started before the review corrections; it spanned several source revisions and was stopped after reporting old-behavior and timing failures. It is not claimed green. Fresh project and recovery suites pass; the recovery suite also passes all 25 cases with the broad runner’s isolated home/config. The Slack timing case passed in isolation. Latest-head CI is the authoritative complete test matrix. - Existing authorization suite — 66 tests passed. - Latest-head CI on `8e5763d915aa6f375bdab6601996899ea01496fc` — 55 checks passed, 4 intentionally skipped, no pending or failing checks. - Greptile — 5/5, zero unresolved threads on the same commit. - `git diff --check` — passed. ## Risks - `list_projects` now returns summaries. Callers must follow `nextCursor` and use the authorized project API for full records. Candidate narrowing is only an optimization: project policy and responsible-user authorization remain authoritative. - Assessment identity changes for workspace finalization. Existing evidence remains intact; no migration is needed. - Cleanup still requires matching identities, settled tool evidence, and verified process ownership. Unknown tool outcomes remain blocked from session reuse. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and GitHub tools. The exact serving model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a386a59998 |
Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native agents receive task constraints and completion tools from Paperclip. > - Completion tools already define the procedure for reporting a result. > - Repeated procedure text adds instructions to each full task turn. > - The final reply must still explain a blocker and link a saved document. > - This pull request removes repeated procedure text and keeps these visible outcome requirements explicit. > - A document receipt supplies the exact link, and stricter evals check the persisted reply and browser navigation. ## Linked Issues or Issue Description Refs: #14961. Related: #14948 and #15007. **What happened?** Native task envelopes repeat completion procedure text. A reduced envelope needs explicit final-reply requirements. The `write_document` receipt also lacks a canonical document link. **Expected behavior** Keep the completion tools as the source of procedure details. Require one accepted completion result before the final reply. A blocked reply must explain the reason, owner and unblock action. A document reply must contain a working link to the saved document. **Steps to reproduce** 1. Run the native assigned-skill document case and native blocker case. 2. Inspect the run-attributed provider final and its persisted comment. 3. Check the blocker explanation or open the final reply's document link. ## What Changed - Remove repeated completion procedure text from the native task constraints and backend instructions. - Keep explicit blocker and document-link requirements in full task turns. - Return a company/task-scoped `documentHref` from `write_document`. Preserve the link in the idempotent mutation receipt. - Repeat canonical links for this run's current saved revisions in accepted completion feedback. Give blocked providers final-response guidance for the cause, owner and unblock action. - Keep internal document/comment anchors when Markdown issue links load cached issue details. - Add a manual six-cell comparison suite with strict source, build, default-instruction and budget admission. - Capture eighteen shared runnerd RPC projections and six direct OpenCode HTTP projections across start, resume and continuation phases, using scripted local transports and no provider execution. - Apply v3 checks only to the manual instruction comparison; preserve v2 checks for the existing native completion suite. Check the actual persisted blocker reason and exact saved-document link. Click the rendered document link and check the original content marker in the classic document card or the new document tab. - Forward exact OpenCode finishing calls through the controller. Wait for acceptance, keep accepted feedback and concrete rejection text, and reject malformed responses. Preserve ordinary dynamic-tool response handling. - Settle the completion decision and tool response before mapping a racing idle/error/abort event or handling explicit close/interruption. Reject a concurrent finishing call before controller admission. - Add a provider-free regression through real runnerd, the OpenCode proxy and a fake provider. Reject the first completion, accept the corrected report in the same turn, and propose one result. - Keep all original verdicts unchanged. Treat replay under new checks as separate diagnostics. ## Verification - `pnpm -r typecheck` and `pnpm build` pass locally. - Native document-authority tests pass, including company/run authorization and idempotent replay. - Native runtime-context, backend and measurement tests pass. - Final-answer calibration, protocol scoring, source-admission and catalog tests pass. Wrong reasons, absent links and wrong link targets fail. - `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six single-attempt local cells with the declared models. - Exported `prepareNativeInstructionPreflight` then `verifyNativeInstructionPreflight` pass on this clean committed source. They build locally and make zero provider calls. - Corrective live confirmation is incomplete. Source |
||
|
|
eb049aebf2 |
feat(skills): let agents update company skills safely (#15049)
## Thinking Path > - Paperclip is an open source control plane for AI-agent companies. > - Company skills give agents reusable work instructions. > - Skill Studio can edit skill files and save version history. > - Agents can create a skill, but they do not have a first-class update tool. > - An agent update needs a version check and safe retry behavior to prevent lost edits. > - This pull request adds `update_skill` through the existing company skill file API. > - The change keeps company policy, version history, and audit records in one path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: the server API, shared validation, and runner tool catalog. **Problem or motivation** An agent can create a company skill but cannot update its `SKILL.md` through a first-class tool. An unguarded retry can also create duplicate versions or overwrite a newer edit. **Proposed solution** Add `update_skill` with a required current version ID and a retry key. Route it through the existing skill file API. Reject stale versions and changed-input retries. Save the version and audit event together. **Alternatives considered** A separate write endpoint would duplicate the Skill Studio mutation path and policy checks. This PR reuses that path instead. **Roadmap alignment** This work extends Skills Manager and Skill Studio, which are listed in `ROADMAP.md`. **Additional context** The tool accepts a complete `SKILL.md`, not a partial patch. Callers must read the current version before they edit it. ## What Changed - Add optional version and retry fields to the existing skill file update contract. - Add a guarded API update with a stable retry receipt and attributed audit event. - Add `update_skill` to native and semantic runner tool catalogs, with mode and policy gates. - Add unit, integration, protocol, and semantic-tool regression coverage. - Document agent use and extend the OpenAPI request contract. ## Verification - Focused tests and direct server and runner TypeScript checks passed before this PR. - `git diff --check` passed after the rebase onto `master`. - CI passed on the latest PR head, including the full test matrix, typecheck, and build. Local full typecheck and build stopped because `cargo` is not installed. The local full test run ended without a verdict. - No dedicated end-to-end eval scenario was added or run. The protocol coverage and semantic-tool test cover the new action deterministically. - Reviewers can read a skill version, call `update_skill`, repeat the same key, then try a stale version and a changed-input key. Only the first edit must create a new version. ## Risks - File writes and database transactions must stay in sync when a write fails. The integration tests cover failed writes and retry behavior, but CI must verify them on the PR head. - Existing Skill Studio callers do not send the new optional coordination fields. Their request shape remains valid. ## Model Used - OpenAI Codex CLI assisted with this change. The runner did not expose the exact model ID or context window. The agent used code execution and repository tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details; exact model ID and context window were not exposed) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused tests; full suite is pending CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (53 pass, 4 skip on the latest head) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b17019e14d |
fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its adapters supply task context and access to Paperclip skills and tools. > - The default hire manual and shared prompts also repeat general work procedures. > - Those procedures overlap with stock provider instructions and the Paperclip skill. > - Existing E2E fixtures supply a QA manual, so they do not qualify the production default. > - This pull request reduces the generic instructions and adds real default-hire coverage. > - The benefit is less competing guidance, with inspectable evidence for preserved skills and task context. ## Linked Issues or Issue Description Refs: #14920. That merged change preserves native Codex base instructions. This PR covers the default manual, shared legacy prompts, operational skill guidance, and the narrowly approved ACP skill-discovery/session-environment repair for measured delivery and credential-persistence failures. **What existing behavior does this improve?** New non-CEO hires without a custom bundle and legacy task/chat startup and continuation prompts. **Current behavior** The shipped default manual contains 602 words. Generic task/chat prompts and ordinary resume deltas repeat work procedures already available through the harness and Paperclip skill. **Proposed behavior** The default manual contains only the eight-word company identity. Shared startup prompts retain identity and connection guidance. Ordinary resume deltas retain current work context without the generic execution contract. **Reason and benefit** Let the stock harness guide general work. Keep Paperclip-specific capabilities and independently test default hires, skills, ordered comments, and chat restart. **Breaking changes** New default hires receive less guidance. Existing saved manuals, explicit custom bundles, CEO templates, and specialized wake contracts retain their behavior. The obsolete includeExecutionContract option remains accepted for source compatibility. ## What Changed - Reduce the default hire manual to one sentence. - Reduce shared task/chat defaults and remove the generic ordinary-resume contract. - Keep connection guidance, auth, skills, custom prompts, and specialized wake context. - Add credential-free instruction-boundary gates and 26 explicit Product E2E cells across eight legacy/native profiles, including two focused Paperclip-storage cases. - Capture public hire receipts before providers run, then grade delivered prompts and independent task/chat outcomes. - Add an early legacy skill API recipe for saving a task document, checking the saved revision receipt and linking the document. Improve stock task/heartbeat skill-selection metadata and show a clickable Markdown UI-link example. Keep native tool completion separate. - Advertise bounded routing descriptions and exact successfully staged SKILL.md paths in legacy ACP Claude; keep full bodies on demand and preserve remote path rebasing. - Remove only the provider environment from copied persisted ACP session records, while loading current run credentials and preserving all other options/conversation state. - Regenerate both capability metadata inventories and reject stale manifests/inventories before provider admission. - Publish the original reduction and focused skill-repair comparisons, preserving all failures, automatic recovery, cost coverage and limitations. ## Verification **Behavioral qualification remains pending.** Original legacy ACP Claude loses the issue document only in the reduced cohort beneath an unchanged credential failure. A source-backed diagnosis finds that neither ordinary assignment reads the staged operational skill, while the runtime persists provider environment in session state. The new common repairs expose skill metadata/path and omit persisted env; strict document and credential guards stay intact. [Inspectable diagnosis and retained hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md). Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb` incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged #14961/#15007). All 222 affected adapter tests, adapter-utils/E2E typechecks, and final 96 variant/grader/retry calibrations pass. The frozen historical comparator is `c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths identical, exactly two production instruction paths plus nine declared unit expectations differ. The operational skill/discovery/environment repairs, selected model/profile/task/core grader/auth/permissions/retry policy are identical. Both actual launcher prepare→verify admissions pass with zero providers. [Immutable manifest and exact receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json). One original legacy ACP Claude cell per variant is authorized, with enforced single campaign attempts, 12-minute deadlines and company/agent 1,000-cent hard stops; every product recovery run/cost is counted. Actual live outcomes are pending. Current normal CI has one failed server shard and failed aggregate verify under diagnosis; other normal gates including typecheck/build/Rust/all eight browser shards pass. Fresh review completed successfully; the valid historical startup/resume masking finding was fixed with per-invocation task/chat checks and strict complete-snapshot capture, calibrated and resolved. Prior heads, failures and campaigns below remain historical evidence, not checks on this repair head. - Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52 current-head checks pass with two intentional Storybook skips, including repository typecheck/test/build and the browser shard. Fresh Greptile is 5/5 with zero unresolved review threads. Exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained receipt verification, zero providers/source errors. Fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck and 26-cell stock discovery pass. Canonical contract/inventory checks and the later issue-derived reference calibration are retained; that reference-only follow-up is not live-qualified by earlier frozen runs. - Prior full repository typecheck/build passed. The complete local Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three timing failures. All three affected files passed unchanged narrow reruns; original failures remain retained. Current-head CI now passes the full general checks; the original local failures remain retained. - The original 24-pair default-manual/shared-prompt comparison has two new overall classic Claude/OpenCode document-delivery failures plus an additional legacy ACP Claude document loss beneath an unchanged credential-guard failure (not closed by later runs), two newly passing OpenCode ordered cases, seven unchanged failures and 13 unchanged passes. Equal 15/24 totals do not establish behavioral equivalence. [Complete original report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md). - The skill-only repair holds the eight-word manual/shared prompts and merged #14920 fixed. All four matched profile configurations and 203 fixture/behavior files match. Candidate `abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill sources against baseline `bc83fe030234439ac51279502a28803958963e2e`. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547) and [baseline workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047) each pass 571 exact-source prerequisites before providers; all eight cells clean up successfully. Failed campaigns publish successfully and remain failed. - Repair pairs: Claude original Fail → Pass; Claude explicit Pass → Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff worsens beneath the unchanged failing UI-link grade: baseline gives a clickable API URL, candidate gives a code-formatted path without an anchor. The request's usable-link wording is narrower in the UI-only oracle. [Complete repair report and safe projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md). - The subsequent narrow stock metadata/link correction has two matched Pass → Pass cases, zero new machine failures/passes and no pending pairs. Both original-case handoff links remain deficient: candidate uses a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the preserved original oracle only requires a durable document. Both explicit clickable UI-link cases pass revision/content/link grading. All four exact-source 587-check gates, single assignment runs and cleanup pass. This does not establish fix causality because baseline also succeeds. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401) freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374) freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only comparison with reduced manuals/shared prompts held constant, not a repeat of the historical-manual comparison. Only original and clarified explicit classic OpenCode cases are selected, two per variant/four expected turns. 8,242 other tracked files and both profile hashes match; protected workflows admit each exact source before credentials. [Complete qualification report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md). Candidate original loads Paperclip/reference before saving publicly; baseline original loads it after writing locally, then saves publicly within the same assignment. Reported cost totals are $0.0107824490 candidate / $0.0107909015 baseline, with unmetered runtime. The later reference-only issue-derived link correction is provider-free calibrated and **not live-qualified** by these frozen runs; no further paid runs. - Retained tool calls show the repaired original OpenCode assignment loads only its assigned output skill before writing locally. Operational Paperclip is first loaded during automatic disposition recovery; its early recipe is visible then, but it never saves the missing document. Explicit candidate loads Paperclip and reads the new reference before saving successfully. All nine actual runs are counted. Reported LLM totals are $0.3802537209 baseline and $0.4918990161 candidate; local runtime is unmetered. - Initial setup, packaging, cancelled/missing-cell recovery, callback test and relative-output attempts remain retained. No completed provider failure was rerun. Frozen measurement branches are unchanged by later canonical metadata maintenance. - Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`, and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly for paid execution; it is excluded from `--all`. Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` replays this PR on merged hiring #14985 (`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four explicit custom-CEO-bundle checks, minimal generic manual boundary, and both suites. The combined fixture catalog and hiring calibrations pass 67 assertions; exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained-receipt verification, zero providers/source errors, fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E typecheck and 26-cell stock discovery pass. Fresh current-head CI passes all 52 checks with two intentional skips, and fresh Greptile is 5/5 with zero unresolved review threads. The prior source-plan browser failure is retained: a deterministic process fixture replayed its last `fixture:plan` command on `chat_task_completed`, writing revision 2 with identical body after the approval handoff. This was not paid provider execution. Rebased current-head CI passes the same assertion without an old-head retry or a change to that browser fixture. The merged hiring change was measured separately on immutable matched unions, with this reduced/shared/operational context and native completion guidance held constant. [Complete original two-profile report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md): [candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208) / [historical baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463), 705 provider-free prerequisites each. Both pairs are unchanged Fail → Fail on the exact-five count, with six core delivery checks passing all four cells; 28 actual successful runs include eight automatic completion wakes, zero retries, four successful cleanups. Source-read coverage is uncomparable, actual model charges unknown. Separately versioned provider-free accounting remains analytical work; original verdicts are preserved. This does not rerun or qualify the completed default-manual or native campaigns. ## Risks - Legacy ACP Claude's additional delivery loss is not closed by any later matched run and blocks the no-extra-failing-behavior merge criterion. Legacy document delivery may have relied on the prior manual/shared prompts. The early skill repair improves Claude in one trial; the later OpenCode pairs pass in both variants and cannot establish causality or robust recovery. Both original-case links remain deficient beneath the storage-only grade. The later issue-derived reference correction has only provider-free validation. Native finish/block descriptions must not be supplied to legacy agents. - The comparison holds merged native Codex fix #14920 constant; it cannot measure that fix's before/after task performance. - These bounded skill/context/chat workflows do not measure general coding quality. Unrepresented providers remain unqualified. - Saved manuals and old Codex sessions are not automatically migrated. Codex through ACP still has a separate base-instruction follow-up. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution, and delegated PR/eval work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (relevant suites and all three unchanged narrow reruns pass; complete-run timing failures retained in Verification) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green on the new repair head (prior-head checks retained above) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the new repair head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dd868ed125 |
fix(runner): share native completion tool guidance (#14961)
## Thinking Path > - Paperclip manages AI agents and their work. > - Native Runner agents report completion through finish and block tools. > - The providers receive different descriptions for those tools. > - Completion guidance belongs with the tools that enforce the result. > - This pull request shares the descriptions and refreshes retained catalogs. > - A separate native suite checks completion and blocking on production defaults. > - Legacy agents retain their separate skill and API paths. ## Linked Issues or Issue Description Refs: #14920, #14948, #14985. **Current behavior** Native Codex and MCP bridges describe finish and block differently. Retained provider sessions can keep old descriptions. **Proposed behavior** Native providers receive the same finish and block descriptions. The descriptions cover report selection, validation feedback, returned outcomes, approval gates and final-answer timing. Retained native sessions refresh from v13 to v14. **Reason and benefit** Put the completion procedure next to its native tool. Preserve stock base instructions, schemas, permissions and terminal semantics. This PR now stands alone on master. It contains no reduced manual, shared prompt or operational-skill changes from #14948. ## What Changed - Add canonical native finish and block descriptions. Use them in direct Codex and both native MCP bridges. - Advance the native tool contract to v14. Cover old-v13 refresh without replacing task identity or prior history. - Check authenticated tool catalogs, provider start/resume frames and serialized daemon catalogs. - Add an independent, explicit-only native completion suite. Preserve the original assigned-skill durable-document journey. Pair it with a concrete whole-task blocker across Codex, ACPX Claude and OpenCode. - Verify the actual public production default bundle and budgets before execution. Require independent durable disposition, native result/terminal receipts and observable provider-final ordering. - Correct the blocker browser oracle to accept the requested explanation. Keep exact owner/action/scope checks. Calibrate positive, missing and contradictory replies. - Preserve only actual `tool_call` terminal names (`paperclip_finish` / `paperclip_block`) in the native compatibility run-log projection. Require the same named call ID through its finishing result; retain all other redaction boundaries. - Admit verified hosted shallow checkout/build hydration and bind the selected runnerd to exact source/archive/binary provenance. Hosted cells truthfully reuse the existing trusted build; local admission executes Rust calibration. Forward only public source/run identifiers through both launcher preflight subprocess paths. - Enforce single attempts in the launcher for opted-in fixtures. Keep ordinary retry policy unchanged. Run exact-source, credential-free admission before credential loading. ## Verification - Frozen candidate: `d6e59e4712a3158ab4cd7d58deff1389b4578c21`, based on master `59c07ede72dc08b8aba149a01cc11e0b7a204621`; historical descriptions: `e74ed61a69fbdd8b3a8f15dd6456bc3140246e33`. Exactly the five original native production files and six unit tests differ. Both carry identical corrected fixtures, strict named finishing-call grader, closed compatibility carrier and admission. Defaults, profiles/models/auth/permissions and manifest bytes match. - Actual launcher `prepareNativeCompletionPreflight` → `verifyNativeCompletionPreflight` admission passes on both exact refs with zero providers: candidate 132 / historical 127 selected TypeScript assertions, 128 Node calibrations and one Rust normalization calibration each; E2E typecheck, manifest checks, selected binary provenance and six-cell discovery pass. Each has 257 explicitly skipped unrelated assertions, not coverage. The credential-free environment calibration exercises both real prepare/verify subprocess options with public hosted identifiers and rejects credential/ambient overrides. Complete actual launcher prepare→verify also passes on both frozen refs with explicitly synthetic hosted metadata/verified archives, separately labeled as calibration rather than a trusted GitHub run. Exact framed provenance parsing and mock source identity are calibrated without relaxing the real verifier. - [Complete matched qualification report](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-master-qualification.md), [immutable manifest](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-manifest.json) and [closed retained audit/hashes](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-results/comparison.json) are inspectable. All six candidate cells pass; historical descriptions pass five. Paired outcomes: **zero new failures, one new pass (Codex blocker), five unchanged passes, zero pending pairs**. [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37098728980) and [historical campaign](https://github.com/paperclipai/paperclip/actions/runs/37098815696) each execute six original attempt-1 native runs, with no campaign retry and successful cleanup. Their trusted workflow revision is `215586d127e97c9301d86e769a39a15c13298ca2`, separate from measured source. [Candidate public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098728980-1/index.html) and [historical public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098815696-1/index.html) retain declared screenshots. - Independent candidate evidence agrees with all original grades: 51 strict native checks, 12 served-default/budget checks and 21 original skill/document checks pass. The historical Codex blocker saves the correct whole-task blocker but omits the required marker from its actual provider final and identical saved reply. This is not semantic-summary fallback. Its original browser/matcher failure stays retained; the additional native snapshot/grade and workspace before/after digest were never written and are not fabricated by the separate API/PRP audit. Historical Codex completion has one failed finish followed by success within the same native run; the public receipt records no failure reason. All twelve runs and their usage remain counted. Reported model-cost subtotals are $0.00421482 historical/$0.00437391 candidate; Codex/Claude zero entries have unknown billing type, actual invoices are unverified and hosted execution cost is unmetered. One matched trial supports no extra failure within these six cases, not broad statistical or coding-quality equivalence. - Initial hosted `e18c2cf9` / `459455ac` and subsequent `0a9c5a7` / `00a761b` cohorts each stopped before providers in all twelve cells. The latter failed a mocked-receipt unit test under ambient hosted metadata; all source/build proofs passed. [All twelve later setup receipts](https://github.com/paperclipai/paperclip/blob/402ee94c52273ad58de355ae9a7d562dd22f8101/doc/plans/2026-10-02-native-completion-qualified-hosted-setup.json) are retained. [Exact failed setup receipts](https://github.com/paperclipai/paperclip/blob/27653eb1a8f8ce839776d760f4563f672e5a706c/doc/plans/2026-10-02-native-completion-master-hosted-setup.json) and the original manifest remain intact. Local sandbox-denied loopback and stale anchor-expectation attempts are retained separately; unchanged appropriate assertions were corrected/admitted before paid dispatch. Old anonymous OpenCode streams are not assigned inferred tool names or retroactively passed. - Full provider-free E2E support previously passed 927 tests in 67 files. Exact-head d6 normal CI run `37098409915`, attempt 1 passes full repository typecheck/build/tests, Runner Rust/static checks, all browser shards/aggregate and canary: 52 check-runs pass, four intentional skips, Snyk passes. Fresh Greptile check `111132956342` is 5/5 with zero unresolved threads. Source-specific deterministic tests do not substitute for the bounded live comparison. - Earlier native source `9138f570c341c251a5727c32d6615ce238bc8e03` is archived. Its [complete reduced-manual-context report](https://github.com/paperclipai/paperclip/blob/9138f570c341c251a5727c32d6615ce238bc8e03/doc/plans/2026-10-02-native-completion-live-comparison.md) remains intact, including original failures, grader limits and provider-free replay. It is not current-master-context qualification. ## Risks Changed tool text can change model behavior. The completed six-pair qualification shows no extra failing outcomes in this bounded trial; other tasks and repeated-run variance remain unmeasured. Observable final ordering does not prove provider feedback consumption. Public evidence can fail closed if a provider does not expose the required result sequence. This slice does not remove native fixed prompts or measure general coding quality. No database, schema, permission or legacy completion changes occur. ## Model Used OpenAI Codex, GPT-6 family, with code inspection, execution and tool use. The exact deployment ID and context-window size are not exposed in this session. They are unavailable rather than inferred from the model menu. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ec82c5774 |
fix(runner): preserve task context when tool connections change (#14963)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use tools through company-scoped connections and provider sessions. > - Resolving a tool connection currently forces a fresh session even when the provider can load new tools into the existing conversation. > - A fresh provider conversation can receive too little history to continue the task. > - This pull request adds explicit tool-refresh capabilities and uses them in both runner paths. > - Fresh attempts receive bounded task history with source IDs and retrieval instructions. > - The benefit is that agents can continue the same task after a connection changes. ## Linked Issues or Issue Description **What happened?** A resolved tool connection forced a fresh provider conversation. The new conversation could lose the original goal and prior answers. Claude also rejected resume when only the MCP server set changed. **Expected behavior** Resume the provider conversation when its harness can refresh tools. When a fresh session is required, supply enough bounded history to continue the task. Preserve company, agent, task, workspace, model, instruction, and skill checks. **Steps to reproduce** 1. Start a conversation and agree on a task and its constraints. 2. Request and connect a tool needed for the task. 3. Continue the conversation after the connection resolves. 4. Check that the agent remembers the task and can use the new tool. **Paperclip version or commit** The bug was reproduced on master at `c46e41e81`. This branch is rebased on current master. **Deployment mode** Self-hosted server. Both legacy adapters and the native runner are affected. Related public work: Refs #13282 for task-backed conversations. Refs #13057 for the broader session-compaction proposal. Refs #14659 for another report about local CLI session continuity. This change fixes tool-connection continuation. Provider authentication repairs keep their existing recovery behavior. ## What Changed - Expose tool-refresh support in native harness descriptors and legacy adapter metadata. - Request tool refresh after connection resolution. Keep provider authentication repair as a fresh-session wake. - Reload current tools and credentials while retaining supported Claude, Codex, Grok, and other provider conversations. - Allow MCP-only changes during qualified native recovery. Keep all other compatibility checks. - Refresh managed-provider and ACPX tool bindings when attaching a new run. - Add a fresh-session handoff for both runner paths. Bound database reads, excerpts, and the final packet to 24,000 bytes. - Include the original request, recent messages, decisions, plans, prior answers, and source IDs. Mark omitted content. Apply reset boundaries, wake cutoffs, quarantine, and secret redaction. - Add regression tests and document the capabilities and handoff behavior. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. Rust formatting passed. - Final review fixes passed 314 server tests, 333 adapter utility tests, 139 native-session runtime tests, and 12 managed-provider Rust tests. They verify historical quarantine, raised budgets across attachment, no history reads on successful resume, and handoff delivery on fresh retry. - Broader branch verification also passed 1,401 adapter utility tests, 1,047 runner TypeScript tests, 43 Grok adapter tests, and 311 Rust core tests. - Live Claude CLI and Grok ACP probes preserved the provider session ID, recalled a prior task constraint, and called a newly added read-only MCP tool. - GitHub CI passed on `b21486d18084a7aa4cafbe8e012f7cad6585d9cc`: 55 successful checks and 4 skipped checks. This includes all test shards, all eight browser shards, runner checks, and the Grok clean public npm install canary. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37057514976). - A full local test attempt encountered a separate Git snapshot timeout. All affected local suites passed after the final edits, and the full CI test gates passed. - Review the capability matrix in `packages/paperclip-runner/README.md`. Repeat the four reproduction steps with a supported provider and with an unsupported harness. ## Risks - Provider tool refresh can fail. Existing recovery falls back to a fresh conversation where policy permits it. - A new transport can replace an old process while preserving the provider conversation. Tests cover current credentials and unchanged identity. - Long history can omit older context. Explicit markers and source IDs let the agent retrieve needed context within task scope. - Unknown and unqualified harnesses use the fresh-session path. No database migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and live provider testing. The exact model ID and context-window size are not exposed in this session. Claude and Grok also ran as test subjects. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d034ba7491 |
fix(interactions): derive question storage from canonical forms (#14946)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents request human input through durable issue interactions. > - A question form has a canonical presentation and a compatibility storage format. > - The creation API required agents to write both formats. > - Tool guidance told agents to split text and choice questions across those formats. > - This pull request accepts one complete canonical form and derives storage fields on the server. > - The benefit is a complete question card with stable answer and retry behavior. ## Linked Issues or Issue Description Related work: Refs #13630 and #14430. PR #13630 addresses the display of historical partial forms. This change fixes creation and keeps the check that rejects conflicting new forms. **What happened?** A question save supplied three compatibility questions and one canonical text question. The API correctly rejected the incomplete canonical form. The Runner's tool description encouraged this split. Sending only a complete canonical form also failed because the API required compatibility questions. **Expected behavior** An agent sends one complete `payload.questionSet` with every text and choice question. Paperclip derives `payload.questions` for storage and answer compatibility. Existing legacy requests remain valid. Explicitly conflicting dual forms remain invalid. **Steps to reproduce** 1. Call `paperclip_request_human_input` with `interactionKind: "questions"`. 2. Send `payload: { version: 1, questionSet: ... }` with a required text question and a required choice question. 3. The old API rejects the missing compatibility questions. With this change, it stores both questions and preserves the canonical form. 4. Retry with the same idempotency key. Confirm that only one interaction exists. 5. Submit both answers. Confirm that the normal resolver and continuation rules apply. **Paperclip version or commit** The branch is based on `cf8ad63c8`. The problem affects the native Runner and the interaction creation API. **Deployment mode** Server deployment with the native Paperclip Runner. Integration tests use the real interaction service and an embedded test database. ## What Changed - Add one shared canonical-to-storage projection. Reuse it for native harness question requests. - Accept canonical-only question creation at the shared validator and server boundary. - Export the input type and update the plugin SDK and its RPC contract. - Advertise a typed, complete question form in the live and scenario tool schemas. - Enforce canonical text and custom-answer constraints before ordinary or native resolution. Preserve harmless display whitespace. - Run regex matching in isolated workers with a deadline and resource limits. Both answer paths await the result before persistence. Saved native delivery uses the validated answer without taking another worker slot. - Update agent guidance and generated Runner contracts. - Test mixed forms, option-ID collisions, retries, answers, legacy requests, and conflicting forms. ## Verification - Interaction service, HTTP route, native bridge, and Runner authority suites: 221 tests passed after correcting an obsolete tool-description assertion. - Shared validator, plugin SDK, CLI, and UI compatibility suites: 67 tests passed. - Runner core tool-contract suite: 20 tests passed. AJV validates live and scenario schemas. - Final review regressions: 172 shared, service, native bridge, and authority tests passed. These cover text length, pattern, numeric limits, whitespace, custom option IDs, and historical pending cards. - Runner session suites: 67 tests passed. Published example tests: 4 tests passed. - Server typecheck and the shared/server builds passed after the compatibility fixes. - Final delivery verification: 35 response-delivery tests passed. The native delivery regression proves saved answers do not enter pattern workers; server typecheck and build passed. - Pattern security and answer-flow verification: 205 tests passed after repairing the child fixture loader. These cover pathological matching, event-loop responsiveness, worker concurrency, slot cleanup, HTTP routes, native delivery, and the full helper in a child process. - `pnpm -r typecheck` passed on the bounded-worker revision. - `pnpm build` passed on the bounded-worker revision. - All 55 GitHub checks passed on `fe457af`; four optional jobs were skipped. An unchanged Cursor adapter test timed out once in CI, passed locally, and passed on one failed-job rerun. - Reviewers can send the canonical-only mixed form above and verify that the saved interaction contains both canonical and compatibility questions. ## Risks - The creation API accepts a new input shape. Stored rows and answer contracts keep the existing shape. - The shared projection must preserve synthetic free-text option IDs. Collision and native round-trip tests cover this behavior. - Historical partial rows remain readable. New conflicting dual forms, including written-answer mismatches, remain rejected. - Existing pending cards retain the written-answer paths offered by their stored options. Canonical text constraints still apply. - Ordinary answers now enforce declared canonical constraints before persistence. Invalid answers leave the card pending. - Regex validation has a one-second deadline and a four-worker capacity limit. A complex pattern or capacity error leaves the card pending with a validation error. - No database migration or change to company authorization is required. ## Model Used - OpenAI GPT-6 through Codex. The session exposes the GPT-6 model family; its exact runtime model identifier and context window size are not exposed. Used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9786f6df56 |
fix(runner): preserve credential content in document saves (#14937)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner sends authorized tool calls to the control plane. > - Agents use these calls to save plans and instruction files. > - The Runner used diagnostic secret detection to reject execution arguments. > - Ordinary credential-related prose could reject a document save before persistence. > - This pull request forwards the original arguments and leaves credential policy to the provider harness. > - The benefit is reliable saves with useful diagnostic records. ## Linked Issues or Issue Description Related foundation: Refs #12415 and #14430. No duplicate save-policy fix was found. **What happened?** A `write_document` call failed before the server saved its plan. The Runner reported `semantic tool input contains credential material; refusing to execute altered arguments`. The detector also masked ordinary phrases such as `secret manager` and `credential handling` in diagnostics. Both TypeScript dispatchers had equivalent execution gates. One dispatcher also rewrote structured approval and question payloads before execution. **Expected behavior** Paperclip forwards authorized arguments unchanged. The provider harness decides credential-content policy. Log and audit redaction does not reject or rewrite save input. **Steps to reproduce** 1. Send an authorized `write_document` call with a plan that discusses credential handling. 2. Include an intentional credential value in the body to exercise harness-owned policy. 3. The old Runner rejects the call. With this change, the document service stores the exact body. 4. Diagnostic records still mask explicit credential values. Qualified credential fields, short bearer values, opaque diagnostic pairs, and valid encoded JSON token headers have regression coverage. **Paperclip version or commit** Reproduced at `c46e41e81c03cd3c8b64cf993615b604d7fe8c62`. The branch is based on current `master`. **Deployment mode** Server deployment with the native Paperclip Runner. Local regression tests use the real document service and an embedded test database. ## What Changed - Remove credential-content vetoes from Rust admission and both TypeScript semantic dispatchers. - Preserve original structured approval and question arguments during execution. - Keep transport bounds, schema checks, authorization, idempotency, and audit masking. - Require explicit credential syntax or recognized formats for diagnostic masking. Preserve ordinary prose, metadata, and dotted identifiers. - Test exact document persistence, replay, nested argument identities, and masked audit copies. - Remove obsolete retry guidance and document harness-owned credential policy. ## Verification - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --lib --test acpx_event_payload --test acpx_provider_state --test acpx_provider_turns`: 355 tests passed. - Focused server and adapter tests: 189 tests passed after rebase. These include the real document save and the complete tool-gateway suite. - Semantic dispatcher and conformance tests: 34 tests passed. - Diagnostic redaction and MCP tests: 46 tests passed, including all six review examples. - `pnpm -r typecheck` and `pnpm build` passed on the repaired branch. - The broad local root suite was interrupted after database fixture setup failures. The focused database suites passed. CI runs the complete configured test lanes. ## Risks - Authorized tool arguments can intentionally contain credentials. The harness must enforce its content policy. - Diagnostic detection is narrower. Explicit assignments, credential fields, and recognized credential formats remain masked. - The change does not add a database migration or change company authorization. ## Model Used - OpenAI GPT-6 through Codex. The session exposes the GPT-6 model family; its exact runtime model identifier and context window size are not exposed. Used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cf8ad63c80 |
build(deps-dev): bump tsx from 4.23.12 to 4.23.15 (#12965)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.23.12 to 4.23.15. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/privatenumber/tsx/releases">tsx's releases</a>.</em></p> <blockquote> <h2>v4.23.15</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.14...v4.23.15">4.23.15</a> (2026-09-20)</h2> <h3>Bug Fixes</h3> <ul> <li>exclude bare builtins from namespace inheritance (<a href="https://github.com/privatenumber/tsx/commit/38e158857e50bca311be7c232a5057e5a2e5347a">38e1588</a>)</li> <li>expose require.cache and require.extensions to tsImport CommonJS modules (<a href="https://github.com/privatenumber/tsx/commit/2da34075afaed43e2b7fd0aca5fbebaaf337ff3a">2da3407</a>)</li> <li>make namespaced register() overloads portable for declaration emit (<a href="https://github.com/privatenumber/tsx/commit/562c434a5c8695e74327bbeb51cfeb9b86fc7e15">562c434</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.15"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.14</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.13...v4.23.14">4.23.14</a> (2026-09-20)</h2> <h3>Bug Fixes</h3> <ul> <li>restore the CJS bridge namespace for Node 24 require(esm) under tsImport() (<a href="https://redirect.github.com/privatenumber/tsx/issues/802">#802</a>) (<a href="https://github.com/privatenumber/tsx/commit/6e5236b065738d3687a06396d064774cfede390f">6e5236b</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.14"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.13</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.12...v4.23.13">4.23.13</a> (2026-08-30)</h2> <h3>Bug Fixes</h3> <ul> <li><strong>cache:</strong> bound shared transform cache memory (<a href="https://redirect.github.com/privatenumber/tsx/issues/835">#835</a>) (<a href="https://github.com/privatenumber/tsx/commit/28e1f12d04cd2afe1db17f8555b14fe5fb567c6e">28e1f12</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.13"><code>npm package (@latest dist-tag)</code></a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/privatenumber/tsx/commit/ca66105a17a2a4c6503fe3a12b5b9ec408286011"><code>ca66105</code></a> test: fix drive-less file URLs in ESM resolver fixtures</li> <li><a href="https://github.com/privatenumber/tsx/commit/2da34075afaed43e2b7fd0aca5fbebaaf337ff3a"><code>2da3407</code></a> fix: expose require.cache and require.extensions to tsImport CommonJS modules</li> <li><a href="https://github.com/privatenumber/tsx/commit/38e158857e50bca311be7c232a5057e5a2e5347a"><code>38e1588</code></a> fix: exclude bare builtins from namespace inheritance</li> <li><a href="https://github.com/privatenumber/tsx/commit/562c434a5c8695e74327bbeb51cfeb9b86fc7e15"><code>562c434</code></a> fix: make namespaced register() overloads portable for declaration emit</li> <li><a href="https://github.com/privatenumber/tsx/commit/edfb1f05a3f40b879a41a03a0801c2abd3a3ecf9"><code>edfb1f0</code></a> build: upgrade pkgroll and externalize CJS loader reference</li> <li><a href="https://github.com/privatenumber/tsx/commit/70e78284837c859f09b96cd10cd71d007aa4b795"><code>70e7828</code></a> test: upgrade tinyspy for disposable API</li> <li><a href="https://github.com/privatenumber/tsx/commit/9ed2022dfa9ea1be9511fe6abcde8110c25055a7"><code>9ed2022</code></a> ci: avoid duplicate release notifications</li> <li><a href="https://github.com/privatenumber/tsx/commit/872e77ffc5e96ca5c4727e74c0694debcb26219b"><code>872e77f</code></a> refactor: use disposables for cleanup</li> <li><a href="https://github.com/privatenumber/tsx/commit/6e5236b065738d3687a06396d064774cfede390f"><code>6e5236b</code></a> fix: restore the CJS bridge namespace for Node 24 require(esm) under tsImport...</li> <li><a href="https://github.com/privatenumber/tsx/commit/28e1f12d04cd2afe1db17f8555b14fe5fb567c6e"><code>28e1f12</code></a> fix(cache): bound shared transform cache memory (<a href="https://redirect.github.com/privatenumber/tsx/issues/835">#835</a>)</li> <li>See full diff in <a href="https://github.com/privatenumber/tsx/compare/v4.23.12...v4.23.15">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
ec3bacc9bd |
fix(chat): hide ignored provider information (#14929)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task and agent chats show agent progress and problems that need attention. > - Codex also sends account, skill, and unrelated thread notifications. > - The runner correctly ignores that information but reports it as a warning. > - Chat then shows an internal diagnostic as an actionable provider notice. > - This pull request keeps the diagnostic in run logs and removes it from chat. > - Real provider warnings, errors, and agent replies remain visible. ## Linked Issues or Issue Description **What happened?** Chat showed “Received a provider update” and a warning with the text “ignored unrelated provider information”. Its details said “User Actionable: Yes” even though no user action was needed. Saved conversations retained the same noise. **Expected behavior** Keep ignored provider information in the run log. Do not show it as chat activity or a user warning. Preserve real warnings and errors. **Steps to reproduce** 1. Start a conversation with the native Codex runner. 2. Have the provider send an account update, skill change, or unrelated thread notification during the turn. 3. Inspect live chat and reload its saved history. The regression tests also reproduce the old stored notice without a live account. **Paperclip version or commit** Source implementation on master at `e00d10d5d`. The duplicate search found no open PR for this fix. Related prior work: #13109 improved provider-notice presentation. #12367 added Codex thread normalization. This change addresses the internal information that those paths still projected as chat warnings. **Deployment mode** Native Paperclip Runner with the Codex app-server provider. The issue was seen in hosted chat and can be reproduced with local provider fixtures. ## What Changed - Map ignored unrelated Codex information to `harness.diagnostic` in the Rust and TypeScript normalizers. - Retain a bounded allowlist of redacted provider method and thread/turn identifiers. - Use the same Unicode character limit and truncation marker in both normalizers. - Share the text redactor through a pure helper. Keep provider connection code out of the standalone demo's source closure. - Omit that diagnostic and the matching legacy notice from live chat. - Omit the matching legacy notice from saved chat history. - Test diagnostic retention, account-notification integration, live and saved chat, and continued visibility of real warnings, errors, and replies. - Document the local run-log event and historical display behavior. ## Verification - Passed: 68 tests in the two affected UI transcript suites. - Passed: 60 TypeScript tests across provider events, transport behavior, and the standalone demo boundary. - Passed: 13 Rust provider-event tests and the Codex account-notification integration test. - Passed: `pnpm check:token-gates` and Cargo formatting checks. - Passed: full `pnpm build` and `pnpm -r typecheck`. After the review fix, the provider package build, typecheck, and both provider-event suites passed again. - Full local `pnpm test:run` failed: 608 files / 10,904 tests passed, 30 server suites failed, and 104 files / 4,012 tests were skipped. Most failures were embedded PostgreSQL startup errors. Two tests timed out in `heartbeat-comment-wake-batching` and `workspace-git-snapshot-streaming`. PostgreSQL startup also failed in `heartbeat-run-event-sequencing` and `native-finalization-migration`. These server files are unchanged by this PR. Isolated heartbeat reruns were skipped locally. The stable test script stopped after this general-server group, so later groups did not run locally. - The original review thread is resolved. Greptile is 5/5 on current head `683dab7cce57187c57e84c83f5e9da4ad75c9c04`. - All current-head CI gates passed, including the full server/chat/workspace test matrix, Rust and TypeScript runner suites, browser E2E, build, typecheck, and release canary. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37021330663). - Replay the exact old warning in either transcript adapter. It must produce no chat row. A genuine provider warning or error must still produce a row. ## Risks - Low risk. The display filter matches one diagnostic code or the complete legacy warning shape. Other provider notices remain visible. - New ignored-information events use the existing harness-diagnostic event type. They retain diagnostic evidence without original account payloads. - No database migration, API permission, provider execution, or recovery behavior changes. This affects the local run log, not Telemetry or OpenTelemetry exports. ## Model Used OpenAI Codex, GPT-6. The exact backend model ID and context-window size are not exposed in this session. Used reasoning, repository inspection, code editing, tool use, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected tests locally and they pass (the broad local run has PostgreSQL startup errors and timeouts documented above; the full CI matrix passed) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
408f70e69f |
fix(runner): preserve stock Codex base instructions (#14920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner connects Paperclip tasks to Codex app-server. > - Paperclip passed its runtime context as `baseInstructions`. > - That field replaces the stock Codex base prompt. > - This pull request sends the same Paperclip context as additive developer instructions. > - Codex keeps its stock prompt and still receives Paperclip task instructions and tools. ## Linked Issues or Issue Description **What happened?** The native Codex driver and Rust provider sent Paperclip context through `baseInstructions` on thread start and resume. Codex used this text in place of its stock base instructions. Direct-chat resume also sent an empty replacement base. The Runner Lab session path used the same replacement field. **Expected behavior** Codex should retain its stock base prompt. Paperclip should add its runtime context through `developerInstructions`. Other provider facades should retain their current instruction handling. **Steps to reproduce** 1. Create a native Codex session through Paperclip Runner. 2. Inspect the `thread/start` request in the native provider trace. 3. Resume the session and inspect `thread/resume`. 4. Before this fix, these paths set `baseInstructions`. After this fix, the Codex paths set `developerInstructions` and omit `baseInstructions`. **Paperclip version or commit** Reproduced against master at `cad26c6bfb736039c8ed5743da650a44792a083c`. **Deployment mode** Built from source. Native Codex app-server and runnerd paths. A local protocol probe used codex-cli 0.153.4 and a localhost Responses stub. No duplicate fix or matching public issue was found in the GitHub search. ## What Changed - Send additive developer instructions on Codex start and resume in the TypeScript driver, Rust provider, and Runner Lab session path. - Carry the additive fragment through runnerd, including runtime asset path mapping. - Preserve existing instruction fields for other provider facades, including OpenCode. - Add start/resume/direct-chat regression coverage and check the actual Rust provider request. - Document the historical option and trace field names. Record progress and follow-ups in the working checklist. ## Verification - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - Targeted Codex driver lifecycle, driver, and live-session Vitest suites — 139 tests passed. - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --test codex_provider` — 91 passed, 2 ignored subprocess helpers. - Real app-server probe: a localhost Responses stub captured identical 14,732-character stock base instructions on fresh start and cold resume. Both requests retained the Paperclip marker in developer input. Both stub turns completed. No paid inference was used. - Runnerd transport Vitest suite — 182 tests passed. - The initial `pnpm test:run` attempt reported local dependency-loading, embedded PostgreSQL startup, and macOS `/var` versus `/private/var` path failures. It was stopped after those failures. Loading-suite reruns passed 1,428 tests after the build; native interaction/finalization reruns passed 38 tests. A seven-suite diagnostic rerun passed 463 tests and isolated the remaining path and PostgreSQL setup failures. - With `TMPDIR=/private/tmp`, workspace, gateway, interaction, and attachment suites passed all 356 tests. The remaining environment-image and native-session-resumption suites passed all 44 tests with the same canonical temp path. All affected suites passed on rerun. The original full local command was stopped after failures and is not claimed as passing. - All 55 PR checks passed at `83281439456181396f3707eecda5d2ebc90bd14d`. Greptile scored 5/5 with no open review threads. - No paid live campaign or Product E2E browser suite was run. This change has protocol and regression coverage; it does not claim improved task quality. ## Risks - Stock Codex behavior may differ from behavior under the previous Paperclip replacement prompt. Restoring that behavior is the intended change. - Existing Codex threads retain their saved replacement base prompt. They need a provider session reset to receive the stock base. This PR does not reset active sessions or alter recovery rules. - The legacy `baseInstructions` option and trace field names remain for compatibility. They now describe the additive Paperclip fragment for Codex. - The separate Codex-through-ACP dependency patch remains a follow-up in the harness coverage checklist. This PR covers native app-server execution. ## Model Used OpenAI Codex, GPT-6. The exact runtime model variant and context window are not exposed in this session. Used reasoning, repository inspection, code editing, shell execution, and test tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018993140f |
feat: let agents name prompt-only tasks (#14761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cbd278dc03 |
fix(interactions): derive chat recipients and validate explicit users (#14742)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agents use saved questions to get human input and continue the same task. > - The standard question example recently told models to copy a user ID. > - A model can omit an identity prefix and create a question its intended recipient cannot answer. > - Agent Chat already knows the conversation owner, so the server can supply that identity. > - This pull request removes the blanket instruction and validates explicit recipients before saving. > - Ordinary questions stay simple, and explicit addressing remains available for decisions that need a particular person. ## Linked Issues or Issue Description Refs #14707, #14188. Related: #14238 handles legacy email recipients; this change prevents invalid recipients in new cards and retains exact ID matching. **What happened?** A model copied a Cloud user ID without its prefix into `addresseeUserId`. Creation succeeded. The intended user's answer then failed the exact recipient check. **Expected behavior** Ordinary chat questions use the saved conversation owner. A task may optionally name a specific recipient. The API rejects an unknown or unauthorized recipient before it creates a card. **Steps to reproduce** Create a chat question for a user whose ID is `paperclip-id:example`. Supply `example` as the addressee. Before this change, creation accepts the invalid recipient and the owner cannot answer. With this change, creation returns 422. Omitting the field saves the full owner ID and allows that owner to answer. ## What Changed - Remove `addresseeUserId` from standard question examples and remove the blanket requester-ID instruction. - Derive the recipient of ordinary chat questions from the persisted conversation owner. Reject conflicting explicit user IDs. - Keep explicit task recipients optional. Validate supplied user IDs with the existing board mutation policy, including company, viewer, and Cloud restrictions. - Preserve explicit agent routing, connector intents, confirmations, exact recipient checks, idempotent retries, and no-login local-board authority in local-trusted mode. - Update the blocker grader to accept an omitted recipient and verify the actual requester answered. - Add database and HTTP tests for prefixed identities, denied recipients, concurrent retries, saved answers, and response delivery. ## Verification - Database interaction service suite: 90 tests passed, including implicit local-board creation/answering and authenticated/Cloud denial. - Interaction HTTP route suite: 84 tests passed. - Affected interaction/native/connector/documentation suites: 231 tests passed across six files after valid-user fixtures were updated. - Resolver and interaction unit suites: 29 tests passed. - Product E2E unit/calibration suite: 793 tests passed; Product E2E typecheck and blocker catalog discovery passed. - Generated API-reference and capability contract checks passed. - `pnpm -r typecheck` and `pnpm build` passed. - Full local `pnpm test:run` did not finish green: its initial general-server pass had 14,416 passing assertions, one unrelated native-resume assertion failure on macOS, and three teardowns from an intermediate fixture cleanup fixed above. Separate broad local groups also encountered timeout/live-port failures under host load. Local UI (7,026), CLI (502), shared (817), and skills-catalog (20) tests passed; the complete final-head CI matrix is the broad verification gate. - After two CI cold-start readiness timeouts, a separate test-only commit gives the first exposure lifecycle fixture the existing normal 30-second readiness budget. Its real HTTP, ordering, and cleanup assertions remain intact; the targeted case and final Linux CI shard passed. Production deadlines are unchanged. - A separate OpenCode fixture failed twice on GitHub-hosted Ubuntu because its cached Node executable was group-writable; the same case passed on AWS runners. The fixture now qualifies its own Linux copy with mode `0500` and the actual copy digest. Host files and production security checks are unchanged. The focused macOS case passed; the new Linux-copy branch also passed on the final AWS-hosted Linux runner (1,125 passing Runner tests, 3 skipped). The final run was not on a GitHub-hosted runner. - Final-head [CI run 36762078176](https://github.com/paperclipai/paperclip/actions/runs/36762078176) passed for `116b968b24fa0a8c5724a7bf96e73a8dda5f0425`: 54 successful checks and two conditional Storybook skips, with no pending or failed checks. The 27 general/serialized test jobs reported 28,635 passing tests. Typecheck, build, Runner, browser E2E, and Canary gates passed. Greptile reviewed that exact head at 5/5; both review threads are resolved, with no open follow-ups. - No live provider replay is claimed by this PR. ## Risks - New explicitly addressed cards reject users who cannot mutate the issue, including viewers, inactive members, and invalid IDs. Callers that supplied invalid recipients must correct their request. - Existing addressed cards are not rewritten. Existing authorization checks remain strict. - Chat inference applies only to questions without an agent addressee. Connector intents and governed confirmations retain their own recipient paths. - No schema change or migration is required. ## Model Used OpenAI Codex, GPT-6 (exact serving variant and context window are not exposed in this environment). Used reasoning, tool use, code editing, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
94e8dec56b |
fix(runner): preserve tool outcomes through shutdown and restart (#14734)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner sends authorized tool calls to the server and saves their results. > - A provider turn can stop while a server write is still running. > - The old shutdown path invented a failed result that could conflict with the real result. > - Truncated execution input and incomplete recovery records made the failure harder to diagnose. > - This pull request preserves exact inputs and actual outcomes through shutdown and restart. > - Tests force the race and crash boundaries so safe retries do not repeat writes. ## Linked Issues or Issue Description **What happened?** Stopping a turn during a server tool call could record a false failure, then reject the actual result as a conflict. The diagnostic input formatter could truncate instruction content before execution. A crash during saved-result delivery could leave that delivery permanently indeterminate. Cleanup could hide the first failure, and a retry could overwrite earlier run logs. **Expected behavior** Keep dispatched tools pending until their actual result is known. Preserve accepted input bytes. Accept identical result delivery without failing the task. Reject conflicting results with enough evidence to diagnose them. Recover saved-result delivery without repeating the business operation. **Steps to reproduce** 1. Hold an instruction update at the filesystem commit barrier. 2. Stop its provider turn before the server returns the result. 3. Release the write, deliver its result, and replay the same result. 4. Repeat with a restart before and after the delivery receipt is saved. 5. Check that there is one write and one audit row, and that the exact result survives. **Paperclip version or commit** The change was developed from `44736c9c7` and rebased onto `0e5830887`. **Deployment mode** Self-hosted server with the native runner. Tests use local runner processes, scripted providers, and PostgreSQL. Related work: #12353 added durable semantic tool receipts; #12384 added durable Codex tool recovery; #12404 bound semantic tools to ACPX sessions. #14633 covers separate native-provider cancellation and qualification work. This PR addresses server semantic-tool outcomes and their durable delivery. No duplicate fix was found. AgentMail discovery is outside this PR. ## What Changed - Close turn admission without inventing results for dispatched tools. Keep pending calls and accept late actual results. - Accept identical result replay with a diagnostic warning. Include call identity and both result hashes in real conflict errors. - Preserve exact execution arguments. Reject prohibited or oversized input before dispatch. Keep diagnostic previews redacted and bounded. - Commit instruction-attempt evidence before the filesystem write. Save completed mutation receipts so concurrent and restarted duplicates return the first result. Recheck authorization before replay. An attempt without a completed result stays unknown and cannot execute again. Definite pre-write failures save and replay their original error without another write. - Recover an interrupted saved-result delivery only for backends with durable result receipts. Never replay an ordinary business operation with an unknown outcome. - Preserve the initiating error when cleanup also fails. Record incomplete settlement evidence. Propagate typed unknown-outcome errors through the native tool wrapper without creating a false completed tool result. - Append run-log attempts and restore the durable log before appending after local file loss. Reject incomplete restores. Publish a restored prefix only if the destination is absent so concurrent attempts cannot overwrite new lines. - Add deterministic race, crash, replay, authorization, exact-content, and log-restoration tests. Document their assertions in `packages/paperclip-runner/docs/durable-recovery.md`. ## Verification - Current head: `7e088f4c7fba8ebabf98ae95485a5753b013d489`. All 55 applicable checks pass; four conditional/manual checks are skipped. This includes build, typecheck, Rust, both runner TypeScript shards, server and workspace tests, all eight browser shards, isolated runner compilation, and the clean-install release dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/36746101110). - Greptile reviewed this exact head at 5/5 with zero new findings. All three earlier review threads are resolved. - Focused local verification includes 11 instruction integration tests, 23 surrounding authority/tool tests, 26 run-log tests, and 169 controller/driver tests. The post-rebase controller/transport/runtime selection passed 415 tests. The full Rust release suite passed 617 tests with two ignored. The real-process SIGKILL recovery test passed three consecutive runs. - The fault matrix in `packages/paperclip-runner/docs/durable-recovery.md` uses explicit barriers, real PostgreSQL rollback, durable journal reloads, and killed runner processes. It covers late results, identical and conflicting replay, exact long content, concurrent log restoration, lost commit acknowledgements, and definite failure replay after the original CAS base becomes valid again. No paid model calls are needed. - Full local recursive typecheck and build passed during implementation. Server typecheck and the runner TypeScript build passed after the review fixes. The broad local repository test run was stopped after repeated database startup timeouts. Four timing/launch failures in an earlier broad runner run passed focused reruns without changed assertions or timeouts. These are local verification limitations; the complete current-head CI suite is green. An earlier CI workspace job received an infrastructure shutdown signal; its current-head replacement passed. ## Risks - A stopped turn can remain blocked when a dispatched operation has no proven result. The system does not guess its outcome or rerun its effect. - Conflicting results still fail settlement. Existing failed or conflicting journals are not repaired automatically. - Accepted semantic input is limited to 480 KiB of encoded JSON to fit the encrypted transport. Larger input fails before execution. - Instruction filesystem writes and database receipts are not one atomic storage operation. A separately committed attempt and audit record survive rollback. An attempt without a completed success or definite pre-write failure receipt remains blocked as an unknown outcome. It is not replayed or reported as success. - Run-log restoration now reads the durable object before appending when the local log is missing. Failed or incomplete reads reject the append. - No schema migration, dependency change, workflow change, or AgentMail change is included. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test analysis. The exact served model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; the broad local run limitation is recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
af5c2d101c |
fix(paperclip-runner): deliver the shutdown settlement event past the terminal-turn gate (#14668)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The OpenCode driver maps provider events to runtime request events > - A provider turn can end before a pending runtime request receives its answer > - The consumer reads one turn's events and stops at that turn's terminal event > - A settlement event that arrives after that terminal event never reaches the consumer > - This pull request settles the request inside its own turn, before the terminal event > - The benefit is reliable request settlement without weakening late-frame protection ## Linked Issues or Issue Description **What happened?** A pending runtime request stayed open after an OpenCode turn failed through `session.error`. Session shutdown then dropped its settlement event as a late provider frame. **Expected behavior** The driver must deliver `runtime_request.expired` with the original `turnId` and `itemId`, inside the same single read pass the consumer performs on that turn. **Steps to reproduce** 1. Start an OpenCode turn that creates a native runtime request. 2. Leave the request pending and fail the turn through `session.error`. 3. Close the session and inspect the emitted events. **Paperclip version or commit** `22d41c6081f05658e0d7c8485d0f22f35af4a79e` **Deployment mode** Built from source with the OpenCode driver test fixture. ## What Changed - Settle a pending runtime request as soon as its own turn goes terminal, before the terminal turn event. - Add an optional `bypassTerminalTurnGate` parameter to the OpenCode session emitter, and set it on the settlement emit. - Keep a settlement loop in session close as a fallback for a request whose turn never went terminal. - Add a fixture trigger and a regression test that reads one turn in a single pass. - Keep the late provider frame gate unchanged for every provider event path. ## Verification - Run `pnpm vitest run packages/paperclip-runner/src/drivers/opencode/opencode-server-driver.test.ts`. - Run `tsc -p tsconfig.json --noEmit` in `packages/paperclip-runner`. - Run `tsc -p tsconfig.surfaces.json --noEmit` in `packages/paperclip-runner`. - Confirm that the regression test receives `runtime_request.expired` with the original identifiers. - Confirm that the late-frame tests still report dropped provider frames. ## Risks Four call sites now reach the settlement path: session close and the three terminal turn paths (completed, cancelled, and failed). At the three terminal turn paths the turn is still the active turn, so the gate admits the settlement event with or without the parameter. Session close is the only place where the parameter changes the result of the gate, and only for a request whose turn already ended. Every provider event path keeps the existing terminal-turn gate. The known driver test failure is pre-existing and does not touch this change. ## Model Used Claude Sonnet 5, with code execution and test assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the changed tests locally; the known pre-existing failure remains documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation, or no documentation change applies - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I have addressed all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3c561642b4 |
fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask for decisions and optional details through cards in chat. > - A clear approval in a message can leave the matching card pending. > - An unanswered question can also block an unrelated later reply. > - Decisions need a saved source message, while optional questions need to remain answerable in history. > - This pull request records conversational decisions and lets users move on from questions and answer them later. ## Linked Issues or Issue Description **What happened?** Native Claude and Codex could act on approval in chat while the original approval card stayed pending. Pending question forms stayed above the composer, were absent from history, and could suppress later chat replies. A late native question answer could wait for a finished run to reconnect. **Expected behavior** The active agent records a clear approval or refusal against the exact card and user message. Ambiguous replies do not grant consent. Users can send another message without answering a question. The question remains pending in history and can be reopened and answered later. The saved answer reaches the agent. **Steps to reproduce** 1. Ask an agent to propose work with a confirmation card, then approve it in chat. 2. Check that the original card records that approval before work starts. 3. Ask an interactive question, send an unrelated message, and reload. 4. Open the unanswered question from history and submit an answer. Related work: #14408 added completion delivery. #14607 tests completion reporting turns. Neither records conversational answers on approval cards. ## What Changed - Add a confirmation endpoint backed by a user comment, with schema validation, OpenAPI discovery, and native Plan-mode access. Ask mode remains read-only. - Check company, active run, actor, current session, message provenance, revision, and resolver policy. Save the decision and audit in one transaction. Retries do not repeat effects. Emit resolution telemetry after commit. - Give fresh and resumed chat turns the actual pending confirmation identities. Teach agents to save clear conversational decisions before acting and to clarify ambiguity. - Keep unanswered Agent Chat questions as compact history entries. A newer user message closes the old form. Question cards never contribute to composer pending counts or navigation, including after dismissing a fresh form. The history card is the sole reminder; clicking it restores that exact form and draft. - Preserve Agent Chat questions when later messages or questions arrive. Historical ordinary inputs no longer gate later chat replies. Current-run requests, task execution, and governed approvals keep their gates. Remove the special acknowledgement-publication proof helpers that this rule replaces. - Route answers to finished native runs through durable fresh-wake delivery, with existing idempotency and source-question context. Settle late replies against contiguous completed conversation turns and freeze their history replay; failed, unhandled, and newly arriving messages remain actionable. - Add real-component Storybook scenarios, database and UI regressions, and a three-turn native Claude/Codex E2E case. Capture distinct, UI-ready screenshots and report the individual assertions. ## Verification - Focused decision/publication/UI regressions after merging master: 288 passed; subsequent UI draft, failed-send, and conversation checks: 199 passed. - Native question and durable delivery regressions: 106 passed, including all four terminal run states and exactly-once late delivery. Seven targeted regressions fail against the original implementation and pass with the fix. - Latest conversation/decision/native-delivery regressions after the master merge: 121 passed. Covers completed progress, missing or failed intervening turns, new messages during a late reply, stale sessions, and frozen retry/replay boundaries. Four new assertions fail before the ordering fix. - E2E support suite after the master merge: 792 passed. Negative controls reject expired cards, wrong questions/answers, stale or missing replies, unrelated clarification forms, and unexpected tasks. - The embedded-browser walkthrough caught one additional defect: dismissing a fresh question still showed a composer badge. Both Cancel and close-button regressions failed before the fix. The fix at `65f2ade12` passes 170 chat-thread tests and 792 E2E support tests. After merging master, 232 chat-thread/confirmation tests, server/UI typechecks, and token gates pass. The preview and two-provider live E2E pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at that commit. All 55 checks are now successful at `e5512a206` (four conditional checks skipped), including the aggregate verification gate and clean-install canary test. The first attempt was interrupted by simultaneous CI worker shutdowns; one failed-job rerun passed without code changes. - [Published Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on): nine real-component scenarios. Manually exercised move on, reopen, preserve draft, answer later, answer one of multiple questions, and a custom mobile answer in the embedded browser. Retested fresh Cancel and close-button dismissal in the updated build, then reopened and submitted the preserved Green selection and inspected its answered receipt. Static preview has no live model/backend; its callbacks are fixture responses. - [First live campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/) reproduced the late-answer completion-state defect on both providers despite correct saved answers and acknowledgements. It also exposed a valid imperative clarification rejected by the old oracle. Both issues are fixed with regression controls; this failing run is retained as evidence. - [Four-cell qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/) passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous confirmation, each on native Claude and Codex. Inspected saved state, source-message decisions, visible cards, and agent replies. Both late-answer chats settled to waiting; no unrequested tasks were created. [Final branch rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/) passed 2/2 at `142630720`: the same unanswered-question journey after merging master, plus an additional screenshot and browser assertion for the actual late-answer acknowledgement. - [Composer-reminder E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/) passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each, with explicit no-badge assertions before and after reload. Inspected saved pending/answered state, both screenshots with a clear composer, and actual Blue acknowledgements; all five behavioral matchers passed per provider and neither created tasks. Cost coverage is partial; this is bounded workflow qualification. - [Fresh-dismissal E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/) passed 2/2 at `e5512a206`: native Claude and Codex, including fresh Cancel, clear composer, reopen, unrelated message, reload, late Blue answer, and actual agent acknowledgement. All five behavioral matchers pass per provider. Inspected the fresh-dismissal screenshots and saved pending/answered identity; neither created tasks. Cost coverage is partial (4/6 runs). - Prior evidence remains available in [the earlier campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/). Its early loading screenshot and overwritten final capture prompted the UI-ready, distinct screenshot fixes. ## Risks - The model interprets intent. The server verifies permission and provenance; it does not infer consent from text. Ambiguous and unrelated replies are not approvals. - Historical questions can accumulate. They remain visible, pending, and answerable; no automatic answer or expiry is invented. - The change to completion gates is scoped to Agent Chat and ordinary historical inputs. Current-turn and governed approvals retain their existing controls. - Live qualification is limited to the selected stories. Broader native onboarding finalization remains separate work. - No database migration. Telemetry adds no fields or values; the contract and README document the commit boundary. Privacy review was requested on the PR. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser-test orchestration. The exact model ID and context-window size are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dd7fc1f90a |
fix: raise the native journal read limit to 256 MiB (#14711)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native sessions persist control-plane state so they can resume safely. > - The state includes committed provider history needed for recovery. > - The server, runnerd recovery, and durable control plane validate this state before trusting its identity. > - Their differing 64 MiB and 192 MiB limits can reject a valid journal before recovery. > - This pull request aligns all three local state limits at 256 MiB. > - Larger files remain bounded, while recovery can read larger valid histories. ## Linked Issues or Issue Description Refs #13882 Refs #14312 ## What Changed - Raise the server and runnerd recovery limits from 64 MiB to 256 MiB, and align the durable control-plane limit from 192 MiB to 256 MiB. - Add coverage for a valid history above 64 MiB and rejection above 256 MiB. ## Verification - Matching server recovery passed with more than 64 MiB of actual committed event payloads (128 events with 512 KiB deltas). - The real runnerd exact-authority resume regression with the test Codex provider passed with 193 MiB of valid JSON whitespace appended. It crosses the former 192 MiB core limit and confirms the same provider identity. This exercises the runner process and durable control plane with a simulated provider, not a live OpenAI API call. This test used approximately 1.15 GiB peak RSS. - The actual runnerd reader accepted valid 256 MiB JSON and rejected valid 256 MiB + 1 byte. The reader call took 231 ms; the fresh process peaked at 1,244 MiB RSS. - Server tests reject mismatched identity above 64 MiB and files above 256 MiB. - `pnpm -r typecheck`, `pnpm build`, and `git diff --check` passed. - Full local `pnpm test:run`: 13,730 passed, 575 skipped, 7 failed across 6 files. All failures were embedded PostgreSQL startup errors after five attempts. They affected agent hiring, instruction revisions, environment images, reviewed chat bindings, issue monitoring, and legacy continuation authority. The focused journal tests passed; the latest pushed head passed all ordinary CI checks. Superagent is the only blocking check. ## Risks - **Open review concern:** Greptile is 5/5, but Superagent is `ACTION_REQUIRED` with two P2 findings on the server and runnerd readers. Both flag the increased synchronous parsing and memory cost. This PR keeps the requested fixed-limit change small. It does not add a worker parser or a process-wide memory budget. This resource tradeoff needs review before merge. - Large state parsing is synchronous and can consume several times the file size in memory. - Remote checkpoint archive and expanded-size limits remain 64 MiB, so this change alone does not make larger remote checkpoint transfers portable. ## Model Used - OpenAI Codex, GPT-6, with delegated assistance from `gpt-6-luna` at high reasoning effort; tool use and code execution. The GPT-6 context window is not exposed in this task runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d30b03bd8c |
test: add persistent E2E coverage for human blocker decisions (#14707)
## Thinking Path > - Paperclip manages work for AI agents. > - Agents use the coordination skill when work needs human authority or a scope decision. > - PR #14188 replaced automatic manager escalation with direct blocker handling. > - This behavior needs real browser, server, database, and provider tests. > - The test must verify saved human input, task ownership, and resumed work. > - This pull request adds six reusable Product E2E cases and improves the skill examples that they exercise. ## Linked Issues or Issue Description Refs #14188. The merged change needs repeatable behavior coverage. The new suite tests missing administrator access, missing hiring permission, and requester scope questions. Searches found no duplicate blocker-guidance suite. This extends the existing eval system described in ROADMAP.md. ## What Changed - Add the explicit-only `blocker-guidance` Product E2E suite. It has three local scenarios on legacy Codex and legacy Claude. - Use the production UI and public APIs to create work, save a human-only question or confirmation, answer it after reload, and resume the same task. - Check requester identity, ownership history, manager activity, hiring, saved answers, and completion. Keep direct text input as a separate UX result. - Save pending and final screenshots, API checkpoints, skill hashes, provider evidence, and billing data through the existing report pipeline. - Isolate the Claude fixture home. Verify the served skill bytes before dispatch so an old installed skill cannot silently replace the evaluated skill. - Improve the coordination and hiring skill examples. Include the human-only policy, requester address, wake behavior, and handling of authorized scope changes. - Grader v5 requires the exact approved public welcome note as a new worker comment. Browser input checks reject unwritable scope cards before clicking, and confirmation direction must be saved in the resolution before the worker wakes. - Add grader calibration and browser-input tests. Update the fixture guide and generated capability inventories. ## Verification - `pnpm build`: passed after rebasing onto current master. - `pnpm -r typecheck`: passed. - `pnpm test:e2e:runner:typecheck`: passed. - `pnpm test:e2e:runner:unit`: 742 tests passed. - `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests passed. - `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells found. - Capability inventory and generated-contract checks: passed. - `pnpm exec vitest run server/src/__tests__/hiring-operational-examples.test.ts`: four tests passed after synchronizing the generated API reference and section anchor. - Full general and serialized test suites: passed in CI on `6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm test:run` was interrupted after complete CI coverage passed; it is not claimed as a completed local full-suite run. - Final GitHub checks: 54 passed, two optional Storybook checks skipped. The runtime-exposure startup test hit a 10-second readiness timeout once, passed a targeted local reproduction, and its CI shard passed the single retry without code changes. - Current-head Greptile: 5/5, clean check, zero unresolved threads. - Historical live measurement on September 29 at `4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex `gpt-5.6-sol` passed 8/9. These runs predate this rebase. - Version 5 changes the scope answer to an exact approved publication draft. The historical runs do not qualify that new requirement; the two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec` passed Codex and failed Claude. Claude posted the correct salary-free sentence but omitted its required reference line from that comment, placing the reference in a separate completion message. The `public-welcome-note` check correctly failed. An earlier Claude database-startup failure was retained separately; its fresh-instance retry reached the model. This pilot is not a six-cell qualification. - The failed Codex scope case omitted `addresseeUserId`. The strict routing check remains. All 18 attempts had clean evidence manifests and passed cleanup. - To repeat with provider credentials: `pnpm test:e2e:runner -- --suite blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is excluded from `--all`. ## Risks - The live suite measures variable model behavior. The retained 17/18 historical result and the current 1/2 scope pilot are not all-pass qualifications. These paid cases are opt-in; their observed model failures remain visible independently of deterministic CI checks. - A separate generic task-replacement diagnostic still exposed a Claude refusal. The ordinary cases use specific business decisions. The diagnostic is not a standalone catalog case in this change. - Earlier measurements included an old installed Claude skill and test defects. Their grades remain retained and are not combined with the three final repetitions. - Skill examples can affect when agents ask for human input. Downstream permission checks still apply. - Native runners, Daytona, agent-requester routing, and real external connection authorization are outside this suite. - Raw provider traces and credentials remain private. No screenshots, raw reports, secrets, workflow changes, or lockfile changes are committed. ## Model Used OpenAI GPT-6 through Codex assisted with this change. The exact deployed variant and context window size are not exposed in this session. The assistant used reasoning, repository edits, tool use, and shell execution. The evaluated models were `gpt-5.6-sol` and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4f9a7c613 |
test(runner): guard continuation after journals exceed 2 MiB (#14312)
Add an actual runner resume regression above the former 2 MiB journal boundary and an explicit-only three-turn Daytona workflow that grows real execution history. Verify journal size and distinct completed tool calls without exporting private payloads. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
eb31b926a1 |
fix(runner): keep the OpenCode session event stream open across turns (#14582)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip Runner keeps provider sessions and their event streams for agent runs > - The OpenCode driver closed its event queue after each terminal turn > - A second turn on the same session then lost its response and completion events > - This pull request keeps the queue open between turns and rejects late events for sealed turns > - The benefit is reliable multi-turn OpenCode sessions with visible diagnostics for late provider events ## Linked Issues or Issue Description **What happened?** The OpenCode driver closed its event queue when a turn completed, was cancelled, or failed. A second turn on the same session then lost its response and completion events. **Expected behavior** The session must keep its event stream open between turns. Each turn must deliver its response and one terminal event. The session must close the stream only during session shutdown or an unrecoverable pump error. **Steps to reproduce** 1. Start one OpenCode session. 2. Run one turn and wait for its terminal event. 3. Run a second turn on the same session. 4. Confirm that the second turn delivers its response and terminal event. **Paperclip version or commit** `c0e1d87ddc181329471fa80a2061b2c538bb6618` **Deployment mode** Built from source with the Paperclip Runner package test suite. ## What Changed - Keep the OpenCode event queue open across completed, cancelled, and failed turns. - Track sealed turn ids and reject later events for those turns with a diagnostic event. - Preserve queue shutdown on session close and unrecoverable pump errors. - Give each simulated fixture turn unique provider event ids. - Add regression tests for completed, cancelled, failed, and closed-session paths. ## Verification - Type check: `cd packages/paperclip-runner && node ./node_modules/typescript/bin/tsc -p tsconfig.json --noEmit` passed. - Driver tests: `cd packages/paperclip-runner && npx vitest run src/drivers/opencode/opencode-server-driver.test.ts` passed except for the known pre-existing flaky test described below. - Consumer tests: `cd packages/paperclip-runner && npx vitest run src/native-session-runtime.test.ts src/backends/harness-driver-backend.test.ts src/cli/opencode-app-server-proxy.test.ts src/conformance/harness-driver.test.ts` passed. - The known flaky test reproduced on unmodified `master` because fixture event order depends on a local MCP HTTP round-trip. - CI must run the full pull request suite. ## Risks - The queue now retains sealed turn ids for the session lifetime. OpenCode does not reuse turn ids, so this set grows with the session. - A late provider event cannot reach a later turn. The driver emits a diagnostic event so the rejection remains visible. - The change does not alter session shutdown or unrecoverable pump error handling. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This pull request fixes an OpenCode Runner bug and does not add a roadmap feature. ## Model Used OpenAI Codex, GPT-5, tool use and code execution; Anthropic Claude Sonnet 5 also assisted with the implementation. The runtime did not provide a context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2de43fc909 |
fix(issues): keep agent mentions as context and defer personal app authorization (#14577)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Each task has one assignee. Explicit assignment and review requests select who should act. > - An agent mention started another agent on a task it did not own. Native attachment staging then rejected that run. > - Allowing that run through startup could also let two agents work on the same task. > - Mentions should identify relevant context. They should not start work or forward comments to other tasks. > - A personal app installed on a shared agent must also wait until tool use to resolve the current user's grant. > - This pull request removes mention dispatch and keeps missing personal app credentials from blocking startup. ## Linked Issues or Issue Description **What happened?** A native agent mentioned on another agent's task failed with `paperclip_runner_attachment_staging_not_authorized`. The source task could already be complete. A nearby optional-app warning was a separate problem: personal app tools were excluded when their shared health state required attention. **Expected behavior** An agent mention is context only. It does not wake the agent, take ownership, or copy a comment onto another task. Normal feedback still reaches the assignee. Assignment and explicit review requests still dispatch work. An unavailable personal app does not block startup or produce a startup warning. Tool use requests the current user's authorization and never uses another user's grant. **Steps to reproduce** 1. Assign a task to agent A. Post a comment that mentions agent B, including a comment that closes A's task or references B's child task. 2. Confirm the comment retains its agent link and B receives no run or deferred wake. A can still receive normal feedback. 3. Install an active personal MCP connection on B. Give only Alice a grant and leave shared health at `error`. 4. Explicitly assign work to B for another user. Confirm it can finish without using the app. 5. Ask B to use the app. Confirm its tool call shows an inline connection request for the current user. Related work: Refs #11144. This change uses the existing execution-time personal grant resolution. ## What Changed - Remove mention dispatch from standalone comments and issue updates. Remove implicit forwarding of parent comments to a mentioned worker's child task. - Ignore new requests with the legacy mention wake reason before creating a run or deferred request. Preserve already accepted queue entries, which can combine assignments and feedback with a later mention. - Remove the native mention admission, staging, and finalization exceptions from this PR. Native task ownership checks remain intact. - Keep active, installed personal app tools available despite shared health errors. Remove optional-app startup warnings. Tool execution retains the current user's grant and policy checks. - Update agent instructions and product/API docs. Refresh generated capability source anchors. ## Verification - Red: comment-route regressions reproduced extra agent wakes and child comment forwarding. A separate regression proved that cancelling by the last coalesced reason could drop an accepted assignment. - Green: the targeted route, wake queue, heartbeat, workspace, responsible-user, MCP discovery, and HTTP gateway suites passed. The final queue and heartbeat rerun passed 104 tests, the restored queue adapter passed 56, and both comment-route suites passed 135. These include accepted assignment preservation, rejection of new mention requests, and normal assignee feedback. - `pnpm -r typecheck` and `pnpm build` passed locally. The full local `pnpm test:run` attempt was interrupted for review/CI fixes, so it is not claimed as a completed local pass. It exposed a cleanup timing race in the concurrent-mention assertion, now fixed and verified across 10 repetitions. CI also exposed an obsolete test waiting for the removed mention lookup; it was reproduced and fixed, then both comment suites passed. Final full-suite verification is through CI. - Final head `bd9ea4cb05a8f081c54e017760a8999f9ea6ef44`: 54 checks passed, 2 Storybook checks intentionally skipped; no pending or failing checks. Full CI includes general and serialized suites, all 8 browser shards, runner verification, typecheck, build, and canary dry run. Greptile is 5/5 on this exact commit, with no unresolved findings. - One unchanged Cursor adapter test hit its 10-second CI timeout. All 5 tests in that file passed locally; one retry of its CI shard passed all 674 tests (3 skipped). The aggregate verification gate then passed. No code or timeout was changed for that retry. - Live browser check: inserted a structured mention with the picker on a human-owned task. The saved link remained visible. Database checks found zero new runs and zero wake requests. - Live Codex runner check: explicitly assigned that task with the unavailable personal app attached. The run succeeded and committed completion without using the app or creating a connection card. - Live browser follow-up: asked the assignee to call PostHog and mentioned another enabled agent as context. Only the assignee ran. It succeeded and displayed the existing inline connection card. Only Alice's grant existed; the run belonged to a different user. - The HTTP regression covers tool discovery with no provider calls or connection cards, first use returning the current user's authorization request, and successful retry after that user's grant exists. - App checks use an isolated local fixture and a fake MCP provider. They do not use production app credentials. ## Risks - Intentional behavior change: workflows that used mentions to wake agents must use assignment, a bounded child task, or an explicit review request. - Already accepted queue entries retain their prior rules. An old entry can combine assignment or feedback with a later mention; its last reason cannot safely identify mention-only work. New mention requests create no run or deferred wake. - Personal apps with a shared health error remain discoverable. Actual tool use still requires the responsible user's grant and existing policy gates. - No database migration or public API schema change. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact serving model ID and context-window size are not exposed in this session. - Live native-run verification used `gpt-6-astra` through the Codex provider. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
da887ea3e9 |
fix(runner): honor Codex effort selected in composer (#14568)
## Thinking Path > - Paperclip manages AI agents that work on assigned tasks. > - The task composer lets a person choose an assignee, model, and effort for the next run. > - A Paperclip Runner agent can use Codex as its provider. > - The composer hid Codex effort for that agent because it checked only the older Codex adapter. > - The native Runner input also did not carry an effort choice to Codex. > - This pull request carries the chosen effort from the composer to each Codex turn. > - People can now select a supported effort and get the effort they selected. ## Linked Issues or Issue Description Refs #14322 **What happened?** The composer showed a model but no effort slider when the assignee used Paperclip Runner with the Codex provider. A task-level model override also did not reach the native Runner input. **Expected behavior** The composer shows effort choices for a known Codex model. The next native Codex turn uses the selected model and effort. **Steps to reproduce** 1. Open a task composer. 2. Select an agent that uses Paperclip Runner with the Codex provider. 3. Select a known Codex model such as `gpt-6-astra`. 4. Open the assignee and model picker. The effort slider is missing before this change. ## What Changed - Show known Codex effort levels for Paperclip Runner Codex assignees. - Save the task effort override in the native run input and send it to Codex on each turn. - Apply the task's merged model and effort overrides when the native run starts. - Apply a task model override for OpenCode Runner without changing the agent's provider. - Add Runner effort tests and desktop and mobile Storybook cases. ## Verification - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm build-storybook` passed. - `pnpm check:token-gates` passed. - Focused UI, server, Runner contract, and Codex driver tests passed. - The full CI test matrix, build, typecheck, and canary dry run passed on the latest head. ## Risks - Native Runner inputs add an optional Codex effort field to the current v5 input. Older inputs keep their previous behavior. - A known model rejects an effort that its catalog does not support. Unknown models do not show a slider. > This fixes an existing composer bug. I checked `ROADMAP.md`; it does not describe this bug as planned work. ## Model Used OpenAI Codex, GPT-6. The exact deployment ID and context window are not exposed in this session. The model used reasoning, code execution, and repository tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |