mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 00:54:38 +02:00
5ee9e751fbb38787c055ee793c3412ecfc89da2e
25
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
db8f8fe5b7 |
fix(evals): select Grok subscription protocol credentials explicitly (#13901)
## Thinking Path > - Paperclip manages agent work through shared runner contracts. > - Direct protocol evals qualify provider behavior against a mock control plane. > - Grok supports API keys and company subscription credentials. > - The hosted protocol workflow selected an API key for every Grok cell. > - Product subscription support did not enable subscription protocol runs. > - This change adds explicit subscription selection and checks the recorded authentication mode. ## Linked Issues or Issue Description Refs #13878, #13882, #12618. The direct Grok protocol roster cannot run with subscription authentication through the trusted default-branch workflow. Add an explicit selector while keeping API-key dispatches compatible. Keep the actor allowlist, protected environment, immutable source revisions, and publication gates. ## What Changed - Add `grok_authentication` with `api_key` and `subscription` choices. Keep `api_key` as the compatibility default. - Deliver the protected `GROK_AUTH_JSON` secret only to a subscription-selected Grok cell. Do not provide an API key to that cell. - Read authentication mode from the pinned eval program's actual roster summary. Retain it in the cell, catalog, campaign roster, and result. - Reject missing or mismatched authentication evidence during aggregation. Preserve cell metadata and an allowlisted failure reason before failing a cell, so malformed evidence cannot hide the retained attempt. - Document credential setup, source separation, and temporary-secret cleanup. ## Verification - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-campaign.test.mjs packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`: 34 tests passed. - Validated all 39 Grok cells at eval revision `3213dbec7e8ca1865ea95e6db7e7d34b095eb47a`; every selected cell requests only the subscription credential. Validation made zero provider calls. - Negative coverage rejects invalid selectors and missing or API authentication evidence in an otherwise passing subscription attempt. - `git diff --check` passed. All 53 current-head checks passed; the unchanged callback-drain timing test passed its bounded rerun, and the failed attempt is retained. Greptile reviewed `ecd3dcc0e998f07cf56fcb1f087946f50f388bec` at 5/5 with no remaining findings. - No Docker or broad builds ran on the developer machine. CI performs repository checks on the configured fleet. ## Risks Grok runs require an eval revision that records `authenticationMode` in the roster summary. Missing evidence fails closed. The credential contains account access and refresh tokens; an owner must approve its delivery to the protected environment before a live run. The change adds no PR trigger or authorization bypass. Live subscription protocol qualification remains pending this workflow reaching master. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b648d8cdda |
fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path > - Paperclip manages AI agents and their provider connections. > - Product E2E checks real tasks through the browser, server, and runner. > - Grok qualification needs separate API-key and subscription evidence. > - Product subscription tests and direct Grok protocol evals need explicit credential delivery. > - This change supplies each credential only to its selected profile and prepares the pinned binary. > - Maintainer authorization and protected-environment gates remain required. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13882. The Grok feature branch has a manual subscription qualification profile. The trusted master workflow must admit its selected credential and prepare the same verified binary and artifact verifier as the API profile. Direct protocol evals also need the selected xAI key and pinned Grok binary. These prerequisites do not register or schedule the new profiles on master. ## What Changed - Deliver `GROK_AUTH_JSON` from the protected paid environment only when the selected profile requests that credential. - Install the checksum-verified Grok binary for the local subscription profile. - Prepare the pinned artifact verifier for the manual subscription suite. - Extend workflow security assertions to cover the new credential and profile. - Add the ACPX Grok credential mapping to the trusted-master catalog, then deliver only the selected `XAI_API_KEY` to direct protocol cells and install the target’s checksum-verified Grok binary before packaging. - Allow a direct-protocol concurrency override from two cases up to the existing configured ceiling; it can only lower concurrency. - Document the Grok protocol workflow and its API-only credential boundary. - Render missing LLM usage and cost as Unavailable, and label partial observations with coverage. Preserve raw records, grades, and actual zero costs. - Preserve measured campaign source metadata during report regeneration instead of inheriting the renderer checkout or CI event; skip empty legacy source records when recovering older provenance. ## Verification - Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54 reported checks successful, two intentional skips, Greptile 5/5, and zero unresolved review threads. [CI run](https://github.com/paperclipai/paperclip/actions/runs/35890978288). - After merging current master, all 17 workflow security/image tests and 23 catalog/workflow policy tests passed. The trusted catalog also generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per shard. The new policy tests execute the concurrency guard against valid, out-of-range, and malformed values. - The Grok branch separately passed 450 Product harness unit tests, including private company credential staging, cleanup, and token-fragment redaction. - The fresh-login native subscription smoke passed three repetitions of tool execution, session resume, restrictive permissions, and cleanup. These are setup evidence; full subscription Product qualification remains pending. - All 72 focused report/billing/history/catalog tests and the Product harness typecheck passed for the report-display change. The initial sandbox run could not open the tsx IPC socket; the permitted rerun passed. A zero-provider-call replay of the actual 16-cell Grok campaign preserved all result records, grades, timing, and source provenance while correcting missing usage labels. - Review the thirteen-file diff. Provider credentials still enter only the selected paid-test step; default-branch, numeric-actor, and environment restrictions are unchanged. ## Risks This admits a refreshable subscription credential to explicitly selected trusted tests. Store it only in `runner-e2e-paid`, use a test login, and remove it after qualification. Unselected profiles receive an empty value. Pull requests cannot trigger the paid workflow. This PR changes no fleet admission or actor allowlist. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b8ec588f3 |
Stop duplicating wake context in adapter environments (#13891)
## Thinking Path
> - Paperclip manages agent work and preserves task context.
> - Built-in adapters already include wake context in the agent prompt.
> - They also copy the full wake JSON into a process environment
variable.
> - A large environment entry can prevent the agent from starting with
`spawn E2BIG`.
> - This change removes the duplicate environment entry and uses the
existing prompt delivery.
> - The agent keeps its context without extra file transport or new
history limits.
## Linked Issues or Issue Description
Refs #13144, #13860, #13872, #13793.
Large wake payloads can exceed the operating system limit for one
environment entry. The launch-envelope fix in #13793 handles the outer
transport but leaves that child environment entry intact.
Credit to @nickyleach for the prompt-only approach in #13144. This PR
applies that part on current master. It does not include that PR's
30-item history limits or recovery-history endpoint. Those behavior
changes can be reviewed separately from the process launch fix.
## What Changed
- Stop exporting `PAPERCLIP_WAKE_PAYLOAD_JSON` in the shared ACP engine
and all ten built-in adapter writers.
- Ignore configured values of the retired variable so saved adapter
settings cannot restore the oversized entry. Also drop inherited copies
in Hermes, which builds its environment directly.
- Keep scalar runtime variables, existing prompt rendering, continuation
history, resume deltas, gateway bodies, and Hermes JSON template
variables.
- Document the prompt delivery contract and the migration for custom
instructions that read the retired variable.
- Test large local and sandbox child-process launches, fresh and resumed
ACP turns, SDK delivery, and configured-variable filtering.
## Verification
- `pnpm -r typecheck` passed.
- Focused adapter utility, ACP, Codex child-process, and Cursor Cloud
suites: 340 tests passed.
- Hermes execution and prompt tests: 18 tests passed using its package
Vitest configuration.
- The child-process tests deliver over 128 KB of context through stdin
and check the complete text. The ACP test retains 50 complete messages
and 50 completed actions, then checks the resumed delta.
- `pnpm build` passed.
- `pnpm test:run` was attempted, then stopped after it reproduced ten
macOS runtime-skill-cache permission failures (also reproduced on
unchanged master) and one HTTPS backfill test failure. The HTTPS test
passed when rerun unchanged on this branch and master. The complete
local suite was not completed; Linux CI provides the full-suite gate.
- CI is green on
|
||
|
|
7944ed3d97 |
fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Native runner agents need governed tools, durable runtime state, and useful execution evidence. > - A first activity trace waited 53.467 seconds even though tool activity took 6.274 seconds; provider input arrived before the server API call executed. > - Native agents also need a safe way to hire teammates without asking the model to rebuild runtime configuration. > - This pull request separates the observed ACP input-stream window from the actual server `tool.execute` span and adds a server-owned native hire contract. > - The benefit is clearer latency evidence and safer native teammates with existing approval, auth, and company boundaries preserved. ## Linked Issues or Issue Description Related Daytona provenance work is in [#13814](https://github.com/paperclipai/paperclip/pull/13814). No duplicate public PR was found for this combined timing and native-hire change. **What existing behavior does this improve?** Native runner agents can use governed tools and request hires. The server did not expose a safe native hire operation that reused the caller's validated runtime settings. First-activity traces also mixed provider input timing with server tool execution timing. **Current behavior** A native hire must construct a separate runner configuration. Full configuration copying could expose paths, instructions, secrets, or sessions. Timing evidence could make a provider or MCP identity join appear proven when the trace did not contain that join. **Proposed behavior** The native `hire_agent` operation accepts identity and persona inputs. The server sends `adapterType: "paperclip_runner"` with `inheritRuntimeFrom: "caller"`, then copies only validated provider, model, permission, lifecycle, and bounded execution settings. It inherits and validates the default environment, derives the managed AI binding through existing normalization, preserves approval and permissions, and creates fresh child instructions. Caller secrets, paths, prompts, and sessions are excluded. Provider events now include the optional boolean `inputUpdated`, with Rust forwarding support. Timing evidence separately records the ACP input-stream window and the actual server `tool.execute` activity. It does not claim a provider or MCP join without matching evidence. **Reason and benefit** Native agents can hire teammates that start with the caller's approved execution policy. Operators retain company boundaries, auth rules, approval gates, and requalification. Reviewers can distinguish provider streaming time from server API execution time when diagnosing first-activity delays. **Breaking changes** None for existing hires or tool calls. `inheritRuntimeFrom` is optional and only applies to same-company native agent callers. Conflicting explicit runtime settings are rejected. The provider event field is optional for existing producers. ## What Changed - Added the native `hire_agent` protocol action, catalog entry, API contract, and runner authority checks. - Added `inheritRuntimeFrom: "caller"` validation and a closed native runtime inheritance allowlist. - Preserved managed AI binding normalization, default-environment validation, approval snapshots, permissions, requalification, and fresh child instructions. - Added provider `inputUpdated` schema support and Rust forwarding. - Added first-activity and server tool timing evidence with conservative identity-join handling. - Added route, authority, provider-event, sidecar, API, catalog, and Rust-focused tests. - Kept private Honeycomb links, raw traces, and local result paths out of this description. ## Verification Focused checks passed: - 458 timing/session checks. - 61 native hire inheritance checks. - 20 hire authority checks. - 1,741 API checks. - 106 catalog checks. - 54 provider sidecar checks. - 12 Rust provider checks. Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2). R1 stopped at missing Docker image setup. The final trace is available at https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces. Latest-head CI passed all required build, typecheck, Rust, static, Vitest, serialized-server, workspace, chat, and E2E jobs. The focused local checks listed above passed; the broad local suite was not run before the live evaluation, while CI provides the full repository verification. ## Risks - Timing fields describe separate observed windows. They do not prove a provider or MCP owner without a valid trace join. - The inheritance allowlist must stay synchronized with native runner configuration fields. - Approval snapshots include resolved safe inherited settings and should be reviewed when native configuration fields change. - The focused local suite is narrower than the full repository suite; latest-head CI covers the broader repository checks. > Roadmap review: `ROADMAP.md` places this work within Paperclip's bring-your-own-agent direction. It extends existing native runner hiring and observability behavior. ## Model Used OpenAI GPT-6 (exact serving model ID is not exposed), with extended reasoning and repository tool use; GPT-5.6 Luna assisted with focused implementation and verification work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
43acbcc398 |
fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task state to provider sessions. > - Follow-up turns must retain provider memory and carry new user direction. > - Lost session IDs caused repeated context and extra input tokens. > - Native question answers and approval races could leave valid work blocked. > - This pull request repairs those paths and adds regression coverage. > - Agents can continue accepted work without repeating the conversation or losing the user's answer. ## Linked Issues or Issue Description Refs #13574. That merged PR shortened continuation prompts and moved question instructions into tool documentation. This change preserves sessions and fixes failures exposed by broader testing. Related runtime work: #13408 and #13410. **What happened?** Native follow-up turns could lose the provider session ID. Completion guidance could replace the original task with its latest comment. Claude native questions could remain pending after the user answered. Approval during a running tool call could suspend the run before the tool response arrived. Onboarding and chat handoff instructions also caused repeated planning or missing plan documents. **Expected behavior** Reuse a valid provider session. Send only new events when that session already has the history. Preserve the task requirements and apply later user direction. Store the question answer and deliver it to the waiting run. Finish governed tool responses before suspending. Execute the accepted plan without asking for the same approval again. **Steps to reproduce** Run the continuation, local-session-integrity, first-task, and agent-chat suites with native Codex and Claude. Include provider-question-bridge, accept-while-running, and plan-handoff. **Paperclip version or commit** This branch is based on master |
||
|
|
11921075a4 |
Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first task helps a new user define and approve useful work. > - That workflow needs reusable instructions and tests against the production experience. > - Native Codex and Claude must load the assigned skill, including after resume. > - Maintainers need recorded conversations and precise failed checks to judge regressions. > - This pull request adds the first-task skill and a suite in the shared Runner E2E harness. > - It keeps behavior results separate from informational quality scores and incomplete recordings. ## Linked Issues or Issue Description **What existing behavior does this improve?** The first onboarding task and the Runner E2E report used to review it. **Current behavior** Onboarding embeds its policy in a hidden brief. Native Codex drops the skill-instructions setting at the Rust boundary. The shared E2E harness has no onboarding suite or full conversation view. **Proposed behavior** Assign and invoke `/first-task` for the onboarding task. Send selected Codex skills as structured protocol inputs. Run twelve scenarios across legacy Codex, legacy Claude, native Codex, and native ACPX Claude. Include all 48 cells in full campaigns. Show recorded chat, question and approval cards, exact checks, instructions, and billing in the shared dashboard. **Reason and benefit** Measure the real onboarding experience before changing prompts. Distinguish infrastructure failures, behavior failures, and unexercised journey steps. **Breaking changes** No database migration or production API change. First-task instructions now live in an assigned skill. The user-edited persona is preserved; the skill includes the maintainer-approved proposal-mode mapping and saved-plan requirement. Related: #11043 is earlier onboarding work. #13422 already fixes native Claude model pinning, context delivery, and read permissions on master; this branch includes those fixes through its base. The new Claude recovery test supplements them. ## What Changed - Extract and assign the first-task skill while retaining the production greeting and opening question. - Carry the Codex skill-instructions flag through thread start and resume. Resolve explicit task skill references only against assigned skills and send native skill inputs. - Invoke an unambiguously selected assigned skill through Claude ACPX’s native slash-command parser on initial and resumed turns, retaining the entire task/wake envelope as its argument. Do not carry that invocation into ordinary tasks. - Restore the saved single-task proposal modes: confirmation card, or saved plan with revision-targeted checkbox approval. Explicit plan requests also require a saved plan. - Add first-response and complete-journey cases with fixed user facts, acceptance checkpoints, durable outcome checks, and accounting for child runs. - Fail the eval when choice questions have fewer than two real options. Recognize planning documents without treating them as completed work. - Add optional, bounded quality judging as explicit post-processing. - Render full conversations and static interaction cards in the shared report. Conversations start folded. Show original and regraded results and incomplete journeys distinctly. - Keep credential-persistence scanning outside the first-task behavioral suite; retain public evidence redaction. - Refresh generated capability references after the API-reference edits. - Correct shared native question guidance and tool schemas: choices need at least two meaningful options; open-ended questions use canonical text fields with the required compatibility payload. Verify both formats through real tool-authority persistence. - Disable announcements automatically for every isolated Runner E2E process and label the gallery environment/provider/target explicitly. - Remove CI races in the GitHub connection browser test and native session recovery test by waiting for the actual async work before asserting its results. ## Verification - `pnpm exec vitest run server/src/services/onboarding-first-task-assets.test.ts server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19 passed. - `pnpm --dir packages/paperclip-runner exec vitest run src/drivers/acpx/runtime-host.test.ts src/drivers/acpx/native-skill-prompt.test.ts src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command forwarding and the 1 MiB input boundary both failed before their fixes and passed afterward. Coverage includes changed skills on reopen, approval context, and an ordinary subsequent task. - Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64 first-task fixture and grader tests also pass. - Full repository typecheck and build passed locally. Server typecheck and Runner build passed again after the native-command change. - Full GitHub Actions CI passed on `23e56447b`: all server/workspace/browser shards, Runner verification, typecheck/release registry, build, canary, policy, and Docker checks. Greptile reviewed this exact head at 5/5 with no unresolved threads. The earlier broad local run had database startup/timing failures that passed isolated retries; the complete remote suite is green. - Merge verification against current master: 312 harness tests and 13 native recovery tests passed. Regenerated semantic contracts and fixture hashes pass their consistency check. Full local typecheck and build also passed on the stacked queue branch. After merging the latest master and preserving the GitHub setup timing regression in the split browser suite, both focused GitHub browser tests passed. Three CI timing/startup flakes passed local verification and one remote retry; all latest-head checks are green. - Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local mock API confirmed that `/skill-name` expands the assigned skill body before the model request and retains the task arguments. A prose mention does not. The probes made no paid model calls. The ACP probe used the current first-task skill body and retained the wake arguments. - [Full 48-case campaign and report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/): 44 passed after three interrupted Codex cases completed in targeted reruns. Original results, regrades, and all 51 executions remain in the report provenance. - [Claude campaign after the shared-question fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/): 10/12 passed with zero single-option failures. All 12 recorded the current assigned skill and corrected guidance. The failures exposed skipped skill invocation and a missing saved plan. This PR adds native command invocation and explicit saved-plan instructions; the subsequent report below still shows behavior failures. - [Fresh 12-case Claude report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/) at `78452129e`: 10/12 pass after correcting two false proposal-matcher failures. The recordings said “Here is the task I will create and run/complete” in approval cards; the old matcher missed that word order. Regression tests failed before the fix and pass after it. Original results and offline regrade provenance remain linked. No agent rerun was needed. Zero single-option-question failures; two behavior failures remain: direct work before acceptance on a plain first message, and an explicit plan request without a saved plan. Neither check was relaxed. The follow-up `82087ac7e` fixes command-prefix size accounting; `94aefb1f3` fixes only that proposal matcher. - Report browser checks confirm folded conversations, rendered cards, explicit Local/Daytona labels, and no page errors. The published-object audit scanned 1,306 text files across 2,154 objects with no credential-format findings or prohibited files. Image pixels and unknown token formats are outside that scan. ## Risks - Model behavior is nondeterministic. One campaign is evidence, not a guarantee. The two remaining Claude behavior failures are visible in the report and require further product work; this PR does not claim all onboarding scenarios pass. - The suite checks persisted Paperclip effects. It cannot prove the absence of arbitrary external effects. - Historical recordings can miss later journey steps. These remain incomplete, never passes. - Native profiles switch runtime after the production onboarding wizard because it does not yet expose a native option. - Quality scores are informational and cannot override behavioral failures. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, and code execution. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9fd2e50310 |
feat: create company skills from runner tasks (#13538)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner gives agents tools to change company resources. > - Users need agents to save reusable skills during a task. > - A saved skill needs a visible result that users can inspect and edit. > - This pull request adds `create_skill` and a task feed card linked to Skill Studio. > - Users can open the saved skill from the task and edit the same resource. ## Linked Issues or Issue Description **Subsystem affected** Runner tools, company skill storage, task feed, and Skill Studio. **Problem or motivation** The Runner has no dedicated tool to create a company skill. A user cannot follow a creation result from the task feed to the saved skill. **Proposed solution** Add a company-scoped `create_skill` tool. Save the skill with the existing company policy. Add one creation card to the task. Open a named sidebar tab from that card. Let the user open the same skill in Skill Studio. **Alternatives considered** An agent can write a local file, but that file is not a company skill. A second document copy in the task would become stale after a Studio edit. The sidebar therefore reads the saved skill directly. **Roadmap alignment** This extends the shipped Skills Manager, Skill Studio, and Skills Store milestone. The maintainer requested and approved this scope. Search found no duplicate `create_skill` PR or issue. Related UI validation work: #8715. This PR does not change that validation display. ## What Changed - Add the real Runner tool, its contract, and its mock implementation. - Validate the complete SKILL.md and derive company, task, agent, and run identity from authentication. - Apply the existing company skill policy. Do not assign the skill to an agent. - Make keyed retries return one skill and one creation event. Reject conflicting retries. - Make concurrent file creation safe. Never replace an existing published skill during creation. - Add a creation card, a named sidebar tab, and an Open in Skill Studio action. - Show saved Studio edits when the user returns to the task. - Add storage, policy, mode, retry, UI, and Product E2E tests. Document the tool. - Fix deleted-name reuse, onboarding panel persistence, immediate feed refresh, and mock validation parity from review. - Serialize Studio file edits and renames with skill deletion and recreation. Reject stale editor requests before they can change a replacement skill. - Generate the standalone mock parser and validator from the production contract. Use portable UUIDs so the browser scenario bundle builds. ## Verification - All latest-head PR checks pass on `145dd76a5`, including all server shards, browser E2E, Runner verification, build, typecheck, and release dry run. Greptile: 5/5 with no open findings. An interrupted CI runner was retried successfully. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Review regressions: 73 storage tests, 6 real API tests, 63 UI tests, and 61 semantic runtime tests passed. Parser synchronization passed. - CI exposed existing fire-and-forget Sentry test races. Reproduced the resumption race locally, then synchronized the related sweep and finalizer assertions on the actual report; all 27 tests across the three affected files pass. - Runner scenario browser build and strict content-security-policy check: passed. - Runner suite: 2,012 tests passed; 10 skipped. - `pnpm test:run`: the general-server batch had 12,416 passes and two failures. The old tool-count assertion was fixed; all 16 authority tests then passed. The chat webhook test had a socket error; it passed four isolated reruns. - Both workspace test groups passed. The isolated route suites completed. Two socket failures in the initial route batches passed on individual reruns; all remaining 61 files passed. - Product E2E `create-skill-studio`: passed with local Codex and local ACPX Claude. - Manual browser test: submit a task, observe the real tool call and creation card, open the sidebar, edit in Studio, save, and return. The task reached Done. The saved second revision and sidebar tab survived a server restart. - The new companion headless Runner Eval passed. Companion coverage PR: https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not run because no immutable runner image was configured. ## Risks - Database writes and local file writes cannot share one transaction. Recovery accepts only an exact file-for-file retry after a database rollback. Conflicting files remain untouched. - The sidebar displays the current skill. The feed card remains the historical creation receipt. - No database migration, dependency, or workflow change is included. - Remote Daytona behavior still needs a run with a configured immutable image. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and browser verification. OpenAI `gpt-5.6-luna` assisted with bounded implementation and eval work. Both used code execution and tool access. The host did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
18989a9e73 |
docs: add eval guide, authoring skills, and public history hub (#13535)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its evaluations test both the Runner and complete product workflows. > - The guides and run histories are in separate places. > - The shared Evalbook viewer can make the test boundary unclear. > - This pull request names the two families and adds a guide, authoring skills, and a public hub. > - Contributors can choose the correct test and inspect its history. ## Linked Issues or Issue Description **Issue type** Missing documentation. **Where is the issue?** Runner and Product E2E evaluation guides, case-authoring procedures, and public result navigation. **What's wrong?** There is no single entry point. A report format can be mistaken for an execution boundary. There are no dedicated case-authoring skills for these two families. **Suggested fix** Add a guide and three skills. Link both existing histories from a public hub. Keep existing campaign URLs and grading unchanged. ## What Changed - Add `doc/evals.md` and links from existing guides. - Add the `paperclip-evals`, `add-runner-eval`, and `add-product-e2e-eval` skills. Install copies in `~/paperclipai/.agents/skills`. - Add a static hub builder that reads the existing public history feeds. - Show a dated snapshot for each family. Label partial campaigns and preserve measurement dates across report refreshes. - Document publication and refresh commands for https://pages.paperclip.ing/evals/. ## Verification - Seven Python summary tests pass: `python3 -m unittest discover -s scripts/evals-hub -p 'test_*.py'`. Run these checks directly; this PR does not modify package scripts. - All three skills pass the skill-creator `quick_validate.py` check with `/usr/bin/python3`. - Build tested with saved history fixtures and the live public feeds. - Desktop and mobile browser checks pass. The mobile page has no horizontal overflow. - Published https://pages.paperclip.ing/evals/. Browser check: HTTP 200, no page errors, all eight links return HTTP 200, no mobile overflow. - Independent skill exercises found the existing Notion-decline case and a direct Runner permission-denial case. Roster validation with an explicit run ID passes. - Missing refresh measurement date: regression fails before the fix and passes after it. - `git diff --check` passes. - No paid evals were run for this documentation and reporting change. The preceding head passed typecheck, build, server/workspace tests, runner verification, browser E2E, and the canary dry run. Checks for the latest commit are pending. Local repository-wide typecheck, test, and build were not repeated because no product code changed. ## Risks The hub is a dated static snapshot. It can lag behind the linked histories until an operator refreshes it. A changed history schema stops the build. Existing archives and grades are not modified. The published guide link is pinned to the reviewed commit so branch deletion cannot break it. Later builds can use master. ## Model Used OpenAI gpt-6-astra for implementation and review. OpenAI gpt-5.6-luna for documentation and independent skill checks. Both used repository tools and code execution. Context window sizes are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (latest commit pending) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (preceding head was 5/5; latest commit pending) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
422287eecd |
fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task messages, provider execution, and task outcomes. > - First-time user tests exposed gaps in recovery, completion permissions, message delivery, and Stop behavior. > - These gaps left usable output hidden, completed work waiting for bookkeeping, or safe work unable to continue. > - This pull request fixes the shared lifecycle and receipt paths while preserving process ownership and action checks. > - Users can continue work with accurate task state and durable messages. ## Linked Issues or Issue Description **What happened?** A stopped local Codex execution could remain blocked even after its processes had stopped and its complete transcript proved that no external action needed replay. Claude under Conservative permissions could fail to call task completion tools. Recovery could reuse an assistant item ID and overwrite prior output. A delivered comment could remain marked uncertain after navigation. Stop could look like Pause or a new recovery incident. Workspace contention could look like cancellation. A direct reply reopening Done could enter a clarification loop. **Expected behavior** Recover automatically only with verified termination and complete action receipts. Preserve answers and messages. Keep task completion available under Conservative permissions without broad tool access. Show crashes as Blocked, actual human decisions as In Review, and ordinary workspace contention as waiting. Stop the current response and allow a new direction. **Steps to reproduce** 1. Create ordinary response tasks with local Codex and Claude Code, then send follow-up messages through the task composer. 2. Interrupt a disposable local Codex runner during text-only work. Verify automatic continuation and retained output. 3. Stop a response, send a new request, answer a clarification, and reopen completed work with another message. 4. Navigate or reload while a comment submission is pending. Confirm the exact persisted request receipt settles it without removing newer draft text. 5. Run two tasks in a shared Daytona workspace. Confirm waiting does not appear as failure. **Paperclip version or commit** Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`. Current integration base: `6cef9743c`. Both operator-interruption and workspace-waiting guards are preserved; native restart and legacy permission rules remain documented. **Deployment mode** An isolated source-built test-drive instance, with real local Codex and Claude Code providers and disposable Daytona environments. Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163. This PR addresses additional failures from ordinary task journeys, including controller restart handoff and repeated warm sandbox setup. Historical task status reconciliation is excluded. ## What Changed - Persist runner ownership immediately at spawn and resume an explicitly adopted runner even when the controller crashed before the first driver checkpoint. Detach the controller safely across graceful restarts, including session startup. Prevent an old finalizer from suspending or signaling an adopted runner. Checkpoint idle warm sessions before shutdown. Preserve the same run and queued follow-up messages. - Scope saved legacy queue successor checks to the queue owner while preserving ordinary task locks, operator identity, assignment gates, and exactly-once delivery. - Preserve managed Codex credential files when an old session is detached for restart; normal owned cleanup still copies refreshed auth back and removes the scoped copy. - Reuse the bound warm shared sandbox and fully verify an existing staged provider pack before using it. This avoids repeated uploads when the pack is already valid. - Add a narrow local Codex replacement path with stopped-process proof, a closed transcript inventory, exact completion receipts, and fresh-session lineage. Preserve no-replay holds when evidence is incomplete. Recovery may clear only the same run's recorded Blocked status version; manual re-blocking and dependency changes invalidate that receipt, while queued comments do not. Later blocks stop scheduled, queued, and final dispatch; queued/final checks re-read dependencies even when the task status stays In Progress. - Permit only task delivery and human-input tools through the isolated Claude runner's exact task bridge. - Scope assistant item identity to the provider turn and ignore only authority-free Codex skill-change notifications during startup. - Reconcile composer submissions by client request ID across response loss, navigation, and reload. Retain text typed during delivery. - Keep acknowledged run-only Stop neutral and show workspace contention as waiting. Project exhausted native failures as Blocked. - Restore the guarded task-page retry action for failed legacy runs, including the server-supported explicit new-attempt path for stopped conversation adapters. Preserve native/process recovery holds and avoid promising Retry while a decision or execution gate hides it. - Refresh delivered artifacts and handle direct user replies that reopen completed work without a clarification loop. - Check the embedded PostgreSQL PID, data directory, and actual port before connecting or migrating. - Document accepted behavior and add focused regressions at lifecycle, route, transcript, and UI boundaries. ## Verification - Final head `fece606ac2` passes the complete GitHub CI matrix: **34 green checks, two expected Storybook skips, no failures or pending checks**, including `ci / verify`, `ci / e2e`, full runner verification, typecheck, build, every server/workspace shard, and all browser shards. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34727183287). Greptile is **5/5 with no open findings**. The final two commits only refine test fixtures; both affected suites pass 24/24 locally and in CI, with server typecheck green. - Complete local Vitest coverage uses the canonical groups/shards: all 635 general server suites, all 145 serialized suites, and all workspace packages. The aggregate began on `0a8001c18` while the final queue fix arrived: 23,903 passed, five failed, 87 skipped. The five port/socket/timing failures passed unchanged in follow-ups (60 tests in the exposure/file suites and 412 tests covering the serialized failures and unrun tails). The final queue/operator-identity suites separately passed 52/52. This is aggregate coverage plus explicit reruns, not a pristine single-command final-head run. - After integration with current master, queue/operator-identity/continuation suites passed 162/162 and affected UI suites passed 140/140. ACP Stop/continuation and legacy task/Inbox/message browser suites passed 9/9, including both task recovery Retry and thread Try again, automatic saved-message delivery, exactly one new run, Done, and retained output after reload. The default process Stop/Pause/Resume browser case passed (the native-provider case is opt-in and skipped by default). The complete Board attachment/receipt browser suite passed 11/11 on a disposable instance, covering both composers, exact receipts after lost responses, no replay, bound attachments, and newer drafts after reload. - Blocking-intent regressions cover pre-existing Blocked, a mismatched run/cause, an explicit manual re-block, changed dependencies, a queued comment after failure, and a block arriving between scheduling and provider dispatch. The negative cases reproduced before the fix. All 478 affected executor/recovery/dispatch tests passed; both database suites ran separately after availability-probe skips in the first combined command. The final late-dependency check passed all 143 affected recovery/dispatch tests (zero skips) after two new negative cases reproduced the bug. - Focused runtime regressions cover awaited runner ownership publication, authenticated adoption before the first checkpoint, old-finalizer detachment, idle and busy warm-session shutdown, rejected checkpoint propagation, provider-pack verification, and managed-Codex credential preservation. Four managed credential detachment cases reproduced the bug before the fix; normal owned cleanup still succeeds exactly once. - Live local Claude: SIGKILL 2.6 seconds into startup recovered the same run automatically in 53 seconds, then a normal follow-up completed in 24 seconds. SIGTERM 2.5 seconds into startup preserved the same run (54 seconds) and its queued follow-up (21 seconds). Answers remained visible and the task reached Done. - Live Claude Daytona: a warm follow-up retained its sandbox and fell from 121 seconds to 44 seconds. A separate cold turn took 127 seconds; after controller shutdown and checkpointing, its follow-up completed in 33 seconds with the same sandbox, workspace, native session, and runner. Both answers remained visible and the task was Done. - Other live journeys covered task completion and follow-up with local and Daytona Codex, local Codex crash recovery, Stop then new direction, clarification response, live artifact refresh, and shared-workspace waiting. - Validation limits: the opt-in native composer Stop/Pause→subtree Resume fixture exposes terminal/result ordering and subtree-cancellation attribution bugs that can leave a child task blocked; that new finding is assigned to a separate follow-up and is not claimed fixed here. Default CI skips this optional native-provider fixture. Managed-Codex credential handoff and the queue-agent integration use automated regression evidence. Cold custom provider-pack uploads still add startup latency. ## Risks - Automatic replacement remains deliberately narrow: local Codex, verified stopped identities, unchanged retained state, and a complete text/completion-only turn. Unknown actions, partial history, or changed ownership remain blocked. - Claude completion permission handling changes an upstream package patch. The exact isolated task bridge must remain pinned; unrelated tools keep their existing permissions. - New task failure projection changes user-visible status. No historical status backfill or database migration is included. - This is a broad lifecycle fix across server and UI. Live proof covers graceful local Claude restart during startup and idle Claude Daytona session recovery across controller shutdown. Live abrupt SIGKILL during local Claude startup also recovered the same run. Unknown ownership or missing action evidence still blocks reuse. Cold custom provider-pack uploads still add startup latency; this change avoids unnecessary repeat uploads. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, browser automation, and tool use. The exact hosted model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7e6d512597 |
fix(onboarding): make chief-of-staff hiring reliable (#13317)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first agent helps the board define work and hire other agents. > - That agent can have the general role while its instructions require hiring skills. > - Missing skills and blocked schema discovery make valid requests fail. > - Repeated confirmation and invalid waiting guidance can turn these failures into extra runs. > - This PR supplies the required skills, opens read-only schema discovery, and corrects the guidance. > - The agent can complete an authorized hire while company approval and duplicate checks still apply. ## Linked Issues or Issue Description Refs #13068 — the first-task onboarding flow that this change repairs. Refs #12029 — related drift between the sandbox allowlist and bundled hiring guidance. This PR adds schema access; it does not replace the earlier hiring-route fix. **What happened?** A general-role onboarding chief received hiring instructions without the core hiring skills. Sandbox requests to the documented OpenAPI endpoint failed. The agent then guessed question and hire payloads. The persona required new confirmation after validation errors and described waiting states that agents cannot set. **Expected behavior** A direct request authorizes the requested hire. The chief asks only for material missing details, uses valid API payloads, and completes the task. Formal company approval gates still apply. A saved human-input card gives the task a valid waiting state. **Steps to reproduce** 1. Create an onboarding chief with role `general` through the board. 2. Ask it to hire a friendly robot with a supplied name and responsibilities. 3. Check its assigned skills, schema requests, question cards, hire requests, and final task state. **Paperclip version or commit** Reproduced on the first-task onboarding implementation after #13068. The live local verification used this branch at `112f44610`. **Deployment mode** The original failure used a hosted sandbox with legacy Codex ACP. Live verification used an isolated local instance and real `codex_local` execution. Queue and HTTP/2 transport access is covered by automated tests. ## What Changed - Give board-created onboarding chiefs the existing core skills regardless of role. Preserve explicit skill version pins, including aliases. Keep ordinary general-agent defaults and authorization checks. - Allow exactly `GET /api/openapi.json` through both sandbox bridge transports. - Publish validator-tested question, free-text, hire, and waiting examples. Regenerate the runner API reference and capability inventory. - Clarify direct authorization, material ambiguity, and correction of confirmed pre-creation validation failures. Preserve uncertain-outcome reconciliation, duplicate protection, and company approval gates. - Align disposition instructions with agent permissions and the saved human-input waiting path. ## Verification - After rebasing onto current `master`: 69 targeted server tests, 110 queue/HTTP2 bridge tests, and 4 capability inventory tests passed. These cover core skill defaults, version pins, actor restrictions, schema access, published examples, hire validation, idempotency, and approval gates. Waiting recovery tests and live question flows also passed before the rebase. - `pnpm -r typecheck` and `pnpm build` passed again after the rebase. Frozen dependency installation and both generated capability checks passed. - Ran the full `pnpm test:run` suite. The initial run had 14 failed server files due to local database resource limits, a missing built test fixture, and socket failures. All 14 files passed after fixture repair and isolated retries. UI, CLI, workspace packages, database tests, and all 145 serialized server files passed. - Real one-request hiring replay: one hire, one successful run, task done in 2m16s. No repeated approval or recovery escalation. - Real two-turn browser conversation: start with an unspecified hire, then supply a name and friendly robot responsibilities. One clarification card, one hire, two successful runs, task done in 3m27s of execution. No failed writes, confirmation cards, or recovery actions. - Assigned the hired robot a welcome-message task through the browser. It produced a warm message under 100 words and finished in one successful 66-second run, with no questions or recovery actions. - The two-turn flow still asked an optional preferences question and gave a technical final reply. These are remaining presentation limits. - Greptile: 5/5 on `b71f83ba2`, with zero unresolved review threads. Fixed its generator finding and passed 1,655 published-example/runtime API tests plus server typecheck. All latest-head CI checks are green (32 passed; 2 unrelated Storybook checks skipped). The signoff-policy browser test initially timed out while waiting for an approver run. Its shard passed on one rerun without code changes. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34698211049). ## Risks - Onboarding chiefs receive more default skills. Ordinary general agents retain existing defaults, and explicit versions take precedence. - Prompt guidance can affect model behavior. The live replays are examples, not a guarantee that every model follows the guidance. - Retry guidance applies only when validation confirms that nothing was created. Uncertain outcomes still require checking existing agents. - No database migration or new public endpoint. Existing company boundaries, approval gates, and bounded recovery remain in force. ## Model Used OpenAI Codex, model `gpt-6-astra`, with reasoning, tool use, code editing, and live browser verification. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
250deab910 |
fix(runner): keep healthy native sessions alive (#13261)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native sessions keep a provider process and its work alive across control-plane operations. > - A hidden 15-minute turn deadline stopped work even when the agent timeout was zero. > - A one-hour runner lifetime and fixed connection lease added two more limits. > - Recovery also rejected goal commands because it reconstructed their startup summary with the wrong protocol version. > - This pull request removes implicit duration limits and renews authenticated leases in the harness. > - Healthy sessions can continue without model action or a user interface change. ## Linked Issues or Issue Description Refs #13092 and #12845. Related: #13163 covers sandbox recovery after app restarts; this change covers session duration and lease renewal. **What happened?** A native Codex session stopped after 15 minutes while a tool was still running. The agent had `timeoutSec: 0`. Recovery then rejected a `session.goal.get` startup command with `invalid provider startup ownership fence`. **Expected behavior** An unlimited session keeps working while its provider and authenticated controller remain healthy. Lease maintenance is transparent. Explicit timeouts, cancellation and revoked authority still take effect. **Steps to reproduce** Start a native session with `timeoutSec: 0` and run a tool beyond 15 minutes. Before this fix, the runtime cancels the turn. A recovery startup that uses a goal command also exposes the protocol-version mismatch. ## What Changed - Honor the agent turn timeout. Zero disables the timer. Long explicit durations use timer chunks to avoid Node timer overflow. - Default native runner lifetime to unlimited. Keep bounded startup, reconnect and control-operation deadlines. - Renew leases over the authenticated connection. Persist renewal before the reply. Validate identity, epoch and expiry. Handle duplicate requests and a lost reply on reconnect. - Freeze renewal during warm ownership transitions and terminal handling. - Validate persisted goal startup commands with protocol v2. - Add duration, renewal, ownership, recovery and real-process regression tests. Update runner protocol and recovery docs. - Add no UI components or controls. Renewal requires no model output or user action. ## Verification Current head: `348e369c35c5da8bb8be378f4b35dcf6f40882e7`. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34649767113). - All 32 checks pass on this head. The two Storybook checks are skipped as expected. CI includes full build, typecheck, runner verification, browser suites, server tests and the canary package dry run. - Greptile reports 5/5 on this head. All review threads are resolved, and the security scan passes. - Passed `pnpm -r typecheck` and `pnpm build` locally. - Passed 219 native-runtime and controller tests, including fake-clock tests for three weeks of renewal and 30-day explicit timeouts. Six denial tests confirm that renewal cannot extend expired, revoked or mismatched authority. - Passed 337 executor, cancellation and restart-recovery tests, plus 278 Rust runner-core library tests. - Passed a real runner with a silent fake Codex provider across its original lease expiry. Runner PID, provider PID, thread and active turn stayed unchanged. Warm-attach recovery tests also pass. - Passed all 83 plugin-worker tests and 159 of 161 workspace-runtime tests locally. The two remaining assertions passed with a canonical macOS temporary directory, as did the changed runtime fixture. The full affected server shard passes in CI. - An unchanged GitHub callback-ordering test failed once in CI, passed locally in isolation, and passed its one test-shard retry. The final CI summary is successful. - The full local `pnpm test:run` sweep was interrupted after dependency setup failures and load-related timeouts. Identified failing suites passed in isolated reruns after the dependency repair. The complete test matrix passed remotely in CI. ## Risks - Deploy the controller and runner together to enable renewal. Older peers keep their existing bounded lease behavior. - Unlimited runtime permits long resource use until completion, explicit cancellation, configured timeout or loss of valid authority. - Lease renewal changes authenticated protocol handling. Regression tests cover stale, revoked and mismatched authority, lost replies and warm handoff behavior. - Simulated multi-week tests and a real lease-boundary test do not constitute a weeks-long production soak. ## Model Used OpenAI GPT-6 through Codex, with repository inspection, code execution and TypeScript/Rust test tools. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
889947c238 |
feat: add experimental native chat connectors (#13038)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also ask agents for work in their existing chat tools. > - Each external conversation needs one task and a current authorized source. > - Retries, Stop, and provider failures must not duplicate work or expose private data. > - The first chat PR establishes the opt-in provider and data contracts. > - This PR adds experimental channel integration and its durable control plane. > - Users can request work from connected channels and inspect delivery in Paperclip. ## Linked Issues or Issue Description Refs #13100 and #13092. This is the second of exactly two chat PRs. Foundation #13100 is merged and changed 143 files. Runner prerequisite #13092 is also merged. This PR changes 400 files against master, below the 500-file review limit. It contains no wireframe images or HTML galleries. ## What Changed - Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat connections. Keep chat disabled unless the operator enables experimental chat connectors. Preserve the production GitHub tool connection and its normal setup path. - Bind each provider bot identity to one immutable Paperclip agent. Bind each admitted external conversation to one task. Paperclip owns tasks, runs, permissions, and audit records. - Add durable admission, per-conversation queues, questions, task controls, progress, final replies, images, files, and delivery receipts. Board comments remain internal unless explicitly sent to the channel. - Check current identity, provider reach, resource access, credentials, runtime generation, and exact source before provider effects. Keep private responses private. Never send raw reasoning, private logs, credentials, or tool arguments. - Hold uncertain sends for explicit audited resolution. Make Board Send-to-channel atomic and idempotent. Keep reconnect and setup credentials in Paperclip secret storage. - Preserve current native-runner authority across retries, lost acknowledgements, and recovery. Keep immutable input and completion contracts separate from newer user input. Receipt reconciliation cannot launch a provider. - Reconcile chat close/new ordering and provider-effect lock order. Audit resource access changes in the same transaction. Submit only the selected resource from each UI toggle so stale pages cannot undo unrelated access changes. - Drain Codex stdout before certifying process exit. Bound the drain with the existing shutdown grace. Preserve observed terminal authority without treating an undrained process as successful or reusable. - Incorporate master `018ca5da` with its ACP Stop, mobile task layout, runner packaging, and official lock changes. Preserve dedicated chat-answer continuations in both directions when ordinary queued comments are adopted after Stop. - Fence late adapter readiness behind an earlier Stop for the same run. Preserve verified cleanup for registered adapters. Handle single Stop, agent pause, duplicate Stops, and failure release without creating a false cancellation receipt. - Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact failed-chat retry authorization and lineage, retired question-source suppression, and the block on generic recovery that would discard the admitted source. Fresh deferred input retains its separate promotion path. - Incorporate master `2a05b5ed3` and its queue-admission extraction, simplified transaction ports, and separate runner CI job. Preserve exact durable receipts, actor separation, and dedicated-answer isolation through the new module. A failed receipt insert rolls back the accompanying deferred-wake merge. ## Verification Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are resolved. This successor fixes two test-harness boundaries exposed by CI: per-case route-module preparation and actual durable-save completion before intentional runner termination. Production code and all existing test/turn deadlines are unchanged. [Exact-head Greptile review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594) is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable findings or open review threads. [Fresh exact-head CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341) passes **all 24 jobs**, including Build and both required aggregates. Normal exact-head guarded merge was attempted and rejected by the remaining branch approval policy: CODEOWNER review is required and no human approval is present. Normal **squash auto-merge is enabled** as of September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified; no approval bypass or self-approval was used. Earlier-head results below remain historical evidence, not qualification of this successor. - Final exact-head Linux evidence: 995/995 chat integration cases; 36/36 agent-skills routes; 35/35 runner live-session cases, including real process kill/resume; 1948 runner Vitest cases with three existing benchmark/platform guards; 870/870 API-authority cases; and 104 browser cases with four existing optional skips. Rust, conformance/replay, full repository build, typecheck, canary, all server/workspace shards, and both required aggregates pass with normal CI concurrency. Earlier failed attempts remain recorded below. - Latest test-only qualification: 141/141 route/permissions/authentication cases pass in separate cold forks, with plain server types and independent review clear. The real-runner suite passes 35/35, with plain runner types and independent review clear. A controlled premature-save acknowledgement fails as expected; matching ownership/effect/process evidence, rejected saves, real turn outcome, test abort, and pre-kill liveness are covered. No local reproduction of the original CI scheduling failure is claimed. The preceding [CI run](https://github.com/paperclipai/paperclip/actions/runs/34479680858) passes 21/24 jobs, including all 995 Linux chat cases and browser aggregate (104 passed, four existing optional skips); only Build, the skills serialized shard, and the required verification aggregate fail. Its exact-head Greptile review was 5/5. Both failed job logs are retained. - Final fixture qualification: all eight focused Discord cases and all 995 chat integration cases pass. The exact modal statement/PID is observed before taking the real connection lock; the test then proves its actual blocking relationship before mutation. Original SQL execution, provider behavior, negative assertions, and 1s/15s timeouts remain unchanged. Independent review is clear and test/production hashes remain frozen. The preceding [CI attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777) passed 22 jobs, including Build/runner, typecheck, canary, all other test shards, and browser aggregate (104 passed, four existing optional skips); the two fixture failures and failed verification aggregate remain recorded, not relabeled as a pass. - Current queue-module composition: 308/308 recovery/batching/queue/Stop tests; 995/995 full chat integration; 89/89 module tests, including real PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary tests; plain server and UI types. All four actual local process/ACP browser paths pass in 1.4 minutes. Fresh databases, no skips or retries, stable reviewed source hashes. The initial boundary failure is retained; its no-op service wrapper was removed without changing recovery context or weakening the check. An exploratory standalone test-directory typecheck fails because its new upstream transformation config is not a standalone typechecking project; standard CI/build does not invoke it, and no configuration was weakened to suppress those diagnostics. - The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed [all 24 CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958) and exact-head Greptile review at 5/5. Required CODEOWNER review prevented its normal merge before master advanced again. - Final extracted-module composition: 307/307 recovery, batching, queue and Stop-control tests; 995/995 full chat integration; 49/49 module tests including eight PostgreSQL adapter cases; and 19/19 issue-update tests. Plain server types pass. All four actual local process/ACP browser paths pass in 1.3 minutes. Fresh databases, no skips or retries in these cohorts, frozen source hashes, and independent review clear. - The preceding head `3e4e1c1c` passes [all PR CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820), including Build and required `ci / verify` and `ci / e2e`. Both the original Rust failure and the previously load-sensitive lineage fixture pass with unchanged Linux concurrency. Master advanced afterward and required this reconciliation. - Final master composition: 448/448 focused UI tests, 186/186 adapter tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI, server, shared, and adapter types pass. Token gates and diff checks pass. Independent server and UI reviews are clear. - Stop-registration regression: both real-service cases fail against exact `a95` source and pass with the fix. The full corrected recovery/control suite passes 265/265. Duplicate-owner and failed-Stop controls also pass. Plain server types pass. The readiness barrier prevents provider startup without adding an acknowledgment to an already terminal run. - Final qualification strengthens terminal-field equality and repeats both affected cases successfully on a fresh database. All four actual local process/ACP browser paths pass again in 1.3 minutes, without skips or retries. The final screenshot shows Cancelled, a paused subtree, retained input, and no error toast. - Two new actual-service regressions fail before the merge fix. They prove that queued-comment adoption could consume a dedicated chat answer or add unrelated input to that answer. The fixed four-case cohort passes, including ordinary upstream continuation and adapter Stop controls. Full recovery passes 257/257. All four actual local process/ACP Stop browser flows pass in 1.4 minutes, without skips or retries, on a fresh database. - The unchanged runner artifact was qualified with 171/171 transport tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11. Six controlled reader tests prove the exit/drain repair. Its local serial Rust workspace passed 546 top-level cases plus two invoked helpers; the later passing Linux CI supplies default-concurrency evidence. - Prior exact-source full chat integration passes 995/995. Settings regressions cover concurrent stale pages, 501 destinations, pending state, rejected updates, and explicit retry. These deterministic tests do not prove live provider behavior. - Retained failed attempts and their causes are in the [qualification log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md). The first merge adapter run timed out while macOS slept for 290 seconds. Its unchanged repeat passed with a temporary sleep guard. No assertion, deadline, or CI gate was weakened. Review commands include `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh disposable databases. See the [browser runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md) for provider setup and separate live acceptance steps. ## Risks - This remains experimental. Deterministic tests and bounded live evidence do not establish every provider feature, tenant, permission layout, or media shape. Teams work-tenant qualification is still open. - Failed and uncertain provider effects remain visible and can require operator action. A transport receipt does not prove recipient visibility. - Native controller and runner artifacts must remain compatible. Preserve lease ownership, terminal authority, source binding, and quarantine during future changes. - Access and audit rows commit together, but activity notifications remain best-effort. This is not a new durable event outbox. - The PR operation does not deploy a live server, replace its runner, or change provider permissions. Remaining live qualification is documented in the [temporary handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md). ## Model Used OpenAI Codex assisted with implementation, tool execution, testing, and review. The work records `gpt-6-astra` assistance. The environment does not report a context-window size. No private reasoning traces are included. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fac07b42ad |
fix(runner): preserve durable native session authority across recovery (#13092)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner carries tool results and task output to the control plane. > - A lost connection must not change which run owns a result. > - A session must not become reusable while provider output is still pending. > - This pull request adds strict recovery evidence and bounded drain barriers. > - It preserves current PRP version negotiation and session-goal support. > - The benefit is safer reuse of native sessions after a transport failure. ## Linked Issues or Issue Description Refs #13038. This is the first of two stacked pull requests. It contains the native runtime prerequisites. The second pull request contains the experimental chat-channel integration. It preserves the provider identity and typed terminal-failure contracts in #13074 and the durable recovery work in #13075. **What happened?** Native session failures could leave retained provider events, incomplete tool results, or warm handoff state that was not safe to reuse. A later run could observe output from an earlier authority. **Expected behavior** Recovery must preserve exact run, tool, process, artifact, and lease evidence. Uncertain or corrupt state must fail closed. A successful close must prove that retained provider output is settled. **Steps to reproduce** Run the transport and control-plane regressions. They hold and drop authenticated frames, fail durable writes, and restart fresh controllers and runner processes with retained state. Provider executables are local test fixtures. ## What Changed - Preserve pending provider cleanup and semantic-result evidence across session close and restart. - Add an authenticated warm handoff with exact old and new identities, durable receipts, and completion acknowledgement. - Drain retained provider events under the cumulative acknowledgement fence. - Reject corrupt tool-result contracts without unsafe provider replay or reusable checkpoints. - Keep ordinary PRP v1 sessions and current session-goal behavior. Require negotiated PRP v2 and acknowledged native session evidence before warm authority rotation. - Preserve late semantic inputs and exact durable result receipts until close can prove settlement. - Add transport, crash-window, artifact, checkpoint, and final-output regressions. - Deduplicate resolved execution delivery under the current issue lock. Reuse the exact existing successor after concurrent scans or a lost acknowledgement. Preserve newer operator evidence. - Persist idle provider integrity/capacity failures before process retirement, retain permanent model-rejection classification, and keep external question identifiers out of task instructions. - Expose only the context source on native status events. Keep thin dispatch projections compatible without exposing the complete context. ## Verification - Review-fix revision: 128 runtime-context/native-session tests, five idle-failure/adjacent Rust cases, 24 warm crash-window cases, three startup-notification/close cases, and five attach/backlog cases passed. The security and idle-failure cases were first reproduced failing. - Prior merged revision: runner production build, TypeScript typecheck, complete Rust workspace tests and formatting passed; 272 focused runner tests and two real PostgreSQL regressions passed. - Earlier full runner runs and CI Build failed on missing semantic-result fixture receipts, stale local provider fixture bytes, startup-notification ordering, and a confirmation-loss fixture that could accidentally send its final ACK. Each cause was reproduced and corrected without relaxing production authority or close assertions. These earlier runs are retained as failures, not represented as passing verification. - The first local repository-wide run failed before later phases because the isolated install omitted PostgreSQL's native-library aliases; it also encountered an unrelated occupied-port fixture. Those results are retained, not represented as a passing run. - Exact `335b2ee52709afb3885d4d6ebb2a3ece4b5864d6`: the complete runner suite passed 1,888 tests, with 10 existing skips. The full Rust release workspace passed with serial test scheduling. The unchanged parallel Rust run hit the five-second 300-descendant fixture deadline; that failure is retained. No deadline or assertion was relaxed. - The resolved-execution regression suite passed 57 tests, including concurrent delivery, lost acknowledgement, superseded authority, and newer operator evidence. Plain server typecheck passed. The duplicate-delivery cases were first reproduced failing. - Prior exact `335b2ee52709afb3885d4d6ebb2a3ece4b5864d6` CI passed all required jobs and Greptile reported 5/5. Its local general-server run passed 7,208 tests but failed one responsibility fixture; later phases did not run. The fixture started the next wake while its bounded handoff was active. It also used nonexistent comment IDs, which hid the current stored-message-author identity rule. The updated tests use real message authors, preserve task ownership, and await exact automatic handoffs. No production identity policy changed. - Current head `aa39275a1f300f7d1a0b16cd0885eea567cff6b0` includes current master and the native context-source projection. The focused identity/status cohort passed 27 tests and plain server typecheck passed. Fresh full repository tests, types, build, required CI, and Greptile review are pending. Final results will be updated before merge. - This is deterministic local-provider evidence. It is not a claim of complete live-provider qualification. ## Risks - This changes authenticated recovery and close ordering. The TypeScript transport and runner binary must be built from the same revision. - Failed or incomplete evidence intentionally prevents reuse and can require a fresh run. - PRP v1 ordinary/cold sessions remain supported. A v1 connection lease cannot upgrade in place. A current v2-capable runner held on a v1 lease was qualified through owned-process retirement/join, fresh bootstrap on the same old authority, v2 observation/ACK, then warm rotation. Legacy binary replacement and adopted-owner migration are not qualified by that test; rollout must not present them as automatic same-lease upgrades. - This pull request has no database migration or chat-channel activation. The second pull request keeps the channel feature experimental. ## Model Used OpenAI Codex assisted with implementation, tool execution, tests, and reconciliation. The existing implementation records OpenAI `gpt-6-astra` assistance. The current environment does not report a context-window size. No private reasoning traces are included. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
54a99d8840 |
fix(evals): make the chat viewer the default published Evalbook (#12952)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Direct Runner evals retain evidence across model configurations. > - Evalbook already has a grid and a read-only Runner Lab chat viewer. > - Public projection stripped the view and selected a second plain result page. > - This change uses the existing viewer for public and private results. > - The data access differs, but the presentation does not. ## Linked Issues or Issue Description Refs #12931, #12945. Related open runtime-contract PR #11634 does not contain this report-only change. **What happened?** The public direct-eval campaign opened plain result pages. The access-controlled artifact used the chat viewer. Users could not follow the same recorded interaction from the published grid. **Expected behavior** Every newly generated Runner Evalbook opens the existing chat viewer. The grid and durable run history remain. Public evidence has explicit redactions. **Steps to reproduce** Open campaign gha-34062394019-1 from the direct-eval history. Click a result, then compare its plain page with the corresponding Actions artifact. **Paperclip version or commit** Reproduced at |
||
|
|
83987210d6 |
fix(runner): align direct eval provider setup with qualified runtimes (#12945)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner direct live evals test semantic tools against a mock control
plane.
> - The first complete AWS campaign exercised 358 cells.
> - It exposed setup differences from the working full-stack harness.
> - This pull request corrects those direct-harness differences.
> - It preserves production permission defaults and the full-stack
workflow.
## Linked Issues or Issue Description
**What happened?**
Native Codex cells could not find a global codex executable. ACPX denied
unattended tool requests and lost valid provider usage receipts.
AgentCore hit a 30-second facade timeout while its worker allows 120
seconds for delivery. Several models reported a native run result
without updating the separate mock task state.
**What did you expect?**
The direct harness should use the pinned executable, explicit test
permissions, and a timeout compatible with the provider delivery
contract. Its instructions should explain which operation changes mock
task state.
**Steps to reproduce**
Run the full Runner Direct Live Protocol Evals workflow. Baseline
campaign:
https://github.com/paperclipai/paperclip/actions/runs/34059009921.
**Paperclip version**
Master at
|
||
|
|
fee8d8dc39 |
fix(runner): repair direct live provider bootstrap (#12932)
## Thinking Path > - Paperclip runs AI agents through qualified provider backends. > - The direct live eval workflow builds one immutable Runner runtime for every matrix cell. > - The workflow reinstalled the packed Runner with npm. > - That install discarded pnpm patches and selected provider dependencies outside the qualified lock. > - The first pnpm deployment model also placed its virtual-store marker at the wrong level; a real deployment keeps `.pnpm` beside the scoped Runner package. > - AgentCore enforced the current context-aware harness but the direct eval CLI did not supply the production v3 runtime context that harness requires. > - This pull request preserves the qualified dependency graph, resolves the real deployment layout, and makes direct evals exercise the production runtime-context contract. > - The benefit is that live eval cells reach their provider turn with the same artifacts and context contract that Paperclip qualified. ## Linked Issues or Issue Description Refs: #12931 **What happened?** The full direct live eval campaign failed every ACPX cell during `session.open`. The portable runtime had an incorrect dependency root. Its npm install also discarded the qualified ACP server patches. AgentCore cells first failed because Runner enforced `aws-agentcore-harness-v1` while the provisioned stack and eval profile use `aws-agentcore-harness-context-v2`; after aligning that revision, the direct eval CLI still omitted the required v3 runtime context. **Expected behavior** The direct eval runtime must preserve the frozen pnpm dependency graph and patched provider bytes. Runner, server validation, OpenAPI, and the deployed AgentCore stack must use one qualification revision. Direct eval attempts must supply the same immutable native runtime-context contract as production. **Steps to reproduce** 1. Dispatch `Runner Direct Live Protocol Evals` from `master`. 2. Select an ACPX Claude, ACPX Codex, or AgentCore roster. 3. Observe a pre-turn provider bootstrap failure. **Paperclip version or commit** `d96452db059338b329b458ba8fe359fef72f1363` **Deployment mode** GitHub Actions on the RunsOn Linux x64 fleet. ## What Changed - Build the reusable direct-eval runtime with `pnpm deploy --prod`. - Resolve ACPX dependencies from the actual scoped-package layout of a self-contained pnpm deployment. - Align AgentCore configuration and qualification checks on `aws-agentcore-harness-context-v2`. - Materialize a minimal immutable v3 runtime context for each isolated direct eval attempt. - Add workflow, package-authority, runtime-context, Rust, and server regression coverage. - Document the qualified packaging, runtime-context, and AgentCore revision contracts. ## Verification - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/runnerd-codex-transport.test.ts` (70 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/cli/eval-session-contract.test.ts` (14 tests) - Focused Runner contract tests (36 tests) - Focused server profile tests (47 tests) - Focused Rust managed-provider and native-selector tests (19 tests) - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs` - `actionlint .github/workflows/runner-protocol-live-evals.yml` - A local `pnpm deploy --prod` produced both qualified ACP server digests. - A Linux reproduction of the first follow-up smoke identified the real deployment root and the missing AgentCore runtime context. ## Risks The AgentCore revision change rejects profiles that still use the obsolete v1 value. This is intentional because the provisioned context-aware harness and current eval profile use v2. Direct eval prompts now receive the same fixed runtime-context preamble as production, so behavior scores may move; that is the intended qualification surface. The workflow package layout changes, but tests assert the new entrypoint and dependency root. This change does not modify the browser full-stack E2E workflow. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The context-window size is not exposed in this session. The model used extended reasoning, repository tools, code execution, Docker-based Linux reproduction, and GitHub Actions diagnostics. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (for example, `docs/...` or `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
af8439a70b |
feat(runner): restore direct live eval campaigns and reports (#12909)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner executes agents through native and managed provider drivers. > - The direct live eval layer had drifted from the current Runner contracts. > - The old local workflow did not provide a complete parallel campaign or durable report history. > - The Runner also needed current native OpenCode and OpenRouter qualification. > - This pull request restores the direct campaign, corrects the runtime gaps that the campaign found, and adds safe hosted Evalbook history. > - The benefit is repeatable model comparison against an immutable Runner and eval source revision. ## Linked Issues or Issue Description Refs #11297 Refs #11634 **What existing behavior does this improve?** This improves the direct live `paperclip-runner` eval workflow, provider execution contract, and static Evalbook reporting path. **Current behavior** The direct evals do not have one maintained full campaign on current `master`. OpenCode has no qualified multi-model OpenRouter roster. Parallel provider bursts can compact committed events before the transport observes them. Local reports do not have a separate safe S3 history index. **Proposed behavior** Run one immutable roster-plus-case matrix. Use the shared paid AWS runner fleet. Keep raw artifacts access-controlled. Publish a sanitized canonical Evalbook report under the separate `runner-protocol-evals` S3 prefix. Keep immutable campaign directories plus root history, latest, and latest-green pointers. **Reason and benefit** Maintainers can compare native Codex, native OpenCode, ACPX, Claude Managed, and AWS AgentCore behavior over time. They can inspect failures without mixing this direct protocol layer with browser full-stack E2E. **Breaking changes** None. The new workflow and S3 prefix are additive. The existing Runner full-stack E2E workflow and report remain separate. ## What Changed - Added a trusted two-shard direct live workflow for up to 393 roster-plus-case cells. - Reused the numeric actor allowlist, protected paid environment, and RunsOn fleet controls from Runner full-stack E2E. - Added immutable Runner and eval revision resolution, exact credential boundaries, bounded retries, and cost ceilings. - Added a public report projection that removes sessions, transcripts, tool payloads, state, traces, raw failures, remote profile identities, and credential-shaped values. - Added additive S3 history under `runner-protocol-evals`, with immutable campaigns and mutable root index pointers. - Added native OpenCode model injection and current OpenRouter pricing contracts. - Fixed direct eval completion, workflow execution, semantic discovery, warm-attach state reset, executable binding, and event-burst handling. - Kept Runner browser full-stack E2E behavior and publication separate. - Documented local and hosted direct eval operation. ## Verification - `pnpm --filter @paperclipai/paperclip-runner test:runner-protocol-eval-publish` — 15 passed. - `pnpm --filter @paperclipai/paperclip-runner build:typescript` — passed. - `actionlint .github/workflows/runner-protocol-live-evals.yml .github/workflows/runner-full-stack-e2e.yml` — passed. - Local current matrix at the revision in [paperclip-evals#17](https://github.com/paperclipai/paperclip-evals/pull/17) — 323 cells across 10 enabled configurations completed. - Final local current matrix — 269 passed, 11 behavior failures, and 43 expected macOS-only ACPX platform failures. - Targeted Runner checks — 13/13 eval-session tests, 15/15 publisher/security tests, and package typecheck passed; complete PR CI is green, including all browser E2E shards. ## Risks - Paid live campaigns can consume provider budget. Actor authorization, exact per-cell ceilings, protected environments, and explicit schedule enablement bound this risk. - Public reports can leak provider data. The workflow publishes only a separately projected report and validates every file before upload. - The new workflow cannot publish until it is present on the default branch. This pull request does not change the existing `runner-full-stack-e2e` publication path. - The campaign is large. It uses two GitHub matrices and caps combined concurrency at the shared fleet limit. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex on GPT-5. The exact deployment ID and context-window size are not exposed. The model used reasoning, code editing, browser inspection, repository tools, and live provider execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
7b094724e6 |
fix(runner): recover native sessions across restarts (#12845)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner keeps durable run and provider state outside one server process. > - A server restart can leave that runner alive or can interrupt it after a provider checkpoint. > - The old startup path used handoff intent and PID evidence, but it did not reconstruct native ownership. > - That gap could block the issue, create a replacement run, or start duplicate provider work. > - This pull request adds durable same-run recovery for coordinated and uncoordinated restarts. > - The benefit is exact recovery of the run, runner, session, provider, steering, and finalization state. ## Linked Issues or Issue Description Refs #9628. That pull request added earlier local-adapter hot-restart work. This change adds native PRP authority reconstruction and same-run provider resume. Refs #10935. That pull request handles missing hot-restart snapshots. This change also supports hard restarts with no snapshot. Refs #11624. That pull request prevents unsafe retry after an adopted legacy process exits. This change reconciles native terminal evidence before provider recovery. Refs #12070. That pull request improves process liveness checks. This change also binds recovery to a process-start fingerprint and fails closed on ambiguity. **What happened?** The server could record hot-restart intent, but startup did not rebuild native runner ownership. A live runner could not re-register its PRP authority. A dead runner could not resume the exact native and provider session on the same heartbeat run. Generic recovery could then block the issue or create replacement work. **Expected behavior** A live native runner must reconnect with the same PID and logical identities. A dead runner must resume the same durable session and heartbeat run with only a new operating-system PID. A proposed or terminal result must finalize once before any provider turn starts. Ambiguous process or session evidence must stay blocked without a signal or duplicate spawn. **Steps to reproduce** 1. Start a Paperclip Runner heartbeat and wait for an active provider turn. 2. Restart only the Paperclip server, with or without a hot-restart marker. 3. Observe that the old startup path does not reconstruct the native control-plane authority. 4. Kill both the server and runner after a provider checkpoint. 5. Observe that the old path cannot resume the exact native session on the original heartbeat run. **Paperclip version or commit** The defect was reproduced from commit `1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto the current `master`. **Deployment mode** Local development and self-hosted server deployments that use the local Paperclip Runner. ## What Changed - Added correlated hot-restart requests and version-compatible native handoff fields. - Added controller boot identity, process-start identity, controller generation, recovery state, request id, and bounded history to the native finalization ledger. - Added transactional recovery claims for live-runner reattach, dead-runner resume, and incomplete bootstrap. - Added fail-closed ownership takeover rules and process identity validation. - Added live runner adoption to the local runner transport without a duplicate spawn. - Added same-run provider checkpoint resume and legacy retry-row compatibility. - Reconciled proposed and terminal results before runner or provider recovery. - Bound the HTTP and PRP listener before startup recovery and delayed scheduling and generic reapers until classification completes. - Added restart-aware health diagnostics, run-log recovery transitions, durable runner diagnostics, and bounded shutdown finalizer draining. - Moved restart-survivable diagnostics into runner-owned, pre-redacted bounded writes; raw stdout and stderr are never persisted. - Added process-start fencing for controller, runner, and provider PIDs; startup classifies every candidate without an implicit cap. - Added crash-recoverable, contention-safe development restart-request coordination and failed-startup listener cleanup. - Added a credential-free real-process restart suite for eight restart, scale, and identity scenarios. - Documented native restart operation, persistence, diagnostics, and verification. ## Verification - The documented native restart commands passed. They ran eight real-process/database recovery scenarios and the live runner adoption transport test. - Native executor tests passed: 111 tests. - Heartbeat recovery tests passed: 124 tests. - Hot restart, health, and shutdown tests passed: 52 tests. - The broader affected server suite passed: 350 tests. - Focused native recovery and startup tests passed: 49 tests. - Runner transport and control-plane tests passed: 63 tests. - Runner-owned diagnostic tests passed for write-time bounding, credential redaction, private file modes, and raw stream non-persistence. - Development restart coordination tests passed: 11 tests. - Database migration checks and the partial-application/replay regression test passed. - Server, database, and Paperclip Runner typechecks passed. - `git diff --check` passed. - Full Paperclip PR CI passed, including build, canary, all five general server shards, all five serialized server shards, all three browser E2E shards, workspace suites, and release-registry verification. - Greptile completed at 5/5 with no outstanding findings, recommendations, follow-ups, or open review threads. ## Risks - Moderate risk. This changes startup ordering and ownership transfer for active native runs. - The migration adds nullable columns and does not rewrite existing rows. - Recovery fails closed when process or durable session identity is incomplete or contradictory. - The first implementation supports the local Paperclip Runner. Remote targets keep their existing behavior. - The real-process suite covers cleanup and asserts that no runner or provider process survives each test. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime did not expose a more specific model revision or context-window size. Repository editing, shell execution, database tests, and real-process test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0f94521017 |
fix(runner): restore local session and task integrity (#12721)
## Thinking Path > - Paperclip is the control plane for agents that perform work. > - Paperclip Runner connects durable provider sessions to individual task runs through PRP. > - Provider continuity and per-run authority are different lifetimes. > - The existing implementation mixed those lifetimes and lost event metadata between provider frames, runnerd, persistence, API sanitization, and the task thread. > - That caused failed continuation, missing progress and Plans, duplicate replies, hidden failures, and unsafe recovery. > - This repair gives every heartbeat fresh authority, preserves qualified provider-session continuity, and restores one lossless presentation path without changing direct adapters. ## Linked Issues or Issue Description **What happened?** A second native heartbeat could reuse tickets, leases, command receipts, sequence state, and run identity from the first heartbeat. Provider phase and item identity could be lost before the UI read them. Redaction could corrupt protocol discriminators while still missing malformed credential tails. The task thread could fold progress into the final response, hide failures, or show more than one final answer. Native Codex also exposed approval modes that do not yet have a durable approval bridge. **Expected behavior** Each heartbeat uses a new PRP authority epoch. Codex and OpenCode preserve exact qualified provider sessions; ACPX emits an explicit continuity event when its qualified process-replacement policy is used. Every accepted provider event is presented, classified as internal, or surfaced as unsupported. The task page shows chronological progress, reasoning summaries, activity, Plans, interactions, terminal failures, and exactly one final reply. Direct adapters retain their existing path. **Steps to reproduce** 1. Enable the unified experimental Paperclip Runner setting. 2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex agent. 3. Run response, Plan, structured-question/resume, restart, cancellation, and failure scenarios. 4. Reload the task while active, waiting, failed, and settled. 5. On the old implementation, observe stale run authority, missing classifications, incomplete output, or duplicated/folded replies. **Paperclip version or commit** The repair is based directly on `master` at `87d05e194b643810d16d20612115acd01d735d43`. **Deployment mode** Local development with the embedded database. Related work: Refs #12616, #12646, #12666, #12685, and #12700. ## What Changed - Rotates PRP control-plane, outbox, ticket, lease, command, receipt, and sequence authority for each heartbeat while carrying forward only a validated provider-session identity. - Reads `control-plane-state.json`, validates both durable schemas and lifecycle values, resumes coherent current runs, archives qualified settled authority, and quarantines malformed or mismatched scoped state without moving ambiguous live legacy state. - Preserves Codex provider phase and stable item identities so commentary remains progress and only `final_answer` becomes final. - Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning lifecycle mapping. - Makes ACPX normalization lossless for visible reasoning, tool lifecycle metadata, stable bounded identities, Plan revisions, structured requests, failures, and qualified process replacement. Only the compatible terminal assistant message is promoted as final. - Applies schema-aware redaction before generic JWT-shaped detection and scans every diagnostic string leaf. Malformed raw/escaped quoted credential tails are redacted in both server and durable Rust state. - Restores snapshot-style chronological task presentation, expandable tool activity, inline Plan cards, visible waiting/resume/cancel/failure states, and exactly one final answer. - Makes `never` the only qualified native Codex permission mode and rejects unsupported persisted native modes with remediation. OpenCode and ACPX policies remain intact. - Keeps the unified experimental Runner setting as the only enablement flag. Onboarding and direct Codex, Claude, and OpenCode stay on their legacy execution/finalization paths. - Adds cross-language goldens, authority/recovery/fault coverage, exact response/count assertions, and native plus legacy acceptance scenarios. ## Verification - Pull-request GitHub Actions run Rust formatting/tests, TypeScript checks, server/UI tests, builds, protocol drift checks, browser E2E, and security scans. - A separate workflow-only validation ref is pinned directly on this PR head and runs the 35-cell paid local matrix: three core scenarios plus structured-question resume and restart/resume for native Codex, native OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode. Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315 - Acceptance requires exact single visible replies, monotonic sequences, matching envelope discriminators, one semantic terminal, one run terminal, no unresolved interaction, no duplicate mutation, no secret leakage, provider continuity, and zero native rows for direct adapters. - Per maintainer direction, tests are running in GitHub Actions rather than on the slower local host. Only formatters and static diff checks were run locally. ## Risks - Recovery from old or partial filesystem state is sensitive. The repair fails closed, preserves active or unverifiable authority, and quarantines only state whose scoped ownership is safe to move. - Provider event formats can change. Closed validators and boundary goldens turn new or malformed events into visible diagnostics instead of silent drops. - Shared task presentation could affect direct adapters. Runtime-fact gating plus the direct-adapter matrix protect the existing path. - Managed and remote providers are not qualified here. Shared code continues to compile and fail safely, but live qualification is deferred. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact deployed snapshot and context-window size are not exposed to this task. It used agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] The paid local-provider matrix is green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5716fe907e |
test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner subsystem executes agent work across local and managed provider backends. > - The lower pull requests restore the task runtime, provider backends, and managed-provider control plane. > - The restored system needs repeatable full-stack checks before it can ship safely. > - Paid live checks also need clear access, cost, and secret controls. > - This pull request adds acceptance, live evaluation, chaos, and release gates for the restored runner stack. > - The benefit is measurable runner parity with safer release decisions. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change covers runner tests, release workflows, server contracts, and evaluation tools. **Problem or motivation** The runner stack did not have one complete acceptance surface for native Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could miss provider drift, task-view regressions, cost-policy errors, and destructive cleanup errors. **Proposed solution** Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid workflows. Add live evaluation, chaos, cost-limit, redaction, and release contract checks. Add AWS AgentCore infrastructure and guarded provisioning tools. Keep the native runner experimental flag off by default. **Alternatives considered** We considered manual smoke tests only. They do not give repeatable evidence and they do not protect release branches. We also considered one large pull request. The stacked pull requests keep each review below the Greptile file limit. **Roadmap alignment** This work supports the shipped Cloud / Sandbox agents milestone and the shipped Agent evals & feedback milestone in `ROADMAP.md`. Related stack: - #12699 adds managed provider backends and lifecycle support. - #12691 adds qualified OpenCode and ACPX provider backends. - #12685 restores task runtime rendering and steering. ## What Changed - Add the runner full-stack harness with 57 catalog cells and 60 unit tests. - Add a Daytona runner image with digest-pinned base images and base-aware image-content checks. - Add guarded live evaluation and chaos workflows with a fixed 40-execution matrix; live and full-stack paid schedules now run only on Sundays or by manual dispatch. - Add in-flight reported-usage cost stops, post-turn cost caps, exact-threshold failure classification, secret redaction, retry classification, and actor authorization. - Reattach stream and hard-budget listeners before restart-recovery continuations so restored paid sessions cannot bypass in-flight interruption. - Preserve OpenCode usage and cost across tool-loop messages and turns while exposing an explicit current-run delta to durable accounting. - Keep PNG/WebM evidence in access-controlled artifacts only, reject SVG, and publish only pruned inert structured per-attempt evidence. - Add AWS AgentCore infrastructure, provisioning checks, and smoke tools; reject unsafe model identifiers, require exact stack ownership markers, and make failed-stack replacement explicit. - Add evaluation-session contracts and capability reports. - Add release workflow checks for immutable action pins, frozen dependency installs, exact weekly cron shape, paid-run guards, provider-secret isolation, and chaos test paths. - Reauthorize the original and triggering numeric actor IDs as the first step of every provider-secret job, including partial reruns, before checkout or provider access. - Give each full-stack matrix cell only its matching provider credential, expose Daytona only to Daytona cells, and disable shared dependency caches anywhere paid credentials or OIDC write access are present. - Protect the legacy manual E2E workflow with the same default-branch, allowlist, environment, and per-job authorization boundary. - Rotate live-eval candidates by week and retain 120 days of compatible history so the seven-week trend window remains viable. - Restore the root runner-acceptance commands and reconcile reported snapshots, raw receipts, and terminal usage without double counting or losing late usage. - Mark ACPX token deltas exact only when every budget field is present, keep cumulative cost/request authority separate, reject non-USD cost labeling, and include thought tokens in output-token budgets. - Keep `enableNativeRunner` off by default. The acceptance harness enables it only in its isolated test instance. ## Verification Passed locally: - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm test:runner-acceptance:typecheck` - `pnpm test:runner-acceptance` (19 tests) - focused OpenCode proxy, driver, runnerd transport, live-session, and turn-stream tests (106 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/clean-room-server.test.ts` (22 tests) - `pnpm test:e2e:runner:typecheck` - `pnpm test:e2e:runner:unit` (62 tests) - `node --test scripts/__tests__/release-verify-workflow.test.mjs` - `pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals` (22 tests) - `pnpm -r typecheck` - `pnpm build` - `node --test packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs` (6 tests) - `git diff --check` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core --lib --locked` (161 tests) - focused ACPX provider-event tests (10 tests) - The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged. I did not run paid live provider jobs or provision AWS resources. Those checks need credentials and can create cost. ## Risks The paid workflows can create provider cost. They require an allowlisted original and triggering actor, the protected `runner-e2e-paid` environment, explicit opt-in variables, and cost limits. The four provider credentials exist only in that master-only environment, which requires allowlisted reviewer approval and disables administrator bypass; repository and organization Actions scopes contain no copies. Provider usage arrives after a billable request, so the live guard cannot prevent one request from crossing a threshold. It interrupts immediately on the first reported threshold hit and permits no continuation. Visual evidence can contain secrets rendered as pixels. PNG/WebM remain only in access-controlled workflow artifacts; SVG and per-attempt XML are excluded, and S3/Pages receive a pruned structured dashboard. The AWS scripts can create cloud resources. They use explicit commands, least-privilege roles, KMS encryption, saved nonsecret metadata, and explicit teardown. This pull request does not enable the experimental native runner for existing instances. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The model used extended reasoning, tool use, code execution, and parallel subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5458940a6e |
feat(runner): add offline evaluation tooling (#12653)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs repeatable evaluation contracts. > - Evaluation code must stay separate from provider launch and production orchestration. > - Offline fixtures need stable compatibility, scoring, traceability, and report rules. > - Published Runner consumers need only the supported evaluation contract surface. > - This pull request adds offline evaluation tooling and a workspace-private matrix kernel. > - The benefit is deterministic evaluation without credentials or paid provider calls. ## Linked Issues or Issue Description Refs #11297 This pull request extracts the offline evaluation unit from the earlier aggregate Runner work. ## What Changed - Add a workspace-private, provider-neutral evaluation matrix kernel. - Add the public `@paperclipai/paperclip-runner/evals` compatibility and native execution contracts. - Add fail-closed runnerd artifact and protocol compatibility checks. - Add deterministic workflow catalogs, scoring, traceability, and report generation. - Add sanitized Codex, OpenCode, and ACPX fixtures. - Add package-boundary and clean-consumer checks. - Add the eval package manifest to the Docker dependency stage. - Add the generated protocol fixture digest without changing the lockfile. ## Verification GitHub Actions must run: - Runner TypeScript and Rust type checks. - Runner unit and protocol tests. - Evaluation kernel tests. - Workflow traceability checks. - Clean-consumer and package-boundary checks. - Repository test, type-check, build, policy, and security gates. No local test command was run. The repository owner requested GitHub-only verification. ## Risks This is a large greenfield review surface with 51 files. The code does not launch a live provider or load credentials. Package and protocol drift fail closed. The workspace lockfile remains under the existing CI-owned process. ## Model Used OpenAI Codex with the GPT-5 agent model. The work used high reasoning, repository inspection, tool use, and parallel code review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0bdbf61564 |
feat(runner): add secure remote transport (#12639)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner gives native runs a durable and governed execution path. > - The lower stack PR adds authenticated remote execution targets and provider ingress. > - The Rust daemon currently accepts only loopback plaintext WebSocket connections. > - Remote Codex needs authenticated WSS dialing and provider-ingress listener mode. > - This pull request adds the bounded Rust transport contract. > - The benefit is a secure transport layer for the Codex remote vertical slice. ## Linked Issues or Issue Description Refs #12638. Refs #12616. Refs #12352. **Subsystem affected** Paperclip Runner Rust transport and remote runner networking. **Problem or motivation** The runner daemon cannot connect to a public control plane with TLS. It also cannot accept a provider preview connection on the run-bound ingress path. **Proposed solution** Add WSS with native trust roots and an optional private CA bundle. Add a fixed authenticated listener mode for provider ingress. Advertise the exact transport contract through build metadata. **Alternatives considered** Plaintext public WebSocket connections would weaken the transport boundary. A general listener would expose more network surface than the run-bound provider ingress requires. **Roadmap alignment** This work supports the Cloud and Sandbox agents milestone. It also supports self-healing native runs. ## Stack - Lower merged PR: #12638. - This PR contains only its 13-file delta against `master`. - Later stack PRs add the task workspace and administrator UI. ## What Changed - Added WSS dialing with rustls and native certificate roots. - Added an optional bounded private CA bundle that augments native roots. - Kept plaintext WebSocket dialing restricted to loopback addresses. - Pinned resolved dial addresses for the process lifetime. - Added a fixed `0.0.0.0:43127` listener with an exact run-bound path. - Rejected listener queries, ambiguous paths, and WebSocket extensions. - Kept frame and message size bounds. - Added bounded reconnect grace and exponential jitter. - Retried bootstrap failures only before authentication proof transmission begins. - Kept post-proof failures fail-closed and bounded the welcome exchange at two seconds. - Added runnerd build metadata for the versioned transport contract. - Updated Rust dependencies and `Cargo.lock` only for TLS and certificate handling. - Did not add provider dispatch, Pi, AWS, `pnpm-lock.yaml`, migrations, or workflows. ## Verification - GitHub Actions will run Cargo formatting, Rust tests, repository tests, typecheck, build, security, and policy gates. - Rust tests cover URL validation, listener path validation, build metadata, durable recovery, and the existing Codex provider path. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check master...HEAD` passes. - The delta contains 13 files. ## Risks - TLS and listener changes affect the runner trust boundary. - Public plaintext transport remains rejected. - The listener uses one fixed port and one exact run-bound path. - PRP authentication remains required after the WebSocket upgrade. - The optional CA file uses the existing private-file checks and a 4 MiB limit. - This PR does not enable another provider or change direct adapters. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
560e7e48b5 |
feat(runner): add SDK and developer tooling (#12608)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |