mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 20:50:08 +02:00
codex/plugin-task-execution
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3b550c80fa |
fix(codex): correct startup trust, history reads, and resume usage (#13110)
## Thinking Path
> - Paperclip runs Codex locally and in remote sandboxes.
> - The runner must preserve startup configuration and session identity.
> - Missing project trust can disable repository configuration.
> - Full-history requests use deprecated provider fields.
> - Resume usage describes old work and must not become new run usage.
> - This change corrects startup trust, state reads, and usage
classification.
## Linked Issues or Issue Description
**What happened?**
Normal Codex runs could show repository-trust and history-deprecation
warnings.
Resume could report the preceding turn's token snapshot as a late-turn
warning.
The historical last-usage value could also be attributed to the new run.
**Expected behavior**
Trust the server-selected startup root in isolated configuration. Read
lightweight
provider state and paginated evidence. Use historical cumulative usage
as a
baseline without a new charge or user-facing warning.
**Steps to reproduce**
1. Start a native Codex task in a selected repository.
2. Finish the turn and resume the provider thread.
3. Inspect provider notices, history requests, and per-run usage.
4. Repeat startup and cold resume inside a Daytona sandbox.
**Paperclip version or commit**
Codex CLI 0.153.4 is the pinned runtime and reproduced baseline.
Replayed onto master at
|
||
|
|
fdf8c8464d |
feat(runner): add managed provider backends (#12699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner provides durable, provider-neutral agent execution. > - The current stack supports qualified local providers but omits the managed provider paths from the integration branch. > - Claude Managed Agents and AWS AgentCore need explicit profile qualification, durable recovery, usage accounting, and cleanup controls. > - This pull request adds those managed backends as the third part of the Runner parity stack. > - The benefit is managed execution without weakening the default-off Runner rollout gate. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: Runner, server orchestration, database profiles, CLI, and adapter configuration UI. **Problem or motivation** The current Runner stack cannot select or execute the managed Claude Agents API or AWS Bedrock AgentCore Harness backends. It also lacks qualified profile storage and recovery checks for those remote resources. **Proposed solution** Add qualified managed and remote profiles, API and CLI management, exact provider selection, durable lifecycle handling, cumulative usage accounting, bounded cleanup, and retention acknowledgement. Keep `enableNativeRunner` default-off. **Alternatives considered** A direct copy of the old integration branch was rejected because its provider contracts, model values, credential flow, and migration history no longer match the current base. A single large parity pull request was also rejected because stacked review keeps each subsystem bounded. **Roadmap alignment** This continues the existing Runner architecture and rollout work. It does not introduce a separate execution system. **Additional context** This pull request is based on the merged #12691 and #12685 stack. It also closes the delayed security-review findings reported on #12691 by binding qualified ACPX and OpenCode launch artifacts to the bytes actually executed. A GitHub search for managed agent, AgentCore, and Claude managed work found no duplicate public issue or pull request. ## What Changed - Add Claude Managed Agents and AWS AgentCore provider executors to runnerd. - Add qualified managed and remote profile storage, routes, OpenAPI contracts, CLI commands, and migration 0237. - Validate profile ownership, enabled state, exact qualified revision, model, agent version, and secret binding before persistence and recovery. - Persist durable provider session and owned skill state for restart-safe cleanup. - Reconcile uncertain create responses and delete remote sessions before owned skills. - Track cumulative provider usage and enforce positive session spend caps. - Recover interrupted AgentCore usage at the next turn boundary by charging the prior invocation ceiling exactly once; keep the session gated until an explicit monotonic budget raise. - Isolate AgentCore AWS configuration from host profiles and credential-process/SSO configuration while preserving workload identity. - Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact paths; remove the ambient executable override. - Snapshot and content-verify ACPX and OpenCode commands, scripts, and provider executables before launch. Linux executes sealed inherited descriptors; macOS uses authenticated private snapshots with retry-safe rematerialization at the spawn boundary. - Persist canonical ACPX and OpenCode launch-profile digests, reject drift across fresh recovery, and make recovery failures sticky. - Close and journal unsafe ACPX active-turn recovery before any provider bootstrap or reconnect. - Add managed provider fields to the Runner configuration UI and permission projection. - Preserve the default-off `enableNativeRunner` experimental flag. ## Verification - `pnpm -r typecheck` - `pnpm build` - Focused managed server, database, CLI, Runner TypeScript, Rust, Claude, AgentCore, ACPX, OpenCode, process-supervisor, and durable-recovery tests passed. - `cargo test -p paperclip-runner-core --lib --locked` (160 tests) - `cargo check --workspace --all-targets --locked` - Native Codex integration tests passed (60 tests); native provider tests passed (7 tests); server native-runtime tests passed (87 tests). - Verified-launch replacement, nested-spawn retry, exact-version, profile-drift, sticky-failure, and no-bootstrap active-recovery tests passed. - `git diff --check` - The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust workspace lockfile adds the approved `rustix` dependency used for safe descriptor handling while `#![forbid(unsafe_code)]` remains enabled. ## Risks - The provider APIs can change while they are in beta. Exact qualification and fail-closed recovery checks limit drift. - Remote cleanup can fail after a partial create. Durable ownership inventories and retry-safe deletion preserve recovery state. - Migration 0237 adds profile tables. The generated migration and snapshot pass the repository migration checks. - Managed execution can incur provider cost. Positive default spend caps and explicit retention acknowledgement limit accidental use. - An interrupted AgentCore invocation without final metadata is conservatively charged to its active session ceiling. This can overstate cost, but cannot undercount it; later work requires an explicit budget increase. - Linux qualified launches use sealed memory descriptors. macOS lacks executable-descriptor APIs, so the runner uses owner-only private snapshots and minimizes linked-path lifetime; hostile same-UID processes remain outside the documented local-host trust boundary. - The global Runner feature remains default-off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0bdbf61564 |
feat(runner): add secure remote transport (#12639)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner gives native runs a durable and governed execution path. > - The lower stack PR adds authenticated remote execution targets and provider ingress. > - The Rust daemon currently accepts only loopback plaintext WebSocket connections. > - Remote Codex needs authenticated WSS dialing and provider-ingress listener mode. > - This pull request adds the bounded Rust transport contract. > - The benefit is a secure transport layer for the Codex remote vertical slice. ## Linked Issues or Issue Description Refs #12638. Refs #12616. Refs #12352. **Subsystem affected** Paperclip Runner Rust transport and remote runner networking. **Problem or motivation** The runner daemon cannot connect to a public control plane with TLS. It also cannot accept a provider preview connection on the run-bound ingress path. **Proposed solution** Add WSS with native trust roots and an optional private CA bundle. Add a fixed authenticated listener mode for provider ingress. Advertise the exact transport contract through build metadata. **Alternatives considered** Plaintext public WebSocket connections would weaken the transport boundary. A general listener would expose more network surface than the run-bound provider ingress requires. **Roadmap alignment** This work supports the Cloud and Sandbox agents milestone. It also supports self-healing native runs. ## Stack - Lower merged PR: #12638. - This PR contains only its 13-file delta against `master`. - Later stack PRs add the task workspace and administrator UI. ## What Changed - Added WSS dialing with rustls and native certificate roots. - Added an optional bounded private CA bundle that augments native roots. - Kept plaintext WebSocket dialing restricted to loopback addresses. - Pinned resolved dial addresses for the process lifetime. - Added a fixed `0.0.0.0:43127` listener with an exact run-bound path. - Rejected listener queries, ambiguous paths, and WebSocket extensions. - Kept frame and message size bounds. - Added bounded reconnect grace and exponential jitter. - Retried bootstrap failures only before authentication proof transmission begins. - Kept post-proof failures fail-closed and bounded the welcome exchange at two seconds. - Added runnerd build metadata for the versioned transport contract. - Updated Rust dependencies and `Cargo.lock` only for TLS and certificate handling. - Did not add provider dispatch, Pi, AWS, `pnpm-lock.yaml`, migrations, or workflows. ## Verification - GitHub Actions will run Cargo formatting, Rust tests, repository tests, typecheck, build, security, and policy gates. - Rust tests cover URL validation, listener path validation, build metadata, durable recovery, and the existing Codex provider path. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check master...HEAD` passes. - The delta contains 13 files. ## Risks - TLS and listener changes affect the runner trust boundary. - Public plaintext transport remains rejected. - The listener uses one fixed port and one exact run-bound path. - PRP authentication remains required after the WebSocket upgrade. - The optional CA file uses the existing private-file checks and a 4 MiB limit. - This PR does not enable another provider or change direct adapters. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
3aa2065d08 |
feat(runner): validate structured question responses (#12420)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Rust runner now owns a bounded ACPX session and a fail-closed turn lifecycle > - Provider questions pause a turn and must return structured answers to the same persisted question set > - JSON Schema validates the wire shape, but it cannot validate identifiers and constraints across two documents > - Unknown questions, invalid choices, and malformed custom answers must fail before any provider receives them > - This pull request adds only the package-local response validator and tests > - The benefit is a small trust boundary that later request-resolution code can use without changing production selection ## Linked Issues or Issue Description Refs #12419 ## What Changed - Validate `paperclip.question_response.v1` against its versioned JSON Schema. - Bound serialized responses to 768 KiB before validation. - Require answer identifiers to match the exact persisted question set. - Require answers for required questions and reject unknown question identifiers. - Enforce text, single-select, and multi-select answer modes. - Match the existing TypeScript numeric syntax, including decimal, exponent, hexadecimal, octal, and binary input. - Match ECMAScript trimming exactly, including BOM whitespace while rejecting Unicode NEL rather than inheriting Rust-specific whitespace behavior. - Enforce known options, custom-answer policy, text length, pattern, and numeric constraints. - Validate duplicate option IDs, inverted bounds, and dynamic patterns before answer lookup so malformed optional questions fail closed even when unanswered. - Match JavaScript UTF-16 code-unit length semantics for text constraints and the 100,000-unit response-field bound. - Preserve the public optional `recommended` question-option field in the versioned schema, generated schema bundle, and Rust validation path. - Return typed validation errors for malformed inputs without panics. - Export the validator from the Rust runner core. - Add table-driven tests for valid, mismatched, malformed, oversized, and numeric-boundary responses. - Document the package-local structured-response boundary. - Add `num-bigint` 0.4 and `num-traits` 0.2 as direct runner-core dependencies for exact arbitrary-length radix parsing and one-step JavaScript Number rounding; update only the package-local runner Cargo lockfile. - Do not change the repository PNPM lockfile, workflows, runnerd selection, server behavior, UI, or migrations. ## Verification - Replay base: `9a9fdf06ee4142f77427db30efccc4c43056f64b` (`master` after #12419 merged). - Exact replay head: `fad92b3fb348b66ddb10dde44b7b060e55c4fe96`. - Stable patch ID: `4d6ffbd519dd081f7ea530977cd965bd4569fc75`; this is the prepared two-commit delta plus the focused cross-language parity fix found during replay review. - The exact delta is 10 files, 712 additions, and 2 deletions, all in `packages/paperclip-runner`. - The package-local `packages/paperclip-runner/runner/Cargo.lock` records the two direct runner-core dependencies; their already-resolved versions and checksums are unchanged. - The question-set schema source, generated TypeScript schema bundle, and protocol manifest hash are updated together; the schema SHA-256 is `42b5441a3d388851dacb6e4500dfd4a17d878eded2e724228078b647e7440d3f`. - GitHub Actions run `33366812025`, attempt 2: **PASSED** on the exact replay head (23/23 jobs passed; a failed-job-only retry cleared one unrelated ACPX runtime-host timeout). - Greptile: **5/5** on the exact replay head with zero unresolved review threads; Superagent, Socket, Snyk, and contributor-trust checks also passed. - No local test result is claimed. GitHub Actions is the authoritative verification environment for this replayed revision. ## Risks - The validator compiles the embedded response schema for each submission. Responses are user-paced and bounded, so this keeps the slice simple without affecting a hot event path. - The persisted question set is the source of truth for identifiers and constraints. A malformed persisted set fails closed. - Numeric input follows the existing structured-question contract, including JavaScript-prefixed syntax. Optional whitespace-only answers are rejected instead of being treated as an omitted value. - Error messages identify the invalid field but do not include answer text. - No production path invokes this validator in this pull request. Request resolution remains the next slice. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5.6, agentic reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
3ba0e7f64f |
feat(runner): add durable semantic tool bridge (#12378)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The runner package has a reviewed semantic action catalog and dispatcher > - The Rust runner process needs the same fail-closed authorization boundary > - Provider calls must remain correlated and idempotent across durable recovery > - Input and result values must satisfy the authorized operation schemas > - This pull request adds a package-local durable semantic tool bridge > - It does not advertise tools to Codex or enable the Paperclip Runner adapter ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner/runner` semantic tool authorization and correlation. **Problem or motivation** The Rust runner needs a durable representation of the run-scoped tools that the control plane authorizes. It must reject unknown operations, catalog drift, invalid values, and conflicting duplicate calls or results before a provider integration can use those tools. **Proposed solution** Add a serialized provider tool bridge. Validate the authorized catalog and its JSON Schemas. Validate each call and result. Keep pending and completed identities so retries are idempotent and conflicts fail closed. **Alternatives considered** Trusting provider arguments would bypass the run-scoped catalog. Validating only in TypeScript would leave the Rust process without a recovery-safe authorization boundary. Adding provider behavior in this pull request would make the review unit too broad. **Roadmap alignment** This adds a package-local safety boundary for the Codex-first runner path. It does not enable a new adapter or change an existing direct adapter path. ## What Changed - Added the versioned authorized-tool, pending-call, and result contracts. - Added canonical SHA-256 catalog binding and drift rejection. - Added JSON Schema compilation and input and response validation. - Added duplicate-call and duplicate-result idempotency with conflict rejection. - Added bounds for catalogs, schemas, values, and retained call identities. - Added the Rust `jsonschema` dependency and its Cargo lock entries. - Added focused tests for authorization, recovery, envelopes, bounds, and conflicts. ## Verification - `cargo fmt --manifest-path packages/paperclip-runner/runner/Cargo.toml --all -- --check` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core` (64 tests) - `cargo clippy --manifest-path packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core --all-targets -- -D warnings -A clippy::manual_is_multiple_of -A clippy::filter_map_bool_then` - `pnpm -r typecheck` - `pnpm build` - The repository test runner also reached unrelated server worktree suites. Those suites fail on the current macOS worktree with database deadlocks and filesystem fixture assumptions. This pull request does not change those files. The applicable GitHub checks remain the handoff authority. ## Risks The main risks are accepting a tool that the run did not authorize and replaying a conflicting provider result. The bridge validates the catalog, operation identity, JSON Schema, call identity, and result identity before it changes durable state. The new Cargo dependency is package-local. This pull request changes no GitHub workflow and no pnpm lockfile. ## Model Used OpenAI Codex with GPT-5 and repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b76e36d6cf |
Add durable PRP transport and recovery (#12100)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The package-local runner can supervise a local process, but it cannot yet survive a broken controller connection. > - A production transport must authenticate both peers without putting the bootstrap secret on the wire. > - Commands and events must remain bounded, ordered, and recoverable across reconnects and crashes. > - Retrying an uncertain side effect is unsafe, so indeterminate outcomes must fail closed instead of running twice. > - This pull request adds those transport and recovery guarantees inside the runner package only. > - The benefit is a durable PRP boundary that can be reviewed before any provider or server integration exists. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This pull request extends private transport infrastructure in `packages/paperclip-runner`. **Problem or motivation** The local runner introduced by #12095 has no authenticated network handshake, durable outbox, reconnect lease, cumulative acknowledgement, or crash-safe command journal. A dropped connection could otherwise lose an event or tempt a controller to repeat a side effect whose outcome is unknown. **Proposed solution** Add an authenticated PRP v1 WebSocket transport, encrypted frames, lease-based reconnects, a bounded durable event outbox, cumulative acknowledgements, and an idempotent command journal. Preserve pending commands before execution and classify the crash window as indeterminate so an uncertain side effect is never repeated automatically. **Alternatives considered** The combined runner branch implements transport together with Codex, semantic tools, and server coordination. That change is too large for one review unit. Keeping transport in memory would make reconnect and crash recovery unverifiable. Re-running a pending command after restart would weaken the at-most-once side-effect boundary. **Roadmap alignment** This work supports the governed tools and self-healing run direction in `ROADMAP.md`. It does not add a production provider, server endpoint, adapter, feature flag, or user-facing behavior. **Additional context** Refs #12095 and #11962. Pull request #12095 was squash-merged first. This branch has been rebased onto the resulting `master` commit, and its current delta is 13 files. ## What Changed - Added a loopback-only WebSocket connection policy with one-time DNS resolution and pinned reconnect addresses. - Added an HMAC mutual-authentication handshake that never sends the bootstrap ticket over the socket. - Added AES-256-GCM secure frames with per-direction keys, monotonic counters, and session-bound authenticated data. - Added one-use bootstrap-ticket handling and lease-based reconnect validation with expiry, revocation, and epoch checks. - Added a private, symlink-resistant state directory with atomic, synchronized state replacement. - Added a bounded durable event outbox, priority-zero reserve, cumulative acknowledgements, and reconnect replay of only the unacknowledged suffix. - Added a bounded command journal with contiguous sequence enforcement, persistent results, and deterministic duplicate responses. Duplicate replay requires a SHA-256 match over the complete canonical command. - Persisted commands before their effects. A crash after persistence but before result storage returns an indeterminate terminal result and does not execute the command again. - Migrated pre-fingerprint command journals by compacting through their persisted controller cursor. Legacy redelivery fails closed instead of reconstructing an incomplete identity or repeating an uncertain effect. - Added strict limits and validation for frames, state, results, outbox entries, command history, and redacted diagnostics. - Added a transport-only `paperclip-runnerd --connect-url` mode. It handles lifecycle commands and rejects provider commands because no provider is present in this pull request. - Added a full disconnect-before-ack fault test that reconnects with the lease, replays identical command and event state, and proves the effect ran once. - Kept provider transports, semantic tools, server integration, and production runtime selection out of this pull request. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` passes. - TypeScript contract tests pass: 8 Node tests and 44 Vitest tests. - Rust tests pass: 33 unit tests, 3 public durable-recovery integration tests, plus the existing 2 local-runner and 3 process-supervisor tests. - The disconnect-before-ack, lease reconnect, duplicate command, malformed state, unknown command, bounds, and crash-window tests pass. - Rust conformance and replay parity checks pass against the shared PRP fixtures. - `cargo clippy --workspace --all-targets -- -A clippy::filter-map-bool-then -D warnings` passes. The narrow allow covers an unchanged replay implementation from the preceding contract pull request. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm check:token-gates` passes. - `git diff --check` passes. - The delta against `master` is 13 files. The package lockfile is unchanged. - `pnpm test:run` completed locally with 4,686 passing tests, 19 skipped tests, and 30 failures in 8 unchanged server test files. The failures reproduce the established local macOS path-alias, listener, port-range, and workspace-runtime baseline. No changed-file test failed; Linux CI remains the repository handoff authority. - Storybook visual regression is not applicable because this pull request changes no UI or story files. ## Risks Production behavior is unchanged because no server code starts or connects to this transport. The main risks are secret disclosure, forged or replayed frames, state corruption, unbounded disk growth, duplicated side effects, and incorrect recovery. Mutual authentication, encrypted counter-bound frames, private atomic state, explicit bounds, cumulative acknowledgements, a durable command journal, fail-closed indeterminate recovery, and fault-injection tests cover these risks. I checked `ROADMAP.md`. This change is private transport infrastructure for planned control-plane work. It does not duplicate a shipped or public product surface. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6b20cc97cc |
Add local fake runner supervision (#12095)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs a small local process model before it can connect to a production provider or server. > - The TypeScript PRP contracts now define the expected replay behavior. > - A second language implementation must produce the same result from the same fixtures. > - Local child processes also need bounded input, bounded output, and complete descendant cleanup. > - This pull request adds a package-local Rust runner, a scripted fake harness, and deterministic parity checks. > - The benefit is a testable process boundary with no production Paperclip behavior change. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This pull request adds private test infrastructure to `packages/paperclip-runner`. **Problem or motivation** The PRP contracts have no second implementation on `master`. There is also no small harness that can prove process cleanup, command idempotency, terminal reconciliation, or bounded JSONL handling without a production provider. **Proposed solution** Add a minimal Rust workspace. Add a local runner process, a scripted fake harness, a bounded process supervisor, and Rust conformance and replay checks. Keep all binaries package-local. Do not connect them to the Paperclip server. **Alternatives considered** The combined runner branch includes provider transports, durable networking, SDKs, labs, and server behavior. That change is too large for this review unit. A TypeScript-only harness would not test cross-language contract parity. **Roadmap alignment** This work supports the governed tools and self-healing run direction in `ROADMAP.md`. It does not add a user-facing runtime, adapter, endpoint, or rollout flag. **Additional context** Refs #12091 and #11962. Pull request #12091 was merged before this branch opened. This branch is based on the current `master`. Its delta is 25 files. ## What Changed - Added a minimal locked Rust workspace with only `serde` and `serde_json` dependencies. - Added a package-local `paperclip-runnerd` local mode and a scripted fake harness. - Added bounded controller input, harness input, subprocess output queues, line sizes, log retention, script sizes, script steps, and command history. - Added contiguous controller and harness sequence checks and equivalent-command replay handling. - Added process-group supervision that cleans up child processes and remaining descendants after forced or natural harness exit. - Added runner-owned terminal reconciliation for success, failure, interruption, cancellation, controller closure, and protocol failure. - Added Rust conformance output and deterministic replay summaries for the shared PRP fixtures. - Added fake scripts for success, failure, interruption, interaction, duplicate terminal output, process cleanup, and oversized output. - Added package scripts and documentation for the Rust and cross-language checks. - Kept provider transport, server integration, semantic tools, and production runtime selection out of this pull request. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` passes. - TypeScript contract tests pass: 8 Node tests and 44 Vitest tests. - Rust tests pass: 20 unit tests, 2 local-runner tests, and 3 process-supervisor tests. - The Rust conformance and replay parity checks pass against the shared fixtures. - The natural-exit and forced-exit tests confirm that the harness and its worker process are stopped. - The oversized-frame test confirms that a harness frame above the configured limit is rejected. - `pnpm -r typecheck` passes after the final rebase to `master`. - `pnpm build` passes after the final rebase to `master`. - `pnpm check:token-gates` passes. - `git diff --check` passes. - The delta against `master` is 25 files. The package lockfile is unchanged. - `pnpm test:run` completed locally with 4,686 passing tests, 19 skipped tests, and 30 failures in 8 unchanged server test files. The failures are local macOS path-alias, listener, port-range, and workspace-runtime baseline failures. No changed-file test failed, and every applicable Linux CI shard passes. - Storybook visual regression skipped intentionally because this pull request changes no UI or story files. ## Risks Low production risk. No server code invokes the new binaries. The package remains private. The main risks are process leaks, unbounded local input, and cross-language drift. Bounded queues and sizes, process-group cleanup tests, fixture manifests, and parity checks cover these risks. I checked `ROADMAP.md`. This change is private test infrastructure for planned control-plane work. It does not duplicate a shipped or public product surface. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |