mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 03:08:10 +02:00
codex/plugin-task-execution
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fc6304dfe5 |
feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path > - Paperclip manages AI agents, tasks, permissions, and execution budgets. > - Paperclip Runner gives each provider the same admitted task and tool authority. > - OpenAI Dot runs outside the local process tree and needs asynchronous work delivery. > - The merged MCP gateway supplies OAuth consent and signed event delivery. > - A personal assistant grant cannot safely stand in for an assigned agent. > - This pull request adds a separate Dot agent connection and a durable Rust Runner bridge. > - The operator can assign work to Dot and inspect its accepted work, tool receipts, and result. ## Linked Issues or Issue Description **Agent or provider** OpenAI Dot, as an experimental provider of the existing Paperclip Runner adapter. **Why this adapter is useful** An operator can assign normal Paperclip tasks to an existing Dot. Dot can read its mailbox, request work on an assigned task, use admitted task tools, and submit a result. Paperclip keeps company scope, checkout, approvals, known budget limits, and activity attribution. **How the agent is invoked** A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the assignment. The Rust Runner owns the durable turn and operation receipts. The first release supports self-hosted instances with a local Runner controller. **Additional context** This extends the merged public MCP gateway from #14846 and the assistant invitation and device-consent work from #14933. This also integrates the merged assistant tool and configuration expansion in #15380. Dot retains its dedicated agent resource and cannot receive personal configuration permission. The public assistant connection remains a personal connection. ## What Changed - Add a durable Rust Dot provider and its TypeScript Runner driver. - Add closed PRP v3 external-provider operations and native execution input v6. - Add company-scoped pairing, mailbox, assignment, and operation records. - Reuse merged browser/device consent, client metadata verification, webhook admissions, refresh, secret rotation, and warm-standby gates. - Keep Dot scopes, issuer, grants, event workers, and tool access separate from personal assistant access. - Add Dot configuration, pairing, readiness, and consent UI. Keep agent grants out of the personal Connections entry. - Regenerate the Dot-only migration after master. Preserve published gateway migrations. Make the new migration safe to reapply. - Document setup, recovery, accounting limits, evidence, and remaining account qualification. - Reverify reconnect callbacks and wake outstanding work with a fresh mailbox reference; preserve the existing assignment and operation receipts. - Clean up Dot bindings and waiting runs on OAuth revoke and refresh-token replay. Old grants cannot revoke replacement bindings. - Restore the pairing reference when an unsaved agent form is reopened; document board-only pairing routes in OpenAPI. - Accept a clean Rust exit after the acknowledged shutdown receipt. Unexpected exits still require recovery. - Clear the cached binding after a successful revoke so a failed connection refresh cannot restore it. - Add production-component Storybook states and screenshots for pairing and connection review. All preview account data is synthetic. - Persist normalized completion, serialize Dot turns and durable work admission, and poll subscription readiness. - Serialize mailbox writes and cursor reads; retain paused fence acknowledgement without task authority. - Authorize admitted review runs without changing the worker assignee. Include the fenced assignment ID in production stop notices. ## Verification - This PR integrates master `4a8178e9c`. Dot migration `0317_messy_famine.sql` follows the published history and is safe to reapply. The merge preserves the reserved migration connection, batch-commit handling, private task checks, task monitors, and native accounting. - Local workspace typecheck, full build, and UI token gates pass. The server typecheck passes after the review fixes. Database and native executor regressions pass. - All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot driver tests pass. The tests cover native document writing and finalization, durable replay, queue admission, mailbox ordering, admitted reviews, stale authority, production stop references, and paused acknowledgements. - Current head `d0e7e0626` passes all 57 checks: 53 pass and four are intentionally skipped. This includes full typecheck, build, tests, Rust Runner verification, browser E2E, release verification, and Canary Dry Run. Greptile rates this exact head 5/5. All review threads are resolved. - The full local root test run is slower than the sharded CI run and has not completed. The full CI test gates pass on the current commit. Focused local regressions pass. - Real-account pairing and event delivery on this base commit remain unqualified. Live account and setup proof are recorded in the follow-up #15414. The following screenshots use synthetic preview data. They show the production pairing component and do not qualify a real account or the full agent setup journey.   ## Risks - This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP and Paperclip Runner. The separate experimental-settings follow-up in #15414 replaces this environment flag with saved operator settings. - Dot does not expose provider token usage or cost. The operator must acknowledge external billing. Known Paperclip budget gates still apply. - Cancellation fences Paperclip authority. It does not confirm that Dot stopped all external activity. - Assigned skill files and third-party MCP bindings are unsupported and reject admission. There is no mounted workspace, model selector, or provider thread identifier. - Hosted agent-broker and remote controller deployments are not qualified. - The new migration follows the merged master history. Existing prototype databases still need the normal master migration history before this Dot-only migration. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window size are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, and test inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
7944ed3d97 |
fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Native runner agents need governed tools, durable runtime state, and useful execution evidence. > - A first activity trace waited 53.467 seconds even though tool activity took 6.274 seconds; provider input arrived before the server API call executed. > - Native agents also need a safe way to hire teammates without asking the model to rebuild runtime configuration. > - This pull request separates the observed ACP input-stream window from the actual server `tool.execute` span and adds a server-owned native hire contract. > - The benefit is clearer latency evidence and safer native teammates with existing approval, auth, and company boundaries preserved. ## Linked Issues or Issue Description Related Daytona provenance work is in [#13814](https://github.com/paperclipai/paperclip/pull/13814). No duplicate public PR was found for this combined timing and native-hire change. **What existing behavior does this improve?** Native runner agents can use governed tools and request hires. The server did not expose a safe native hire operation that reused the caller's validated runtime settings. First-activity traces also mixed provider input timing with server tool execution timing. **Current behavior** A native hire must construct a separate runner configuration. Full configuration copying could expose paths, instructions, secrets, or sessions. Timing evidence could make a provider or MCP identity join appear proven when the trace did not contain that join. **Proposed behavior** The native `hire_agent` operation accepts identity and persona inputs. The server sends `adapterType: "paperclip_runner"` with `inheritRuntimeFrom: "caller"`, then copies only validated provider, model, permission, lifecycle, and bounded execution settings. It inherits and validates the default environment, derives the managed AI binding through existing normalization, preserves approval and permissions, and creates fresh child instructions. Caller secrets, paths, prompts, and sessions are excluded. Provider events now include the optional boolean `inputUpdated`, with Rust forwarding support. Timing evidence separately records the ACP input-stream window and the actual server `tool.execute` activity. It does not claim a provider or MCP join without matching evidence. **Reason and benefit** Native agents can hire teammates that start with the caller's approved execution policy. Operators retain company boundaries, auth rules, approval gates, and requalification. Reviewers can distinguish provider streaming time from server API execution time when diagnosing first-activity delays. **Breaking changes** None for existing hires or tool calls. `inheritRuntimeFrom` is optional and only applies to same-company native agent callers. Conflicting explicit runtime settings are rejected. The provider event field is optional for existing producers. ## What Changed - Added the native `hire_agent` protocol action, catalog entry, API contract, and runner authority checks. - Added `inheritRuntimeFrom: "caller"` validation and a closed native runtime inheritance allowlist. - Preserved managed AI binding normalization, default-environment validation, approval snapshots, permissions, requalification, and fresh child instructions. - Added provider `inputUpdated` schema support and Rust forwarding. - Added first-activity and server tool timing evidence with conservative identity-join handling. - Added route, authority, provider-event, sidecar, API, catalog, and Rust-focused tests. - Kept private Honeycomb links, raw traces, and local result paths out of this description. ## Verification Focused checks passed: - 458 timing/session checks. - 61 native hire inheritance checks. - 20 hire authority checks. - 1,741 API checks. - 106 catalog checks. - 54 provider sidecar checks. - 12 Rust provider checks. Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2). R1 stopped at missing Docker image setup. The final trace is available at https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces. Latest-head CI passed all required build, typecheck, Rust, static, Vitest, serialized-server, workspace, chat, and E2E jobs. The focused local checks listed above passed; the broad local suite was not run before the live evaluation, while CI provides the full repository verification. ## Risks - Timing fields describe separate observed windows. They do not prove a provider or MCP owner without a valid trace join. - The inheritance allowlist must stay synchronized with native runner configuration fields. - Approval snapshots include resolved safe inherited settings and should be reviewed when native configuration fields change. - The focused local suite is narrower than the full repository suite; latest-head CI covers the broader repository checks. > Roadmap review: `ROADMAP.md` places this work within Paperclip's bring-your-own-agent direction. It extends existing native runner hiring and observability behavior. ## Model Used OpenAI GPT-6 (exact serving model ID is not exposed), with extended reasoning and repository tool use; GPT-5.6 Luna assisted with focused implementation and verification work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7ed122911b |
Add end-to-end session goals to Paperclip Runner
Add capability-aware slash-goal controls, durable provider goal state, PRP v2 negotiation, autonomous goal execution, and safe local session recovery. Integrate with current master, preserve provider session identity, and verify the browser goal/chat/replacement/clear workflow and unsupported-agent rejection. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
560e7e48b5 |
feat(runner): add SDK and developer tooling (#12608)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1ee738cf48 |
feat(runner): project durable ACPX events (#12424)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Rust runner now owns validated, scoped ACPX reducer events and safe session suspension > - Durable PRP transport must receive provider-neutral events rather than sidecar-native envelopes > - Semantic calls and questions must retain the exact run, session, turn, item, and provider-request authority used by the durable command stream > - Terminal, result, assistant, process, and diagnostic events also need one reviewed projection boundary > - Permission requests remain impossible under the pinned Codex policy and must fail closed if they reach projection > - This pull request adds only that package-local projection without selecting ACPX in runnerd ## Linked Issues or Issue Description Refs #12422 ## What Changed - Add a validated durable ACPX event projection context bound to one run, normalized session, turn, and item. - Pass already normalized activity events through without reintroducing provider-native envelopes. - Project authorized tool calls into canonical semantic input receipts with exact correlation and content digests. - Project structured questions into provider-neutral `paperclip.runtime_request.v2` events. - Preserve both the public projected request identity and the original provider request identity so responses resolve the exact sidecar request. - Project dynamic semantic operation results as `semantic_tool.result`; only reserved finish/block operations may propose the run result. - Project semantic completion results into `run.result.proposed`. - Project terminal-flushed assistant messages on the final channel and turn terminal states into existing provider-neutral event families. - Project sanitized process metadata and diagnostics into bounded harness diagnostics. - Validate runtime-request origins against their strict durable shape and fall back from empty optional titles to a valid question prompt. - Reject invalid identities, projected-identity collisions, unstable semantic receipt identities, permission requests, and cross-turn projection fail closed. - Add integration coverage across reducer event families, correlation, identity validation, projected question resolution, and pinned-policy denial. - Document the durable projection boundary. - Do not change dependencies, lockfiles, workflows, runnerd selection, server behavior, UI, or migrations. ## Verification - Replay base: `80639f4f69c8938eb74bdc0833df93e0ed91dab3` (`master` after #12422 merged). - Exact replay head: `3cb29581d2bcbc4b47f8069baffd721c6ce4e444`. - Stable patch ID: `92910b56575e67ae83960177d467a565019ba282`. - The exact delta is 22 files, 1,206 additions, and 59 deletions, all in `packages/paperclip-runner`; it contains no lockfile, workflow, server, UI, dependency, or migration change. - `git diff --check` and the Cargo formatting check pass on the replayed delta. - Exact-head GitHub Actions run `33372209037` (attempt 2): **PASSED** with 23/23 jobs passed. - Greptile reviewed exact head `3cb29581d2bcbc4b47f8069baffd721c6ce4e444`: **5/5**, with zero unresolved review threads. - Superagent, contributor trust, Socket, and Snyk security checks: **PASSED**. - No local test result is claimed. GitHub Actions is the authoritative verification environment for this replayed revision. ## Risks - This function accepts reducer output, not raw sidecar frames. Callers must preserve the existing scope-first decode and reduction order. - Semantic input includes the already sanitized provider input while its content receipt uses the same canonical digest. - Structured input preserves the validated provider-neutral question set and sanitized origin. - Noncanonical provider request identities are deterministically projected for PRP while the original identity remains authoritative for the sidecar resolution command. - Existing PRP v1 identifiers remain schema-compatible; the only public ID-schema change widens turn/item limits from 160 to 240 characters. The internal ACPX sidecar wire schema now mirrors the stable IDs its Rust transport already enforced. - The projector verifies event-carried terminal and assistant turn identifiers against the durable context. - No production path invokes this projector in this pull request. Durable command execution remains the next slice. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5.6, agentic reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the preceding public PR or described the issue in-PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
3aa2065d08 |
feat(runner): validate structured question responses (#12420)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Rust runner now owns a bounded ACPX session and a fail-closed turn lifecycle > - Provider questions pause a turn and must return structured answers to the same persisted question set > - JSON Schema validates the wire shape, but it cannot validate identifiers and constraints across two documents > - Unknown questions, invalid choices, and malformed custom answers must fail before any provider receives them > - This pull request adds only the package-local response validator and tests > - The benefit is a small trust boundary that later request-resolution code can use without changing production selection ## Linked Issues or Issue Description Refs #12419 ## What Changed - Validate `paperclip.question_response.v1` against its versioned JSON Schema. - Bound serialized responses to 768 KiB before validation. - Require answer identifiers to match the exact persisted question set. - Require answers for required questions and reject unknown question identifiers. - Enforce text, single-select, and multi-select answer modes. - Match the existing TypeScript numeric syntax, including decimal, exponent, hexadecimal, octal, and binary input. - Match ECMAScript trimming exactly, including BOM whitespace while rejecting Unicode NEL rather than inheriting Rust-specific whitespace behavior. - Enforce known options, custom-answer policy, text length, pattern, and numeric constraints. - Validate duplicate option IDs, inverted bounds, and dynamic patterns before answer lookup so malformed optional questions fail closed even when unanswered. - Match JavaScript UTF-16 code-unit length semantics for text constraints and the 100,000-unit response-field bound. - Preserve the public optional `recommended` question-option field in the versioned schema, generated schema bundle, and Rust validation path. - Return typed validation errors for malformed inputs without panics. - Export the validator from the Rust runner core. - Add table-driven tests for valid, mismatched, malformed, oversized, and numeric-boundary responses. - Document the package-local structured-response boundary. - Add `num-bigint` 0.4 and `num-traits` 0.2 as direct runner-core dependencies for exact arbitrary-length radix parsing and one-step JavaScript Number rounding; update only the package-local runner Cargo lockfile. - Do not change the repository PNPM lockfile, workflows, runnerd selection, server behavior, UI, or migrations. ## Verification - Replay base: `9a9fdf06ee4142f77427db30efccc4c43056f64b` (`master` after #12419 merged). - Exact replay head: `fad92b3fb348b66ddb10dde44b7b060e55c4fe96`. - Stable patch ID: `4d6ffbd519dd081f7ea530977cd965bd4569fc75`; this is the prepared two-commit delta plus the focused cross-language parity fix found during replay review. - The exact delta is 10 files, 712 additions, and 2 deletions, all in `packages/paperclip-runner`. - The package-local `packages/paperclip-runner/runner/Cargo.lock` records the two direct runner-core dependencies; their already-resolved versions and checksums are unchanged. - The question-set schema source, generated TypeScript schema bundle, and protocol manifest hash are updated together; the schema SHA-256 is `42b5441a3d388851dacb6e4500dfd4a17d878eded2e724228078b647e7440d3f`. - GitHub Actions run `33366812025`, attempt 2: **PASSED** on the exact replay head (23/23 jobs passed; a failed-job-only retry cleared one unrelated ACPX runtime-host timeout). - Greptile: **5/5** on the exact replay head with zero unresolved review threads; Superagent, Socket, Snyk, and contributor-trust checks also passed. - No local test result is claimed. GitHub Actions is the authoritative verification environment for this replayed revision. ## Risks - The validator compiles the embedded response schema for each submission. Responses are user-paced and bounded, so this keeps the slice simple without affecting a hot event path. - The persisted question set is the source of truth for identifiers and constraints. A malformed persisted set fails closed. - Numeric input follows the existing structured-question contract, including JavaScript-prefixed syntax. Optional whitespace-only answers are rejected instead of being treated as an omitted value. - Error messages identify the invalid field but do not include answer text. - No production path invokes this validator in this pull request. Request resolution remains the next slice. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5.6, agentic reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
e18632ebcb |
Add durable semantic tool receipts (#12353)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Semantic tools cross a trust boundary between a provider and the control plane > - Durable runs need exact input, result, denial, duplicate, and reconciliation receipts > - Replay must reject unsupported required versions and mismatched receipt pairs > - This pull request adds the receipt builders and deterministic replay fixtures > - It keeps newer sequence and gap safety limits from the current stack > - The benefit is auditable semantic activity before more providers use it ## Linked Issues or Issue Description **Subsystem affected** packages/paperclip-runner **Problem or motivation** Semantic tool calls have basic authorization records, but durable replay does not yet cover reconciled calls, denial redaction, duplicate receipts, governance targets, or artifact references. **Proposed solution** Add bounded semantic receipt builders, a reconciled phase, strict pair binding, fail-closed version checks, and generated replay oracles for the important lifecycle cases. **Alternatives considered** The runner could store provider-native tool payloads. That would weaken protocol portability and make redaction and retry behavior provider-specific. **Roadmap alignment** This supports the existing experimental Paperclip Runner rollout. It does not enable a production adapter. ## What Changed - Add semantic input and result receipt builders. - Add optional reconciliation receipts for pending calls. - Reject unsupported semantic receipt versions. - Validate receipt correlation, operation, idempotency, and digest bindings. - Add deterministic replay fixtures and generated golden outputs. ## Verification - `pnpm --filter @paperclipai/paperclip-runner test:typescript` - `pnpm --filter @paperclipai/paperclip-runner typecheck:typescript` - `pnpm -r typecheck` - `pnpm build` - Replay golden and protocol manifest checks pass. - The branch changes 27 files relative to its declared base. ## Risks The main risk is accepting a receipt that belongs to another call or replaying a duplicate as a new mutation. Binding checks compare correlation, operation, idempotency, and content digest fields. Fixtures cover denials, duplicates, governance chains, optional fields, artifacts, and unsupported versions. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5`, with agentic reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing public issue or described the issue in-PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b1cd261212 |
feat(runner): normalize provider event contracts (#12350)
## Thinking Path > - The runner already persists PRP events, but provider-native activity needs one bounded, provider-neutral vocabulary before additional providers can be added safely. > - The protocol catalog must describe capabilities without enabling or authorizing a provider. > - Provider normalization must not require an ACPX runtime dependency merely to compile the shared event layer. > - This pull request adds the event contract and pure normalizers only; provider transports and production selection remain unchanged. ## Linked Issues or Issue Description This is the first follow-up stacked on #12321. Codex, OpenCode, and ACP runtimes expose different activity shapes. Without canonical normalization, downstream task threads and traces would need provider-specific branching and could retain unbounded or unsafe payloads. ## What Changed - Expand the PRP provider descriptor and canonical activity event families. - Add bounded Codex, OpenCode, and ACP event normalizers for plans, tools, research, delegation, artifacts, review, safety, waits, and notices. - Preserve strict schema validation and regenerate the checked-in schema bundle and manifest. - Use a structural ACP event input so the provider-neutral layer does not introduce or authorize an ACPX runtime dependency. - Export the provider-event contract from the existing package root. ## Verification - `pnpm --filter @paperclipai/paperclip-runner typecheck:typescript` - `pnpm --filter @paperclipai/paperclip-runner test:typescript` — 11 files and 88 tests passed. - `pnpm -r typecheck` - `pnpm build` - `git diff --check` - The local full repository runner reached unrelated macOS workspace-path fixture failures; the affected runner suites pass and the repository CI shards are the handoff authority. - Diff against the declared base: 8 files. ## Compatibility Boundary - No provider transport, adapter, server route, feature flag, or runtime selection changes. - Catalog presence does not authorize discovery or execution. - Existing Codex execution continues through its current path. - No dependency, migration, workflow, or lockfile change. ## Risks The main risk is accepting malformed or unbounded provider payloads. Schema validation remains fail-closed, text/output fields are bounded and redacted, unsafe paths and URLs are discarded, and representative variants for every declared event family are covered by tests. ## Model Used OpenAI Codex, GPT-5 family. The client does not expose the exact deployment ID or context window. Agentic reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal task identifier - [x] I have run the affected local tests and they pass - [x] I have added or updated tests where applicable - [x] I have updated the compatibility notes for this change - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open actionable comments - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bc9ba7cd26 |
feat(runner): project native runs into task threads (#12321)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The experimental Paperclip Runner can execute a guarded Codex run and persist provider-neutral events. > - The task page still reads direct-adapter transcripts and cannot present those native events. > - Structured runner questions must also use the existing task interaction experience. > - Runtime selection must use the persisted run mode, not an adapter name or a current feature flag. > - This pull request projects native events and questions into the existing task thread. > - Direct adapters keep their existing transcript, composer, interaction, and finalization paths. > - The benefit is a complete native Codex task thread without a behavior change for existing adapters. ## Linked Issues or Issue Description Refs #12202. This pull request replaces that stale implementation on current `master`. **What happened?** The server persists native runner events and structured input requests. The task page only consumes direct-adapter transcripts. A native run therefore cannot present a complete transcript, usage, or question flow through the normal task experience. **Expected behavior** Native runs project persisted provider-neutral events into the existing task thread. Native structured questions use the existing interaction card. Direct adapters retain their current behavior. **Steps to reproduce** 1. Enable the experimental runner. 2. Start a native Codex run that emits progress, usage, a structured question, and a final reply. 3. Open the task page. 4. Observe that the direct-adapter transcript path cannot project the native event records. **Paperclip version or commit** `master` at `67f9867bc`. ## What Changed - Add the canonical structured-question validator and shared contract exports. - Materialize native input requests as existing task interactions. - Validate native answers and deliver them through the durable question-response receipt. - Resume the original PRP request with an idempotent `request.resolve` command. - Project native messages, tool activity, cumulative usage, and final replies into the existing transcript model. - Propagate persisted `runtimeMode` to the task page and select native handling only for `runtimeMode: "native"`. - Expire pending interactions through the shared issue service on every terminal transition, including decisions, stalled reviews, tree control, and pipeline retry cleanup. - Queue native run cancellation while a transaction is open and execute it only after the owning transaction commits. - Keep nonterminal and non-runner issue paths on their existing service call shapes and behavior. ## Verification - `pnpm --filter @paperclipai/server typecheck` — passed, including the Rust runner release build and protocol/catalog drift gates. - Focused native-thread and lifecycle suites — 18 files and 481 tests passed during review. - `issue-execution-policy-routes.test.ts` — 19/19 passed after the final transactional-queue expectation update. - `issue-agent-mutation-ownership-routes.test.ts` — 87/87 passed in the final isolated compatibility rerun. - GitHub Actions — policy, build, canary, typecheck/release registry, 5 serialized server shards, 8 general-test shards, 3 browser shards, and both aggregate gates passed on `7793f3193`. - Security — Snyk, Socket Project Report, Socket PR Alerts, and Superagent passed. - Greptile — 5/5 on `7793f3193`; all actionable review threads resolved. - `git diff --check` — passed. - Diff against `master`: 44 files. ## Compatibility Boundary - Native transcript polling only runs when the persisted run reports `runtimeMode: "native"`. - Missing or legacy runtime modes continue through `useLiveRunTranscripts`. - Legacy questions keep the existing optional free-text choice. - Native closed select sets can suppress that legacy fallback. - Terminal cleanup uses the same issue service for native and legacy interactions; only a bound native question schedules a native run cancellation. - Native cancellation happens after transaction commit, so failed or rolled-back writes do not cancel a still-valid run. - The durable delivery service checks the original native request before it considers a continuation run. - This pull request adds no migration, dependency, workflow, manifest, or lockfile change. ## Risks The main risk is routing a direct-adapter task through native handling or changing terminal issue behavior. The implementation selects the native path only from persisted runtime facts, retains the existing nonterminal call shape, and schedules native cancellation only for a validated bound native question after commit. Focused and repository-wide tests cover both paths. Native requests remain bound to the company, issue, run, and agent; answers are validated, durable, and idempotent across reconnects. ## Model Used OpenAI Codex, GPT-5 family. The client does not expose the exact deployment ID or context window. Agentic reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected local tests and they pass - [x] I have added or updated tests where applicable - [x] I have updated the compatibility notes for this change - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I addressed all Greptile and reviewer comments before requesting merge |
||
|
|
b2d1673b9e |
Add TypeScript PRP replay contracts (#12091)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs one typed interpretation of the language-neutral PRP contract. > - The JSON Schemas and fixtures now exist, but TypeScript consumers cannot validate or replay them yet. > - A deterministic reducer must define how duplicate delivery and source gaps affect the projected session. > - Result and question contracts must also validate untrusted provider and user input before later runtime code uses it. > - This pull request adds those TypeScript contracts and replay oracles without adding a process, provider, endpoint, or production behavior. > - The benefit is a reviewable and testable TypeScript foundation for the local runner and transport pull requests. ## Linked Issues or Issue Description **Subsystem affected** This change affects the private `@paperclipai/paperclip-runner` package. It does not change an existing server or adapter execution path. **Problem or motivation** The PRP v1 schemas do not yet provide TypeScript types, runtime validators, normalized result handling, or a deterministic session projection. Later Rust, transport, provider, and server work needs one tested TypeScript oracle instead of separate interpretations. **Proposed solution** Generate a checked-in TypeScript schema bundle from the PRP v1 sources. Add derived types, AJV validation, result and question validation, deterministic replay, a reducer, and generated golden snapshots. Export only these implemented root-package surfaces. **Alternatives considered** The combined runner branch adds the TypeScript contracts together with Rust, providers, semantic authorization, SDKs, labs, and server behavior. That delta is too large for normal review. Handwritten duplicate protocol types would also create a drift risk. **Roadmap alignment** This work supports the governed tool and control-plane direction in `ROADMAP.md`. It does not enable a new production adapter or endpoint. **Additional context** Refs #12087 and #11962. This pull request was prepared on #12087, then rebased onto its squash merge before opening. The current delta against `master` is 37 files. ## What Changed - Added JSON-Schema-derived PRP v1 types and AJV runtime validation. - Added fail-closed required-version checks and cross-envelope binding checks. - Added provider-neutral completion-result and structured-question contracts. - Added normalization for accepted legacy provider result aliases before strict validation. - Added a deterministic session reducer for replay, duplicate delivery, source gaps, requests, items, results, and terminal state. - Added generated replay snapshots and compact parity summaries for six accepted fixtures. - Added schema-bundle, manifest, and replay-golden drift gates. - Added only the root package export. Deferred testing, SDK, evaluation, lab, provider, and browser entry points remain unavailable. ## Verification - `pnpm --filter @paperclipai/paperclip-runner test` passed with 8 protocol tests and 44 TypeScript tests. - `pnpm --filter @paperclipai/paperclip-runner typecheck` passed. - `pnpm --filter @paperclipai/paperclip-runner check:replay-goldens` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm check:token-gates` passed. - `git diff --check` passed. - The delta against its declared base is 37 files. - `pnpm test:run` was executed locally. The package tests pass, while the macOS repository run retains the unchanged local-environment failures documented on #12087. The complete Linux CI matrix must pass on this commit. - A scoped scan found no secret-like values, internal references, or deferred-provider file names. - Greptile found an unbounded sequence-gap allocation. Commit `4a405c17` caps detailed missing IDs at 256, records the full missing count and truncation state, and rejects sequence values above the exact JavaScript integer range. The focused tests, workspace typecheck, build, and token gates pass after this fix. ## Risks Low production risk. The package remains private. This change adds no process, network endpoint, provider bridge, server integration, database change, or execution selection. The main risk is protocol interpretation drift. Generated schema and replay gates detect that drift. Browser and CSP-specific validator packaging remains deferred to its later package boundary. I checked `ROADMAP.md`. This change defines contracts for planned control-plane work and does not add overlapping product behavior. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fdbc69172d |
feat(runner): add PRP v1 schemas and fixtures (#12087)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs a language-neutral contract between the server and the runner process. > - A shared contract must exist before TypeScript, Rust, transport, or provider implementations can depend on it. > - Required protocol versions must fail closed, while safe optional fields must remain compatible. > - The contract also needs deterministic fixtures and a drift gate for later cross-language work. > - This pull request adds that contract without adding runtime behavior. > - The benefit is a small, reviewable source of truth for the next implementation pull requests. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This pull request adds a private package contract for later server, TypeScript, and Rust work. **Problem or motivation** Paperclip Runner does not have a small language-neutral protocol boundary on `master`. A runtime implementation without this boundary can drift between languages, accept unsupported required versions, or silently change canonical fixtures. **Proposed solution** Add PRP v1 JSON Schemas, accepted and rejected fixtures, a Codex structured-question fixture, and a generated SHA-256 manifest. Run compatibility and manifest checks during the package build. Keep the package private and export nothing in this pull request. **Alternatives considered** The combined runner branch contains schemas together with providers, SDKs, labs, and server behavior. That change is too large for normal review. Generating TypeScript validators in this pull request would also cross into the next review unit. **Roadmap alignment** This contract supports the governed tool and control-plane direction in `ROADMAP.md`. It does not enable a new production adapter or endpoint. **Additional context** Refs #12084 and #11962. This pull request was reviewed as a stack on #12084, then rebased and retargeted to `master` after #12084 merged. The current delta is 38 files. ## What Changed - Added 20 PRP v1 JSON Schemas with stable identifiers and resolved references, including explicit cross-language conformance input and output schemas. - Added canonical replay, cross-language, and Codex question fixtures. - Added accepted cases for additive optional fields and a rejected case for an unsupported required protocol version. - Added a deterministic manifest with SHA-256 digests for every schema and fixture. - Added package-local schema-instance, schema-reference, compatibility, question-ID, conformance-pair, and drift checks. - Added a private workspace package with no public exports and no production runtime behavior. - Added the package manifest to the Docker dependency-stage inventory required for every workspace package. This does not copy or build runner runtime code into the production image. - Kept the provider descriptor and question fixture Codex-only. No deferred provider package or dependency is present. ## Verification - `pnpm install --frozen-lockfile` passed with Node 24.19.0 and pnpm 9.15.4. No lockfile change is committed. - `pnpm --filter @paperclipai/paperclip-runner check:protocol` passed with 8 tests. - The committed AJV 2020-12 gate accepted every canonical v1 replay, question, and cross-language conformance fixture. It rejected the required v2 fixture, a replay fixture with a missing required command ID, and conformance output with a missing session ID. - `pnpm -r typecheck` passed. - `pnpm build` passed and ran the protocol manifest drift check. - `pnpm check:token-gates` passed. - `node ./scripts/check-docker-deps-stage.mjs` passed. - `git diff --check` passed. - The delta against its declared base is 38 files. - `pnpm test:run` completed with 4,687 passing tests, 19 skipped tests, and 29 failures across 9 unchanged server files. The failures reproduce macOS path aliases, local listener probes, workspace-runtime assumptions, and one connection-retry timeout. No changed-file test failed. Linux CI must pass before this pull request is ready. - `pnpm check:tokens` reports existing personal-name references outside this pull request. A scoped scan of `packages/paperclip-runner` found no secret-like values, internal references, or deferred-provider names. - PR #12084 was squash-merged, and this branch was rebased onto that merge and retargeted to `master`. The first master-base policy run correctly caught the missing Docker dependency-stage manifest copy; commit `4fa1ea7c` fixes that gate, and the complete Linux matrix is green. - Serialized server shard 1 initially hit an unchanged heartbeat test-harness timeout and a later assertion in the same file. Its isolated rerun passed in 3m57s. All other shards passed on their first attempt. - Greptile reviewed the final commit at 5/5 with no blocking failure. Both earlier actionable validation threads are resolved, and no review thread remains open. ## Risks Low production risk. The package is private and has no exports, server adapter, endpoint, or process. AJV is a package-only development dependency that the server workspace already uses. The main risk is contract churn before the TypeScript and Rust consumers land. The generated manifest and compatibility fixtures make that churn explicit. I checked `ROADMAP.md`. This change defines a contract for planned control-plane work and does not add overlapping product behavior. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |