mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 05:31:46 +02:00
codex/hermes-qualification
82
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ef0d855034 |
test(hermes): enforce native image evidence gate
Require saved image checks to pass after account and billing evidence is collected. Advance the explicit image definition to version 2. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9478ead982 |
test(hermes): add native image acceptance journey
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2400ecdb8d |
test(hermes): require the persisted native cancellation receipt
Bind the expired card to its original unanswered request and require the canonical zero-answer cancellation result. Preserve the failed live measurement and advance qualification definitions. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a14a752ec8 |
fix(hermes): preserve cancelled usage at the ACPX wire boundary
Bind usage to the admitted native session and turn before cancelled prompt settlement. Check its durable PRP carrier and harden qualification receipt identity and budget capture. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
36125adb4e |
test(hermes): pin lower qualification campaign budgets
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0bda651f21 |
Qualify browser Stop at an unanswered Hermes callback
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b0bbaffec6 |
Clarify the Hermes native question fixture objective
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3c7e0ef3a1 |
Test Hermes native question batches through the browser
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e3be06f25e |
fix(hermes): retain usage after input-yield shutdown
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e5afbd7107 |
fix(hermes): retain per-turn reported OpenRouter billing
Capture pinned SDK wire attempts, carry closed versioned receipts through ACPX and PRP, and bind reported subtotals to native turn accounting. Keep incomplete attempts unpriced and require settled cost plus budget health in the live OpenRouter oracle. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
01855199ad |
fix(hermes): verify approved qualification lockfile
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6d81d5c0dd |
fix(hermes): record checked-out qualification source before credentials
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
af50ff553d |
test(hermes): add bounded managed Bedrock qualification
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
08d9803b8f |
test(hermes): verify the expected account owner independently
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
4313a55dde |
test(hermes): bound API qualification before task dispatch
Require one attempt and public readback of scoped 200-cent company and agent budgets before the paid API task is created. Reject unlimited or foreign budget records. Preserve unknown usage as unknown. Record all five successful paid Linux browser workflows and verified temporary AWS host/access cleanup. The complete support suite passes 1,833 Vitest tests and 128 Node assertions, with one skipped Vitest test. Product E2E TypeScript and diff checks pass. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b10746b4e6 |
test(hermes): qualify managed API accounts through the runner
Add ten explicit pending API-account browser cells and independently grade the selected account, native harness, and exact model. Keep fixture credential staging on the shared Connections path. Carry candidate metadata into CI so every selected Hermes profile receives verified runtime assets before secrets. Record the private AWS image build and three actual Linux browser passes without promoting Hermes. Product E2E support typecheck and all 102 affected tests pass; the broader support suite passed 1,826 Vitest tests and 128 node assertions before the final oracle refinements. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5c2a78fa45 |
ci(hermes): provision pinned runtime for selected paid campaigns
Select candidate assets consistently for immutable image identity and remote packs; provision local Linux assets before the protected paid step exposes credentials. Preserve default-branch dispatch and qualification gates. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
4a0b327727 |
test(hermes): add paid browser qualification campaigns
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
eda7dd0ab7 |
docs(runner): pass resolved lock digest in image commands
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
68b1a39459 |
fix(hermes): keep native candidate independently restorable
Reject replaced learned-state parents, include authenticated run attachment and the advertised native transport fixtures in this candidate. Keep paid browser campaign definitions in the qualification PR. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9e07748e91 |
feat(runner): add native Hermes ACP qualification candidate
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d0db8820db |
fix(connections): repair native baseline and approval continuations (#15420)
## Thinking Path > - Paperclip manages AI agents and the tools they may use. > - Connection setup separates provider preference from permission to use a tool. > - The first native connection baseline could not exercise its intended decisions. > - The browser used mutable task titles, and the provider fixture already granted access. > - Native provider-choice instructions also disagreed with the preferred question format. Schema rejection gave no field guidance. > - This pull request repairs those test preconditions and native guidance, then fixes restart/approval defects exposed by the corrected baseline. It also restores missing OpenCode tool-error evidence. > - The benefit is an inspectable baseline before any further instruction reduction. ## Linked Issues or Issue Description Refs #15407. The original 15-cell baseline remains 0 PASS / 15 FAIL. Ten cells stopped on stale titles, two Codex cells had schema denials, two OpenCode cells used already-granted tools, and one Claude cell returned no native result. No intended user decisions were submitted. The exact invalid Codex field and underlying Claude failure cause remain unknown. [Original campaign](https://github.com/paperclipai/paperclip/actions/runs/37562577199) · [Original report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37562577199-1/index.html) ## What Changed - Match the browser's task route and visible identifier instead of a title the agent can change. - Start native provider-choice fixtures with no agent tool access. Verify the public effective-access records. - In the positive case, select Arcade, then grant its exact HubSpot tool through the real access card. Require both saved decisions and exactly one observed call. - Return a canonical `providerQuestionSet` for native input and retain the equivalent legacy `providerQuestion`. - Keep invalid input rejected. Return bounded schema locations and required field names without submitted values. - Preserve Claude's exact session/content identity while allowing authenticated registered instruction-copy paths to rotate on a new run. - Reject duplicate approval reports for an existing exact tool-action card before they create another human review. - Wait for a recorded service-approval continuation within the existing deadline; retain missing or failed continuation grades. - Forward OpenCode tool activity through the runner facade, preserving bounded errors and execution-part identity without inventing host-call joins or exposing arguments. - Preserve original grades, costs, scope limits and diagnoses in the dated repair report. ## Verification - Eval typecheck passes. Support suite: 1,801 PASS, one intentional skip; Node checks: 128 PASS. - Connection/schema tests: 51 PASS. Real-server public fixture setup: one PASS with zero providers. - Browser support regression: five PASS, including renamed and wrong tasks. - Focused Rust safe-feedback test: one PASS. - Repository typecheck and build pass before the latest master replay. Post-replay connection/shared/real-server fixture checks: 52 PASS; eval typecheck passes. The browser review fix additionally passes all five browser checks and seven suite checks. - The full local repository run was interrupted incomplete after about 45 minutes, with five integration failures retained. All five pass in a separate targeted invocation (1,250 unrelated tests skipped). No full local-suite pass or root cause for the initial local failures is claimed. - Corrected frozen source `162cc90fdabe7f505b88ae095044531b82784c92`: **10 PASS / 5 FAIL** across the [passing Codex canary](https://github.com/paperclipai/paperclip/actions/runs/37575158761) and [remaining 14 cells](https://github.com/paperclipai/paperclip/actions/runs/37576261807). The canary passes all 17 checks. Claude's two provider-choice continuations fail on restart, Claude service approval exposes an early evaluator rejection, Codex service approval creates a duplicate approval, and OpenCode provider-second times out after both decisions with no HubSpot call. No original result is regraded. - Final ledgers count 31 actual runs: 27 succeeded, two failed, two cancelled during cleanup. All 15 cleanup/budget checks pass. The late Claude continuation is absent from its earlier workflow snapshot; it remains in the result/API/final ledger. Original evidence retains 279 hashes. Recorded LLM subtotal $0.04553787 is incomplete billing, not actual total cost; local runtime is unmetered. - New repair regressions reproduce the Claude attach failure, duplicate approval acceptance and dropped OpenCode tool events before their respective fixes. Nine Rust attachment checks, 127 ACPX host/adapter tests, 33 completion/control-plane checks, nine eval deadline tests, 59 OpenCode proxy/driver tests, one Rust tool-error/redaction check, and TypeScript/Rust composer parity pass. Eval typecheck, repository typecheck and build pass. Existing support coverage is 1,802 PASS plus 128 Node PASS, one intentional support skip; two additional deadline tests also pass. - New-source full CI/review and live canaries are pending. The next bounded selection is Claude provider-decline, Codex service-approve and one OpenCode provider-second diagnostic with repaired event evidence. No broader campaign or instruction-reduction qualification is claimed. - Initial corrected campaign [37574251834](https://github.com/paperclipai/paperclip/actions/runs/37574251834) was cancelled during shared build after review found the breadcrumb whitespace assumption. Its matrix job has zero steps and no provider execution. The real adjacent-span browser regression now reproduces the old failure and passes after the fix. ## Risks - The corrected baseline remains 10/15. The new restart/approval fixes require live qualification; OpenCode evidence forwarding does not itself establish or fix its prior behavioral failure. - The positive provider case now expects three runs, including separate access approval. Its new results are distinct from the original invalid fixture. - The old Claude missing-result cause and rejected Codex field are unknown. These repairs do not retroactively explain or erase either failure. - Path rotation must preserve prompt, custom instruction, skill/content identity and protected provider settings; regression checks reject stale or changed content. No connection authorization, JSON schema, budget, cleanup, or final-result requirement is relaxed. Historical Everyday prompts and gateway setup remain unchanged. ## Model Used OpenAI Codex, GPT-6. The exact deployment variant and context window are not exposed in this session. Used repository inspection, code editing, test execution and retained-evidence analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
caf120105c |
test: prepare neutral native connection guidance evals (#15407)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need to discover connections, obtain consent, and continue from saved decisions. > - We want to reduce repeated instructions only when measured behavior supports the change. > - The existing decline tasks tell the model not to retry. One provider-decline check can pass without an explanation or an observed service counter. > - This PR adds neutral tasks and stricter saved-evidence checks before any connection instruction reduction. > - Production instructions remain unchanged. The new cells are configured, not live-qualified. ## Linked Issues or Issue Description Refs #15218. Refs #15389. **What existing behavior does this improve?** The Product E2E connection workflow evaluation and its instruction measurement provenance. **Current behavior** Some decline prompts supply the policy they intend to test. The provider-decline workflow does not require a saved post-decision explanation. Its old no-call check can use a missing fixture counter as zero. OpenCode has no connection cases in the original Everyday matrix. **Proposed behavior** Add an explicit-only suite with five connection stories on native Codex, ACPX Claude, and OpenCode. Require an explanation attributed by exact run ID after a saved decline. Observe the provider fixture counter. Preserve the original cases and grades. ## What Changed - Add fifteen configured cells with one attempt, twelve-minute deadlines, and verified 1,000-cent company and agent budget stops. - Remove procedure hints from the three new decline prompts. Keep a user-permitted explanation fallback and the existing positive controls. - Require saved decline state, one decision, unchanged connections, observed zero service calls where applicable, and a post-decision explanation from a successful run on the same task. - Add negative grader calibration and test the actual fixture budget payloads. Exclude the suite from default and generic selection. - Extend the existing full-catalog measurement source manifest with connection descriptions and schemas. Add an audit of fixed text, tool descriptions, returned instructions, and unqualified behavior. - Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and preserve its new Cursor suites. No production, credential, workflow, or lockfile change. ## Verification - Before rebase: Product E2E support passed 1,424 TypeScript tests and 128 Node checks. Six catalog measurement tests, repository typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery passed. - The full pre-rebase repository test run was stopped when master advanced. Its partial result is not a pass. - After rebase and the review correction: repository build/typecheck, Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128 Node checks, six measurement tests, and exact fifteen-cell discovery pass. The duplicate local full-suite run was stopped incomplete after about 20 minutes once complete CI passed; no local full-suite pass is claimed. - Review found that the initial grader read `runId` instead of public `createdByRunId`. A regression calibration reproduced both rejection of valid public comments and acceptance of the wrong alias. The fix uses the actual field and binds the evidence type to the shared `IssueComment` contract. A subsequent type-only import path correction passes Product E2E typecheck. - Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes [complete CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545): 51 successful checks and two intentional Storybook skips, plus separate Snyk success. Fresh Greptile review is 5/5 with the single review thread resolved and no new findings. The PR is clean and mergeable. - Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm test:e2e:runner -- --list --suite native-connection-guidance`. The measurement uses `PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json pnpm exec vitest run --project @paperclipai/server server/src/__tests__/native-procedure-measurement.test.ts`. Validation used pinned pnpm 9.15.4. - No paid provider campaign was started. There is no baseline/candidate behavior result for these new cells. - The audit records 654 UTF-8 bytes of fixed connection guidance. A clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41 supplied tools, 53,341 normalized bytes at start/resume, 50,949 at compact continuation, and a 48,195-byte authenticated OpenCode MCP catalog. These are byte counts, not tokens, bills, vendor-private prompt sizes, or savings from this PR. ## Risks - This is eval preparation. Passing support tests do not establish live model behavior or qualify an instruction reduction. - The explanation oracle checks attributed saved output. It does not prove cognition or arbitrary prose truthfulness. One saved interaction also does not prove the absence of repeated idempotent tool calls. - Successful new authentication and tool refresh, existing-connection agent grants, independent work while waiting, explicit retry after decline, and blocking when mandatory work remains still need separate coverage. - Notion setup decline does not execute a real Notion service. Positive service approval uses an already installed deterministic service; it does not qualify new connection creation. - The original historical failures remain unchanged. Future comparisons must freeze source, fixture, model, input, and grading controls and retain every actual attempt. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code execution. The exact serving model ID and context-window size were not exposed in this session; they are not inferred. No model provider was invoked by the eval suite in this PR preparation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a6306ba606 |
feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e38d6d16b6 |
feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect accounts and choose an agent harness and model. > - The runtime change in #14970 supports custom providers on those connections. > - Normal setup must stay simple while advanced users can choose a compatible gateway. > - Shared connector rows and access controls keep these choices consistent. > - This pull request refines the agent setup UI and adds review stories and repeatable browser qualification. > - The qualification checks real tools and downloaded outputs, not only a successful run status. ## Linked Issues or Issue Description Refs #14970, #37, #13083, #14104, #14565, #12692. The core implementation in #14970 is merged. This branch incorporates its squash commit and targets `master`. Both PRs contain our implementation. #14016 is a reference only and is not a dependency. This PR has 96 changed files. ## What Changed - Complete model-provider connector presentation beside other connectors. Each row uses the existing Connect action and connection list. Tags are stored without category UI. The base PR includes the provider forms and routes. - Show persistent Subscription, API Key, and Advanced choices. Label Advanced as Custom Gateway. Reuse provider logos, connection lists, and permissions controls. Default access to the organization and all agents when permitted; keep narrowing controls under Advanced. - Keep Configure reachable before subscription sign-in, so users can select a supported environment when the default cannot sign in. Testing and saving still require a connection. Show the execution environment in Configure. Preserve the confirmed Connect choice. Editing a method, credential, saved account, or advanced choice requires that current choice to connect before testing or saving. Use matching model and thinking-effort dropdowns and retain connection icons in selected values. - Preserve the new harness model default when switching an existing OpenCode agent to Codex or Claude, and resolve user-selected model names with the effective harness. - Load popular OpenRouter models through the shared connection-model discovery path. Keep explicit model lists and manual model entry available. - Group onboarding, connection setup, agent runtime, management, recovery, and production-component stories under AI Connections / Provider routing. - Add an explicit-only provider-connections browser suite for managed local or existing local/staging targets. Use private browser profiles and credential handoffs. Support human-assisted subscription sign-in without sharing passwords or tokens in reports. - Verify persisted connection identity, runtime probes, tool execution, exact artifact bytes, completion, and context-dependent follow-up. Retain source/model provenance, cost bounds, closed error diagnostics, original failures, and cleanup evidence. - Add Gemini startup-model and skill-root fixes, Grok private-history detection, ACP filesystem regression fixtures, selected-workspace handling for local Hermes, and artifact-helper workspace fallback. - Keep managed Grok runtime homes disposable. Remove host-side transcript retention/restoration because private file modes do not isolate same-user agent processes. Ignore earlier development archives and use a fresh task handoff when history is unavailable. Verify the absence of restored transcripts with a separate same-user process. - Capture stopped-run diagnostics before deleting an attached-company fixture agent. Track creation and owned sign-in receipts; revoke only this attempt's accounts and never adopt a concurrent campaign's newly created account. Preserve failure signals and final status through cleanup. - Require the requested environment in the saved agent and every run, including follow-ups. Reject a forced incompatible target. Keep one cancellation state through startup, every cell, reporting, and teardown for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after interruption. Document qualification limits. ## Verification - Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes master `d9f600043`. The security fix in `a758fde31` passes full workspace typecheck, production build, and 119 connection/Grok regressions. The unchanged UI passes all 126 configuration/model-discovery tests and token gates. The final published-guide correction passes Grok adapter typecheck. Earlier head `eebd8225c` passed the complete deterministic runner suite (1,404 Vitest tests and 128 Node tests) and all CI jobs. Current-head CI run `37520147514` passed all 47 jobs, including the full sharded Vitest and browser matrix, production build, and canary dry run. All 55 checks completed: 53 successes and two expected skips. The current-head security scan passed, Greptile is 5/5, and no review threads remain open. - A separate same-user process reproduced reading a restored Grok transcript before the security fix. The regression now finds no transcript. Existing fresh-session fallback and ordinary session metadata behavior pass. - The final account-choice and cleanup fixes pass 85 setup tests and 26 qualification-harness tests. Regressions verify that editing a connection invalidates confirmation, Configure remains reachable before sign-in, diagnostics are captured before fixture deletion, and concurrent campaigns cannot adopt or revoke each other's accounts. UI and E2E typechecks pass. - The Storybook build and actual Chromium production-component stories passed during this change. Review the neighboring AI Connections / Provider routing stories, regular connector rows, three connection modes, model discovery, and the single execution-environment control in Configure. - Cancellation smoke verified authenticated cleanup before browser close for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during startup and reporting, missing-file ACP resource errors, and preserved permission denials. Both ACP runtime versions and 54 ACPX/Grok regressions passed. The deterministic connection-intent browser suite passed two tests. - Historical local qualification retained 43 passing API/gateway cells out of 46, with downloaded outputs and follow-up receipts. These attempts span earlier builds; they do not qualify this exact commit or staging. Subscription combinations, Gemini overloads, and the unresolved follow-up failure remain recorded rather than counted as passing. - Use `pnpm test:e2e:runner -- --list --suite provider-connections` to inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md` for credentials, target URL, sign-in assistance, budget, evidence, and cleanup. Paid live tests remain opt-in. ## Risks - The core implementation in #14970 is merged. This PR adds no database migration of its own. - Subscription login needs an interactive provider session. Dedicated accounts and staging qualification remain follow-up work; this PR does not certify every login combination for production. - Managed Grok transcript resume is deferred until provider history has an OS isolation or authorized broker solution. Follow-ups start fresh with Paperclip task context; earlier live Grok results do not qualify this behavior. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Live overloads and one unresolved follow-up timeout remain recorded. The stock CLI is unchanged, and those cases are not marked as passing. - Real-provider tests spend credits and use private credential/evidence directories. The launcher requires explicit selection and checks target ownership. It must not attach to a developer's database by accident. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local remain outside custom provider setup. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e34abee670 |
feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path > - Paperclip gives teams durable tasks, agent execution, budgets, and approvals. > - People also use assistants in Codex, Claude, and other MCP clients. > - Those assistants need a scoped connection that preserves the person’s permissions and attribution. > - Delegating a task must not turn the assistant into the assigned agent. > - This PR adds opt-in user OAuth, ten first-party tools, browser consent, and workflow packages. > - Paid product evals verify the resulting tasks, documents, attribution, retries, and access boundaries. > - The team keeps working after the assistant conversation ends. ## Linked Issues or Issue Description **Problem or motivation** A person cannot connect an external assistant to an existing team through browser consent and safely delegate durable work as themselves. **Proposed solution** Expose an opt-in `/mcp/paperclip` endpoint with individually described first-party operations. Bind every connection to a person, client, company, resource, and scopes. Reuse domain authorization and scheduling. Package shared team-review, delegation, and follow-up workflows for OpenAI/Codex and Claude. **Alternatives considered** Related PRs #9393 and #12549 cover earlier remote MCP and board-operator approaches. This change uses user OAuth and a bounded public catalog. It does not expose a generic executor, operator administration, static shared board credentials, or external agent execution. Registry listing work in #9851 is a separate distribution step. **Roadmap alignment** This maintainer-requested implementation extends the governed MCP gateway, activity attribution, durable work products, and hosted deployment direction in `ROADMAP.md`. It implements the first release of the saved design plan; external agent participation and granted third-party tools remain later releases. ## What Changed - Add MCP 2.0 discovery and task status/comment/document Events on the same authenticated endpoint. Persist subscriptions and delivery receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt callback material, recheck permissions/Cloud membership, and bound retries/expiry. Older MCP clients keep their existing tools. - Add discovery, dynamic client registration, S256 PKCE, resource validation, rotating refresh tokens, revocation, and company consent. Store credentials as hashes and recheck membership at execution. - Add tools for connection identity, agents/projects, task search/read/create, human comments, documents/deliverables, and pending-approval links. Preserve current domain permissions and scheduling. - Add durable mutation receipts across reconnects. Matching retries replay results; uncertain outcomes keep the same request ID and require inspection. - Add consent and connection-management pages, OAuth log redaction, shared plugin workflows, and separate OpenAI/Codex and Claude package outputs. - Add eight paid Product E2E cases across three models, independent durable-state grading, usage evidence, cleanup, and report integration. Add task-document guidance and regenerate the runner capability inventories. - Add migrations 0301 and 0302, the dated implementation plan, result notes, and direct-client setup instructions in `doc/public-mcp.md`. ## Verification - Merge integration `e180b1948`: resolved conflicts with current master, preserved both eval registries, regenerated capability catalogs, and regenerated migrations as 0301/0302 while keeping the original replay-safe SQL byte-identical. Local migration safety/snapshot tests (26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval catalog/grading tests (198) pass. Token and capability gates pass. Full recursive typecheck passed. Fresh Greptile review is 5/5 with no unresolved findings. CI is green on this exact head (55 successes, two intentional skips, one neutral result): one unchanged Cursor sandbox test timed out at 10 seconds, then passed locally in 856 ms. A single retry of that failed shard and the aggregate workflow passed. Merge remains blocked on the repository code-owner approval rule. Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55 successes, two intentional skips and one neutral result. [The earlier CI run](https://github.com/paperclipai/paperclip/actions/runs/36901592350) includes all test shards, browser tests, typecheck, build and canary dry run. Greptile was 5/5 on that commit with no unresolved review threads. GitHub still requires code-owner review under the repository merge rules; passing checks do not bypass that approval. Paid source fingerprints remain separate below and in the dated result note. - Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku 4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback, signature verification and report retrieval in a fresh conversation. A final Mini regression passes after the quota/status fixes. All evidence validates. Bounded tunnel startup retries occur before provider calls and remain visible; failed earlier attempts retain their original grades. - The earlier complete seven-case matrix passes **21/21**, with a separate **3/3** delegation regression. Two preceding matrices also passed 21/21 each. A complete 24-cell matrix including Events has not been run. [The dated results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain exact source fingerprints, failures, model IDs and partial costs. - Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass after merging master. Server typecheck passes after the final quota/status changes. Eval typecheck and all 892 eval-support tests pass. - All 33 real MCP/OAuth tests pass. The preceding combined MCP, redaction, private-address and DNS-rebinding run passed 129 tests; two later MCP regressions cover quota reuse and unchanged-status suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI then found a null checkout result in the existing concurrent-workspace path; logging now uses optional status access. All 12 closed-workspace tests and all 33 MCP tests pass after that correction. The exact-start event calibration exposed a timestamp gap; scanning now includes the subscription start, with all 33 MCP tests and server typecheck passing. These two narrow corrections follow the paid regression. - A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0 subscription/delivery, current membership loss, unsubscribe, legacy SDK tools, refresh and revocation. Its callback transport is a fixture with independent HMAC verification. The paid Events campaigns separately prove public HTTPS delivery. - Earlier component qualification passed UI 7,117 tests, CLI 502, shared 832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module boundaries, migration order and plugin regeneration passed. CI covers general/serialized suites, eight browser shards, runner checks, typecheck, build and canary dry run. - **Local full-suite limitation:** the earlier monolithic run was not clean. It encountered overlapping schema rebuilding, Mac database shared-memory limits and isolated CLI/fixture failures. Targeted reruns passed. The existing >32 MiB Git filename stress test still hit its 300-second Mac timeout. The additional serialized sweep stopped after 62 passing suites once CI passed. Original failures and partial logs remain; this PR does not claim a wholly green local monolithic run. - Local Codex CLI and Claude Code OAuth login and MCP SDK interoperability were verified. Public-store installation, actual ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted newcomer provisioning remain release gates. Enablement is moving to **Settings → Experimental → Assistant connections (MCP)** in the stacked follow-up [#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge both for the intended setup experience. This foundation branch alone still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set `PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and connect to `/mcp/paperclip`. Select a team and allow writes in browser consent. Configure an available agent and budget, then delegate and retrieve results later. For Events, rescan the deployed plugin catalog in ChatGPT Work Cloud; the host supplies its webhook credentials when the user asks to watch a task. See [the setup runbook](doc/public-mcp.md). ## Risks - Events are at-least-once and may arrive out of order. No replay cursor is advertised. Clients must refresh finite subscriptions, read current state and avoid comment feedback loops. Callback material uses the instance secrets master key; hosted subscriptions require the updated Cloud broker and are bounded to five minutes/the access proof expiry. - ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging subscription remain deployment gates. Local signed-webhook and paid model evidence does not claim those surfaces have been exercised. - Disabled by default. Merging adds schema and opt-in code; it does not deploy a public endpoint, publish a store listing, create a team, or start paid agents. - Migrations 0301 and 0302 are additive and idempotent. Their SQL is unchanged from the earlier preview numbers, so hash-aware upgrade reconciliation preserves prior staging applications. Normal instance upgrades must apply it before enabling MCP. - Task creation and comments can schedule paid agent work. Consent and tool descriptions disclose that effect. Revocation blocks future calls but does not undo delegated work. - Public deployments need edge rate limits and credential-safe logging. Internal dispatch is restricted to the closed catalog and carries a request-local verified actor. - Hosted onboarding requires the companion Cloud broker, encryption-key configuration, and tenant rollout. Self-hosted direct connections can use this PR alone. - Store acceptance and agent-mode participation are not claimed. Checked-in plugin endpoints are development defaults; rebuild packages for a real deployment before installation. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A more specific serving version and context-window size were not exposed by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`, `claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted/component checks; full local-run limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
16b7db35ff |
Shorten planning skills and measure task decomposition (#15296)
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
72ff3a9f27 |
Measure native tool context and expand bounded workflow evals (#15218)
## Thinking Path > - Paperclip manages persistent agents and their assigned work. > - Agents receive both fixed instructions and tool definitions. > - Moving a procedure into a tool description still adds model context. > - We need to measure the complete delivered catalog and test real outcomes. > - The tested reductions saved little space and introduced failing outcomes. > - This PR keeps measurement, bounded eval coverage and the original evidence. > - Production prompts, tools and runtime behavior stay unchanged. ## Linked Issues or Issue Description Refs: #15151, #14961, #14948, #14985. This adds the measurement and eval coverage needed to assess further native instruction changes. The attempted hiring and dependency reduction failed qualification and is excluded from the final diff. ## What Changed - Measure the actual standard-mode tool authority, including all 39 tools and their input schemas. Capture scripted native start, resume and continuation payloads and the OpenCode MCP declaration list. - Add OpenCode to the two explicit-only local hiring/reuse and delegation/feedback stories. Preserve the original task requests and independent oracles. - Apply one attempt per selected story and explicit company and lead-agent budget stops. - Select managed hiring credentials from the requested profile, including OpenRouter. - Run the existing Node test files under Node instead of collecting them as Vitest suites. - Retain sanitized comparison reports, original failed grades, source hashes and evidence gaps. ## Verification Final source: `2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8`. Local verification passes: - Full repository `pnpm -r typecheck` and `pnpm build`. - Eval-support unit tests: 1,252 Vitest checks and 128 Node checks. - Eval TypeScript check and six final-source measurement tests. All 36 normalized components across nine scripted deliveries and the OpenCode MCP catalog match the baseline exactly; the complete standing projection is 49,200 bytes. Fresh review of this exact head is [5/5 with no remaining findings](https://github.com/paperclipai/paperclip/pull/15218#issuecomment-5996723170); all three review threads are resolved. [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/37354539126) passes all 47 jobs, including repository typecheck, build, tests, runner checks, browser shards and the canary check. One initial annotation-test timeout is retained in attempt 1; its seven-test suite passed locally, and the affected CI lane passed on one targeted retry. No source or paid eval rerun was needed. Both ready-transition security scans passed with zero annotations. The PR is out of draft and conflict-free. GitHub still requires code-owner approval for the `package.json` test-script change; its requested reviewers are already set. A byte-for-byte comparison against master context `a65ca0950834a85bb93bcc4b4042ecacdebfef53` confirms no production changes under `packages/` or production `server/` paths. The only server addition is a measurement test. No new paid rerun is needed to compare unchanged production bytes. This does not claim that existing product defects have been fixed. The rejected corrected experiment had baseline **5 PASS / 1 FAIL** and candidate **3 PASS / 3 FAIL**, including **two newly failing pairs**. The later readiness experiment had candidate **3 PASS / 3 FAIL** and baseline **3 PASS / 2 FAIL / one setup cell without a behavioral grade**. Its five comparable pairs had two new failures, two new passes and one unchanged pass. The missing baseline Codex cell never reached its provider step because Docker setup timed out. The observed missing behaviors include parent continuation, waiting for the latest child revision, revised ZIP delivery and OpenCode credential persistence. Those failures remain failures. Source review and passing CI do not regrade them. The original reduction saved 460 bytes; its first repair saved only 125 bytes, and the larger unqualified runtime repair increased the full projection. None of those production changes is shipped here. Read the [report](https://github.com/paperclipai/paperclip/blob/2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8/doc/plans/2026-10-05-native-procedure-guidance.md) and its linked sanitized receipts for exact sources, original campaign links, pair-level results and evidence limitations. ## Risks - Full-catalog bytes are not model tokens, invoices, private vendor prompts, lazy-loading behavior or truncation proof. - The two stories are explicit-only and do not prove general coding quality or arbitrary resume behavior. Single trials do not establish causation or performance trends. - The retained OpenCode candidate credential guard failed. Cleanup removed the original provider database, so the precise persistence mechanism remains unknown. The guard is unchanged. - ACPX provider-execution IDs lack proven host-call mapping. No extra-work or feedback-consumption claim is inferred by matching names, order or counts. - PostgreSQL cannot start locally while the host's shared-memory slots are exhausted. Hosted CI must supply the full database checks; the full local database suite is not claimed green. ## Model Used OpenAI Codex, based on GPT-6, with repository tools and code execution. The exact deployment model ID and context-window limit are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a386a59998 |
Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native agents receive task constraints and completion tools from Paperclip. > - Completion tools already define the procedure for reporting a result. > - Repeated procedure text adds instructions to each full task turn. > - The final reply must still explain a blocker and link a saved document. > - This pull request removes repeated procedure text and keeps these visible outcome requirements explicit. > - A document receipt supplies the exact link, and stricter evals check the persisted reply and browser navigation. ## Linked Issues or Issue Description Refs: #14961. Related: #14948 and #15007. **What happened?** Native task envelopes repeat completion procedure text. A reduced envelope needs explicit final-reply requirements. The `write_document` receipt also lacks a canonical document link. **Expected behavior** Keep the completion tools as the source of procedure details. Require one accepted completion result before the final reply. A blocked reply must explain the reason, owner and unblock action. A document reply must contain a working link to the saved document. **Steps to reproduce** 1. Run the native assigned-skill document case and native blocker case. 2. Inspect the run-attributed provider final and its persisted comment. 3. Check the blocker explanation or open the final reply's document link. ## What Changed - Remove repeated completion procedure text from the native task constraints and backend instructions. - Keep explicit blocker and document-link requirements in full task turns. - Return a company/task-scoped `documentHref` from `write_document`. Preserve the link in the idempotent mutation receipt. - Repeat canonical links for this run's current saved revisions in accepted completion feedback. Give blocked providers final-response guidance for the cause, owner and unblock action. - Keep internal document/comment anchors when Markdown issue links load cached issue details. - Add a manual six-cell comparison suite with strict source, build, default-instruction and budget admission. - Capture eighteen shared runnerd RPC projections and six direct OpenCode HTTP projections across start, resume and continuation phases, using scripted local transports and no provider execution. - Apply v3 checks only to the manual instruction comparison; preserve v2 checks for the existing native completion suite. Check the actual persisted blocker reason and exact saved-document link. Click the rendered document link and check the original content marker in the classic document card or the new document tab. - Forward exact OpenCode finishing calls through the controller. Wait for acceptance, keep accepted feedback and concrete rejection text, and reject malformed responses. Preserve ordinary dynamic-tool response handling. - Settle the completion decision and tool response before mapping a racing idle/error/abort event or handling explicit close/interruption. Reject a concurrent finishing call before controller admission. - Add a provider-free regression through real runnerd, the OpenCode proxy and a fake provider. Reject the first completion, accept the corrected report in the same turn, and propose one result. - Keep all original verdicts unchanged. Treat replay under new checks as separate diagnostics. ## Verification - `pnpm -r typecheck` and `pnpm build` pass locally. - Native document-authority tests pass, including company/run authorization and idempotent replay. - Native runtime-context, backend and measurement tests pass. - Final-answer calibration, protocol scoring, source-admission and catalog tests pass. Wrong reasons, absent links and wrong link targets fail. - `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six single-attempt local cells with the declared models. - Exported `prepareNativeInstructionPreflight` then `verifyNativeInstructionPreflight` pass on this clean committed source. They build locally and make zero provider calls. - Corrective live confirmation is incomplete. Source |
||
|
|
b17019e14d |
fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its adapters supply task context and access to Paperclip skills and tools. > - The default hire manual and shared prompts also repeat general work procedures. > - Those procedures overlap with stock provider instructions and the Paperclip skill. > - Existing E2E fixtures supply a QA manual, so they do not qualify the production default. > - This pull request reduces the generic instructions and adds real default-hire coverage. > - The benefit is less competing guidance, with inspectable evidence for preserved skills and task context. ## Linked Issues or Issue Description Refs: #14920. That merged change preserves native Codex base instructions. This PR covers the default manual, shared legacy prompts, operational skill guidance, and the narrowly approved ACP skill-discovery/session-environment repair for measured delivery and credential-persistence failures. **What existing behavior does this improve?** New non-CEO hires without a custom bundle and legacy task/chat startup and continuation prompts. **Current behavior** The shipped default manual contains 602 words. Generic task/chat prompts and ordinary resume deltas repeat work procedures already available through the harness and Paperclip skill. **Proposed behavior** The default manual contains only the eight-word company identity. Shared startup prompts retain identity and connection guidance. Ordinary resume deltas retain current work context without the generic execution contract. **Reason and benefit** Let the stock harness guide general work. Keep Paperclip-specific capabilities and independently test default hires, skills, ordered comments, and chat restart. **Breaking changes** New default hires receive less guidance. Existing saved manuals, explicit custom bundles, CEO templates, and specialized wake contracts retain their behavior. The obsolete includeExecutionContract option remains accepted for source compatibility. ## What Changed - Reduce the default hire manual to one sentence. - Reduce shared task/chat defaults and remove the generic ordinary-resume contract. - Keep connection guidance, auth, skills, custom prompts, and specialized wake context. - Add credential-free instruction-boundary gates and 26 explicit Product E2E cells across eight legacy/native profiles, including two focused Paperclip-storage cases. - Capture public hire receipts before providers run, then grade delivered prompts and independent task/chat outcomes. - Add an early legacy skill API recipe for saving a task document, checking the saved revision receipt and linking the document. Improve stock task/heartbeat skill-selection metadata and show a clickable Markdown UI-link example. Keep native tool completion separate. - Advertise bounded routing descriptions and exact successfully staged SKILL.md paths in legacy ACP Claude; keep full bodies on demand and preserve remote path rebasing. - Remove only the provider environment from copied persisted ACP session records, while loading current run credentials and preserving all other options/conversation state. - Regenerate both capability metadata inventories and reject stale manifests/inventories before provider admission. - Publish the original reduction and focused skill-repair comparisons, preserving all failures, automatic recovery, cost coverage and limitations. ## Verification **Behavioral qualification remains pending.** Original legacy ACP Claude loses the issue document only in the reduced cohort beneath an unchanged credential failure. A source-backed diagnosis finds that neither ordinary assignment reads the staged operational skill, while the runtime persists provider environment in session state. The new common repairs expose skill metadata/path and omit persisted env; strict document and credential guards stay intact. [Inspectable diagnosis and retained hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md). Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb` incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged #14961/#15007). All 222 affected adapter tests, adapter-utils/E2E typechecks, and final 96 variant/grader/retry calibrations pass. The frozen historical comparator is `c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths identical, exactly two production instruction paths plus nine declared unit expectations differ. The operational skill/discovery/environment repairs, selected model/profile/task/core grader/auth/permissions/retry policy are identical. Both actual launcher prepare→verify admissions pass with zero providers. [Immutable manifest and exact receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json). One original legacy ACP Claude cell per variant is authorized, with enforced single campaign attempts, 12-minute deadlines and company/agent 1,000-cent hard stops; every product recovery run/cost is counted. Actual live outcomes are pending. Current normal CI has one failed server shard and failed aggregate verify under diagnosis; other normal gates including typecheck/build/Rust/all eight browser shards pass. Fresh review completed successfully; the valid historical startup/resume masking finding was fixed with per-invocation task/chat checks and strict complete-snapshot capture, calibrated and resolved. Prior heads, failures and campaigns below remain historical evidence, not checks on this repair head. - Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52 current-head checks pass with two intentional Storybook skips, including repository typecheck/test/build and the browser shard. Fresh Greptile is 5/5 with zero unresolved review threads. Exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained receipt verification, zero providers/source errors. Fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck and 26-cell stock discovery pass. Canonical contract/inventory checks and the later issue-derived reference calibration are retained; that reference-only follow-up is not live-qualified by earlier frozen runs. - Prior full repository typecheck/build passed. The complete local Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three timing failures. All three affected files passed unchanged narrow reruns; original failures remain retained. Current-head CI now passes the full general checks; the original local failures remain retained. - The original 24-pair default-manual/shared-prompt comparison has two new overall classic Claude/OpenCode document-delivery failures plus an additional legacy ACP Claude document loss beneath an unchanged credential-guard failure (not closed by later runs), two newly passing OpenCode ordered cases, seven unchanged failures and 13 unchanged passes. Equal 15/24 totals do not establish behavioral equivalence. [Complete original report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md). - The skill-only repair holds the eight-word manual/shared prompts and merged #14920 fixed. All four matched profile configurations and 203 fixture/behavior files match. Candidate `abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill sources against baseline `bc83fe030234439ac51279502a28803958963e2e`. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547) and [baseline workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047) each pass 571 exact-source prerequisites before providers; all eight cells clean up successfully. Failed campaigns publish successfully and remain failed. - Repair pairs: Claude original Fail → Pass; Claude explicit Pass → Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff worsens beneath the unchanged failing UI-link grade: baseline gives a clickable API URL, candidate gives a code-formatted path without an anchor. The request's usable-link wording is narrower in the UI-only oracle. [Complete repair report and safe projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md). - The subsequent narrow stock metadata/link correction has two matched Pass → Pass cases, zero new machine failures/passes and no pending pairs. Both original-case handoff links remain deficient: candidate uses a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the preserved original oracle only requires a durable document. Both explicit clickable UI-link cases pass revision/content/link grading. All four exact-source 587-check gates, single assignment runs and cleanup pass. This does not establish fix causality because baseline also succeeds. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401) freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374) freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only comparison with reduced manuals/shared prompts held constant, not a repeat of the historical-manual comparison. Only original and clarified explicit classic OpenCode cases are selected, two per variant/four expected turns. 8,242 other tracked files and both profile hashes match; protected workflows admit each exact source before credentials. [Complete qualification report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md). Candidate original loads Paperclip/reference before saving publicly; baseline original loads it after writing locally, then saves publicly within the same assignment. Reported cost totals are $0.0107824490 candidate / $0.0107909015 baseline, with unmetered runtime. The later reference-only issue-derived link correction is provider-free calibrated and **not live-qualified** by these frozen runs; no further paid runs. - Retained tool calls show the repaired original OpenCode assignment loads only its assigned output skill before writing locally. Operational Paperclip is first loaded during automatic disposition recovery; its early recipe is visible then, but it never saves the missing document. Explicit candidate loads Paperclip and reads the new reference before saving successfully. All nine actual runs are counted. Reported LLM totals are $0.3802537209 baseline and $0.4918990161 candidate; local runtime is unmetered. - Initial setup, packaging, cancelled/missing-cell recovery, callback test and relative-output attempts remain retained. No completed provider failure was rerun. Frozen measurement branches are unchanged by later canonical metadata maintenance. - Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`, and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly for paid execution; it is excluded from `--all`. Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` replays this PR on merged hiring #14985 (`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four explicit custom-CEO-bundle checks, minimal generic manual boundary, and both suites. The combined fixture catalog and hiring calibrations pass 67 assertions; exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained-receipt verification, zero providers/source errors, fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E typecheck and 26-cell stock discovery pass. Fresh current-head CI passes all 52 checks with two intentional skips, and fresh Greptile is 5/5 with zero unresolved review threads. The prior source-plan browser failure is retained: a deterministic process fixture replayed its last `fixture:plan` command on `chat_task_completed`, writing revision 2 with identical body after the approval handoff. This was not paid provider execution. Rebased current-head CI passes the same assertion without an old-head retry or a change to that browser fixture. The merged hiring change was measured separately on immutable matched unions, with this reduced/shared/operational context and native completion guidance held constant. [Complete original two-profile report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md): [candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208) / [historical baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463), 705 provider-free prerequisites each. Both pairs are unchanged Fail → Fail on the exact-five count, with six core delivery checks passing all four cells; 28 actual successful runs include eight automatic completion wakes, zero retries, four successful cleanups. Source-read coverage is uncomparable, actual model charges unknown. Separately versioned provider-free accounting remains analytical work; original verdicts are preserved. This does not rerun or qualify the completed default-manual or native campaigns. ## Risks - Legacy ACP Claude's additional delivery loss is not closed by any later matched run and blocks the no-extra-failing-behavior merge criterion. Legacy document delivery may have relied on the prior manual/shared prompts. The early skill repair improves Claude in one trial; the later OpenCode pairs pass in both variants and cannot establish causality or robust recovery. Both original-case links remain deficient beneath the storage-only grade. The later issue-derived reference correction has only provider-free validation. Native finish/block descriptions must not be supplied to legacy agents. - The comparison holds merged native Codex fix #14920 constant; it cannot measure that fix's before/after task performance. - These bounded skill/context/chat workflows do not measure general coding quality. Unrepresented providers remain unqualified. - Saved manuals and old Codex sessions are not automatically migrated. Codex through ACP still has a separate base-instruction follow-up. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution, and delegated PR/eval work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (relevant suites and all three unchanged narrow reruns pass; complete-run timing failures retained in Verification) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green on the new repair head (prior-head checks retained above) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the new repair head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
78e0034498 |
fix(evals): account for hiring completion notifications (#15007)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Product E2E evals check real hiring and delegated task completion. > - The hiring fixture requires three requested CEO turns and two coder executions. > - The server can also wake the CEO when each delegated task completes. > - Two exact-five-run guards rejected these valid completion turns in all four retained cells. > - This pull request validates bounded completion turns in both guards. > - The benefit is accurate workflow grading while all actual runs and coverage failures remain visible. ## Linked Issues or Issue Description Refs: #14985, #14948, #14961. **What happened?** The original hiring comparison reports Codex Fail → Fail and Claude Fail → Fail. Each cell has seven successful runs. The five requested work turns are accompanied by two server task-completion notifications. All six other delivery checks pass. **Expected behavior** Require exactly three distinct user-requested CEO turns and one coder execution for each of two known tasks. Admit at most two strictly attributed server completion turns, including one turn that batches both tasks. Reject unknown, duplicate, failed, retried or extra-work runs. **Steps to reproduce** Inspect the retained four-cell report linked below. Each original result fails `five-successful-turns`. The same exact count was also enforced by the final chat-flow guard. ## What Changed - Add one typed lifecycle helper shared by the hiring scorer and the hiring-only final chat guard. - Validate public run ledgers, company/user/account identity, request attribution, task origins, completion deliveries, timing and replies. - Keep exactly five required work turns; declare seven maximum total turns for cost and timeout planning. - Count all actual runs, including notification runs and unexpected resets. Keep other chat count guards unchanged. - Version the hiring grader as v3 (turn accounting v2) and include the helper and chat guard in its definition digest. - Keep source-read, exact coder-body and all six other delivery checks unchanged. - Add 144 focused helper/scorer/settlement calibrations and separately versioned exact retained-input replay reports. - Retry complete bracketed observations, await both owed callbacks and attributed replies, and refresh the final guard consistently. - Reject unrelated completion writes and failed mutation attempts using exact canonical/native action IDs. Missing identity mapping is uncomparable action coverage. ## Verification - All 977 credential-free E2E support tests pass across 64 files, including 144 focused lifecycle/action/scorer/settlement calibrations. - E2E typecheck, ordinary plugin SDK and Runner TypeScript dependency builds, capability contract/inventory checks and the existing two-cell hiring discovery pass. - [Executable replay report](https://github.com/paperclipai/paperclip/blob/fed1729018cc100f5f4bbfb692777e49009c423b/doc/plans/2026-10-02-hiring-executable-accounting-replay.md) pins current code revision `e4077ade1818d98b9862ae79ee1d49a007dcf9c1`, v3 definition digest, exact source/input hashes and each original/new check. - The stricter replay verifies both Codex variants through both executable guards. ACPX Claude action attribution remains unresolved/uncomparable because provider execution IDs cannot be exactly joined to native request IDs; guards fail closed. No notification writes are observed. All six other outcomes and every original source/template coverage check stay unchanged. Original files and Fail → Fail machine verdicts remain preserved; zero providers are called. - Full attempts remain uncomparable in both profiles. Historical Claude also keeps its six-backtick exact-template mismatch. This grading repair does not prove model-performance equivalence. - The limited sidecar-v1 and initial executable-v2 passes checked notification-created tasks but could miss unrelated document writes. Those assessments remain preserved and do not prove harmless notifications. The stricter v3 replay is separate. - [Original measurement and separate sidecar](https://github.com/paperclipai/paperclip/blob/8eb517ca1497687237163bdef4dfc4d3332ea916/doc/plans/2026-10-02-hiring-template-live-comparison.md) retain 28 actual runs, eight automatic notifications, four successful cleanups and unknown actual model charges. No models are rerun. - The branch is replayed on master `59c07ede7`. Intervening master changes are UI-only; eval source bytes and replay verdicts match. The four-cell provider-free replay was repeated against the reachable code revision. - Initial-head normal CI retained browser failures in agent-run denial feedback and touch-picker scroll position. Those browser paths and imports were unchanged, but their cause was not established. The necessary review-fix head passes both browser checks; no blind rerun was requested. - Local full repository typecheck/test/build were not repeated. Exact-head normal CI passes the required repository gates, including typecheck, tests, build and browser shards. Fresh Greptile review completed on `fed1729018cc100f5f4bbfb692777e49009c423b` with 5/5 and zero unresolved threads. An independent rerun of the 144 focused helper/scorer/settlement tests passes on the unchanged head. **Merge readiness:** This PR repairs the evaluator. Its positive and negative calibrations pass, both guards reject missing action attribution, current-head CI and review pass, and there are no merge conflicts. The retained ACPX cells remain uncomparable because their action IDs cannot be joined. That coverage limit remains a separate follow-up; it does not require relaxing this grader or changing the old results. No model calls, production instructions, carrier changes, or historical regrades are part of this readiness update. ## Risks - Missing or inconsistent public lifecycle evidence fails the bounded helper. The focused calibrations reject plausible false positives and malformed observations. Unmatched action IDs fail closed and are reported as uncomparable rather than a model task regression. - Source-read evidence remains incomplete. This PR does not change provider event carriers or relax the coverage oracle. - The versioned count check differs from original v1 results. Reports retain both versions and exact input hashes. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution and delegated calibration work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip normal CI gates are green (exact head `fed1729018cc100f5f4bbfb692777e49009c423b`; fresh review tracked separately below) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (completed exact-head review; zero unresolved threads) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dd868ed125 |
fix(runner): share native completion tool guidance (#14961)
## Thinking Path > - Paperclip manages AI agents and their work. > - Native Runner agents report completion through finish and block tools. > - The providers receive different descriptions for those tools. > - Completion guidance belongs with the tools that enforce the result. > - This pull request shares the descriptions and refreshes retained catalogs. > - A separate native suite checks completion and blocking on production defaults. > - Legacy agents retain their separate skill and API paths. ## Linked Issues or Issue Description Refs: #14920, #14948, #14985. **Current behavior** Native Codex and MCP bridges describe finish and block differently. Retained provider sessions can keep old descriptions. **Proposed behavior** Native providers receive the same finish and block descriptions. The descriptions cover report selection, validation feedback, returned outcomes, approval gates and final-answer timing. Retained native sessions refresh from v13 to v14. **Reason and benefit** Put the completion procedure next to its native tool. Preserve stock base instructions, schemas, permissions and terminal semantics. This PR now stands alone on master. It contains no reduced manual, shared prompt or operational-skill changes from #14948. ## What Changed - Add canonical native finish and block descriptions. Use them in direct Codex and both native MCP bridges. - Advance the native tool contract to v14. Cover old-v13 refresh without replacing task identity or prior history. - Check authenticated tool catalogs, provider start/resume frames and serialized daemon catalogs. - Add an independent, explicit-only native completion suite. Preserve the original assigned-skill durable-document journey. Pair it with a concrete whole-task blocker across Codex, ACPX Claude and OpenCode. - Verify the actual public production default bundle and budgets before execution. Require independent durable disposition, native result/terminal receipts and observable provider-final ordering. - Correct the blocker browser oracle to accept the requested explanation. Keep exact owner/action/scope checks. Calibrate positive, missing and contradictory replies. - Preserve only actual `tool_call` terminal names (`paperclip_finish` / `paperclip_block`) in the native compatibility run-log projection. Require the same named call ID through its finishing result; retain all other redaction boundaries. - Admit verified hosted shallow checkout/build hydration and bind the selected runnerd to exact source/archive/binary provenance. Hosted cells truthfully reuse the existing trusted build; local admission executes Rust calibration. Forward only public source/run identifiers through both launcher preflight subprocess paths. - Enforce single attempts in the launcher for opted-in fixtures. Keep ordinary retry policy unchanged. Run exact-source, credential-free admission before credential loading. ## Verification - Frozen candidate: `d6e59e4712a3158ab4cd7d58deff1389b4578c21`, based on master `59c07ede72dc08b8aba149a01cc11e0b7a204621`; historical descriptions: `e74ed61a69fbdd8b3a8f15dd6456bc3140246e33`. Exactly the five original native production files and six unit tests differ. Both carry identical corrected fixtures, strict named finishing-call grader, closed compatibility carrier and admission. Defaults, profiles/models/auth/permissions and manifest bytes match. - Actual launcher `prepareNativeCompletionPreflight` → `verifyNativeCompletionPreflight` admission passes on both exact refs with zero providers: candidate 132 / historical 127 selected TypeScript assertions, 128 Node calibrations and one Rust normalization calibration each; E2E typecheck, manifest checks, selected binary provenance and six-cell discovery pass. Each has 257 explicitly skipped unrelated assertions, not coverage. The credential-free environment calibration exercises both real prepare/verify subprocess options with public hosted identifiers and rejects credential/ambient overrides. Complete actual launcher prepare→verify also passes on both frozen refs with explicitly synthetic hosted metadata/verified archives, separately labeled as calibration rather than a trusted GitHub run. Exact framed provenance parsing and mock source identity are calibrated without relaxing the real verifier. - [Complete matched qualification report](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-master-qualification.md), [immutable manifest](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-manifest.json) and [closed retained audit/hashes](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-results/comparison.json) are inspectable. All six candidate cells pass; historical descriptions pass five. Paired outcomes: **zero new failures, one new pass (Codex blocker), five unchanged passes, zero pending pairs**. [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37098728980) and [historical campaign](https://github.com/paperclipai/paperclip/actions/runs/37098815696) each execute six original attempt-1 native runs, with no campaign retry and successful cleanup. Their trusted workflow revision is `215586d127e97c9301d86e769a39a15c13298ca2`, separate from measured source. [Candidate public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098728980-1/index.html) and [historical public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098815696-1/index.html) retain declared screenshots. - Independent candidate evidence agrees with all original grades: 51 strict native checks, 12 served-default/budget checks and 21 original skill/document checks pass. The historical Codex blocker saves the correct whole-task blocker but omits the required marker from its actual provider final and identical saved reply. This is not semantic-summary fallback. Its original browser/matcher failure stays retained; the additional native snapshot/grade and workspace before/after digest were never written and are not fabricated by the separate API/PRP audit. Historical Codex completion has one failed finish followed by success within the same native run; the public receipt records no failure reason. All twelve runs and their usage remain counted. Reported model-cost subtotals are $0.00421482 historical/$0.00437391 candidate; Codex/Claude zero entries have unknown billing type, actual invoices are unverified and hosted execution cost is unmetered. One matched trial supports no extra failure within these six cases, not broad statistical or coding-quality equivalence. - Initial hosted `e18c2cf9` / `459455ac` and subsequent `0a9c5a7` / `00a761b` cohorts each stopped before providers in all twelve cells. The latter failed a mocked-receipt unit test under ambient hosted metadata; all source/build proofs passed. [All twelve later setup receipts](https://github.com/paperclipai/paperclip/blob/402ee94c52273ad58de355ae9a7d562dd22f8101/doc/plans/2026-10-02-native-completion-qualified-hosted-setup.json) are retained. [Exact failed setup receipts](https://github.com/paperclipai/paperclip/blob/27653eb1a8f8ce839776d760f4563f672e5a706c/doc/plans/2026-10-02-native-completion-master-hosted-setup.json) and the original manifest remain intact. Local sandbox-denied loopback and stale anchor-expectation attempts are retained separately; unchanged appropriate assertions were corrected/admitted before paid dispatch. Old anonymous OpenCode streams are not assigned inferred tool names or retroactively passed. - Full provider-free E2E support previously passed 927 tests in 67 files. Exact-head d6 normal CI run `37098409915`, attempt 1 passes full repository typecheck/build/tests, Runner Rust/static checks, all browser shards/aggregate and canary: 52 check-runs pass, four intentional skips, Snyk passes. Fresh Greptile check `111132956342` is 5/5 with zero unresolved threads. Source-specific deterministic tests do not substitute for the bounded live comparison. - Earlier native source `9138f570c341c251a5727c32d6615ce238bc8e03` is archived. Its [complete reduced-manual-context report](https://github.com/paperclipai/paperclip/blob/9138f570c341c251a5727c32d6615ce238bc8e03/doc/plans/2026-10-02-native-completion-live-comparison.md) remains intact, including original failures, grader limits and provider-free replay. It is not current-master-context qualification. ## Risks Changed tool text can change model behavior. The completed six-pair qualification shows no extra failing outcomes in this bounded trial; other tasks and repeated-run variance remain unmeasured. Observable final ordering does not prove provider feedback consumption. Public evidence can fail closed if a provider does not expose the required result sequence. This slice does not remove native fixed prompts or measure general coding quality. No database, schema, permission or legacy completion changes occur. ## Model Used OpenAI Codex, GPT-6 family, with code inspection, execution and tool use. The exact deployment ID and context-window size are not exposed in this session. They are unavailable rather than inferred from the model menu. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
862a5758ba |
fix(agents): reduce hiring templates to role descriptions (#14985)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - New agents receive role instructions from onboarding, the hiring skill, or a team package. > - These sources repeat harness procedures and impose generic work policies. > - They can crowd out the task and the harness instructions. > - This pull request reduces those sources to short role descriptions. > - It preserves configuration, skills, authentication, reporting lines, and approval controls. > - The benefit is less repeated instruction text with explicit coverage for the default hiring path. ## Linked Issues or Issue Description Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection. Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes. ## What Changed - Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets. - Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples. - Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog. - Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions. - Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents. - Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse. - Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison. Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing. ## Verification - PASS: 99 focused server tests and eight shipped-catalog tests. - PASS: catalog generation and validation for four shipped teams. - PASS: hiring skill validation. - PASS: `pnpm -r typecheck`. - PASS: `pnpm build`. - INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419. - PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests). - PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog. - PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck. - PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads. - PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge. The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence. ## Risks - New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome. - Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break. - Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review. - The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback. - Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
018993140f |
feat: let agents name prompt-only tasks (#14761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cbd278dc03 |
fix(interactions): derive chat recipients and validate explicit users (#14742)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agents use saved questions to get human input and continue the same task. > - The standard question example recently told models to copy a user ID. > - A model can omit an identity prefix and create a question its intended recipient cannot answer. > - Agent Chat already knows the conversation owner, so the server can supply that identity. > - This pull request removes the blanket instruction and validates explicit recipients before saving. > - Ordinary questions stay simple, and explicit addressing remains available for decisions that need a particular person. ## Linked Issues or Issue Description Refs #14707, #14188. Related: #14238 handles legacy email recipients; this change prevents invalid recipients in new cards and retains exact ID matching. **What happened?** A model copied a Cloud user ID without its prefix into `addresseeUserId`. Creation succeeded. The intended user's answer then failed the exact recipient check. **Expected behavior** Ordinary chat questions use the saved conversation owner. A task may optionally name a specific recipient. The API rejects an unknown or unauthorized recipient before it creates a card. **Steps to reproduce** Create a chat question for a user whose ID is `paperclip-id:example`. Supply `example` as the addressee. Before this change, creation accepts the invalid recipient and the owner cannot answer. With this change, creation returns 422. Omitting the field saves the full owner ID and allows that owner to answer. ## What Changed - Remove `addresseeUserId` from standard question examples and remove the blanket requester-ID instruction. - Derive the recipient of ordinary chat questions from the persisted conversation owner. Reject conflicting explicit user IDs. - Keep explicit task recipients optional. Validate supplied user IDs with the existing board mutation policy, including company, viewer, and Cloud restrictions. - Preserve explicit agent routing, connector intents, confirmations, exact recipient checks, idempotent retries, and no-login local-board authority in local-trusted mode. - Update the blocker grader to accept an omitted recipient and verify the actual requester answered. - Add database and HTTP tests for prefixed identities, denied recipients, concurrent retries, saved answers, and response delivery. ## Verification - Database interaction service suite: 90 tests passed, including implicit local-board creation/answering and authenticated/Cloud denial. - Interaction HTTP route suite: 84 tests passed. - Affected interaction/native/connector/documentation suites: 231 tests passed across six files after valid-user fixtures were updated. - Resolver and interaction unit suites: 29 tests passed. - Product E2E unit/calibration suite: 793 tests passed; Product E2E typecheck and blocker catalog discovery passed. - Generated API-reference and capability contract checks passed. - `pnpm -r typecheck` and `pnpm build` passed. - Full local `pnpm test:run` did not finish green: its initial general-server pass had 14,416 passing assertions, one unrelated native-resume assertion failure on macOS, and three teardowns from an intermediate fixture cleanup fixed above. Separate broad local groups also encountered timeout/live-port failures under host load. Local UI (7,026), CLI (502), shared (817), and skills-catalog (20) tests passed; the complete final-head CI matrix is the broad verification gate. - After two CI cold-start readiness timeouts, a separate test-only commit gives the first exposure lifecycle fixture the existing normal 30-second readiness budget. Its real HTTP, ordering, and cleanup assertions remain intact; the targeted case and final Linux CI shard passed. Production deadlines are unchanged. - A separate OpenCode fixture failed twice on GitHub-hosted Ubuntu because its cached Node executable was group-writable; the same case passed on AWS runners. The fixture now qualifies its own Linux copy with mode `0500` and the actual copy digest. Host files and production security checks are unchanged. The focused macOS case passed; the new Linux-copy branch also passed on the final AWS-hosted Linux runner (1,125 passing Runner tests, 3 skipped). The final run was not on a GitHub-hosted runner. - Final-head [CI run 36762078176](https://github.com/paperclipai/paperclip/actions/runs/36762078176) passed for `116b968b24fa0a8c5724a7bf96e73a8dda5f0425`: 54 successful checks and two conditional Storybook skips, with no pending or failed checks. The 27 general/serialized test jobs reported 28,635 passing tests. Typecheck, build, Runner, browser E2E, and Canary gates passed. Greptile reviewed that exact head at 5/5; both review threads are resolved, with no open follow-ups. - No live provider replay is claimed by this PR. ## Risks - New explicitly addressed cards reject users who cannot mutate the issue, including viewers, inactive members, and invalid IDs. Callers that supplied invalid recipients must correct their request. - Existing addressed cards are not rewritten. Existing authorization checks remain strict. - Chat inference applies only to questions without an agent addressee. Connector intents and governed confirmations retain their own recipient paths. - No schema change or migration is required. ## Model Used OpenAI Codex, GPT-6 (exact serving variant and context window are not exposed in this environment). Used reasoning, tool use, code editing, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b54b2dc35c |
fix: preserve warm Codex turns with incremental managed file checkpoints (#14735)
## Thinking Path > - Paperclip manages AI agents and keeps their instructions and files durable. > - Native Codex runners can keep a process alive between compatible turns. > - Managed file collection stopped that process after each turn, which defeated warm reuse. > - Agent folders can contain large images and other files, so full copies on every turn are expensive. > - This change keeps one managed directory for the live session and saves only file changes after each turn. > - Ownership, authorization, instruction changes, and process retirement still control when reuse is safe. ## Linked Issues or Issue Description Related: #13710 introduced native warm session reuse. This fixes managed file collection that still forced those sessions to stop. No duplicate open PR or issue was found. **What happened?** With managed instructions and warm native Codex enabled, consecutive turns reused a Daytona sandbox but started a new runner process each time. The managed directory collector required process termination before saving files. **Expected behavior** Compatible turns keep the same process and managed `AGENT_HOME`. Each completed turn saves added, changed, and deleted files before the next turn starts. Unchanged large files do not transfer again. **Steps to reproduce** 1. Use a native Codex agent with managed instructions and a reusable Daytona environment. 2. Enable warm session reuse and run three turns on the same task. 3. Write a large binary on the first turn, edit a small note on each turn, and delete a file on the second turn. 4. Compare process identity across turns and read the canonical files through the public agent-files API. **Paperclip version or commit** Reproduced on `d30b03bd8c17604cdab1533eeeeb087aba30e8b1`. **Deployment mode** Local server with remote Daytona execution; cloud native runner uses the same path. ## What Changed - Retain the managed directory only for the verified owner of a live native Codex session. - Checkpoint each completed turn before releasing the session for reuse. Retry unstable captures, then stop and collect when a warm checkpoint cannot be validated. - Compare metadata and cached hashes, stream only changed file payloads, record deletions, and validate path, content, quota, and authorization before saving. - Rotate sessions when canonical files, loaded instructions, credentials, or launch policy change. Fence stale collection and cleanup callbacks from later owners. - Keep cleanup and recovery aware of the current session owner. Recheck canonical files under the writer lock at handoff, attach the successor collector before fallible bookkeeping, and emit one final save receipt on checkpoint fallback. Preserve storage warnings across unchanged checkpoints. - Add regression coverage and a three-turn Daytona test with independent public API file checks, an unchanged 8 MiB binary, deletion checks, and strict process identity checks. - Document checkpoint consistency, lifecycle behavior, and local run-log counters. - Replace a timing assumption in the Daytona teardown test with explicit transfer-arrival gates after CI exposed an unset release callback. ## Verification - Full local `pnpm -r typecheck` and `pnpm build` passed. Server checks were repeated after the final storage-warning fix. - Runner E2E typecheck and 749 runner E2E unit tests passed. - Focused file checkpoint, directory ownership, instruction collection, native session, and merge tests passed. After review fixes, the managed-directory and native-session suites passed 550 tests, including intervening canonical edits, same-run fresh restore, failed handoff collection, and one-call fallback collection. Server typecheck passed again. The Daytona plugin suite passed 218 tests. The quota-warning regression failed before the fix and passed afterward. - Three real Daytona campaigns passed before the final handoff review fixes. The latest kept PID 547 across all three turns. The first checkpoint copied 8,388,635 bytes; the next two copied 36 and 54 bytes. Public API reads verified the binary, note contents, and deletion after every turn. Test cleanup deleted the sandbox. - The final head was also deployed to an isolated cloud staging instance and passed three UI-triggered native Codex turns with managed instructions. All three retained the same process ID/start time, native session, provider session, runner instance, and Daytona sandbox. Checkpoints copied 8,388,643 bytes on turn 1, then only 52 and 78 bytes on turns 2 and 3; those warm captures also hashed only 52 and 78 bytes. Independent canonical API reads verified every byte of the unchanged 8 MiB binary and the exact note contents after every turn; the deleted file returned 404 after turns 2 and 3. After restoring the original lifecycle and agent-auth configuration, removing the temporary secret, pausing the test agent, and deleting both test sandboxes, independent canonical API reads still verified the entire binary, the final 78-byte three-line note, and the deletion. The native runner flag remained enabled and the final serving revision remained the PR head. - Two earlier staging attempts are preserved as failures and are excluded from the acceptance result: a saved ChatGPT login failed with a provider routing 401, and its subsequent stopped-sandbox retry failed before provider startup with a closed-lease admission error. The successful campaign used a fresh sandbox and a temporary encrypted API-key binding. The stopped-lease retry remains unexplained; this campaign does not establish recovery of that failed sandbox. - All [Paperclip CI gates](https://github.com/paperclipai/paperclip/actions/runs/36750397355) pass on `26ef2ef56a389259246809805c0b34a4747eb86b`, including full test partitions, build, typecheck, runner verification, E2E shards, and the Canary clean public-npm install. Greptile reviewed that exact head at 5/5 with no unresolved review threads or outstanding findings. - Full local repository coverage used the existing CI partitions, but the 40,000-file Git streaming stress test timed out and its local retry was interrupted by macOS thermal emergency sleep; this is not a green full local suite claim. The exact stress test passed on the final head in [CI server shard 2/12](https://github.com/paperclipai/paperclip/actions/runs/36750397355/job/110008294290), in 111.9 seconds. - Repeat the live test with configured credentials and a Linux runner artifact: `pnpm test:e2e:runner -- --id daytona-warm-continuity.runner-codex.daytona.warm-three-turn`. ## Risks - This is a file-level checkpoint, not an atomic snapshot of the whole folder. Background writes after a capture are saved by the next checkpoint or final stopped collection. - Metadata scans still visit all paths. Modified files transfer in full; unchanged files do not rehash or transfer. - Incorrect ownership or reuse could collect the wrong directory. Run ownership fences, current authorization, stable capture validation, and stopped collection fallbacks are covered by tests. - Warm reuse remains opt-in. No database migration or fleet default changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code editing, tool use, and test execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3c561642b4 |
fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask for decisions and optional details through cards in chat. > - A clear approval in a message can leave the matching card pending. > - An unanswered question can also block an unrelated later reply. > - Decisions need a saved source message, while optional questions need to remain answerable in history. > - This pull request records conversational decisions and lets users move on from questions and answer them later. ## Linked Issues or Issue Description **What happened?** Native Claude and Codex could act on approval in chat while the original approval card stayed pending. Pending question forms stayed above the composer, were absent from history, and could suppress later chat replies. A late native question answer could wait for a finished run to reconnect. **Expected behavior** The active agent records a clear approval or refusal against the exact card and user message. Ambiguous replies do not grant consent. Users can send another message without answering a question. The question remains pending in history and can be reopened and answered later. The saved answer reaches the agent. **Steps to reproduce** 1. Ask an agent to propose work with a confirmation card, then approve it in chat. 2. Check that the original card records that approval before work starts. 3. Ask an interactive question, send an unrelated message, and reload. 4. Open the unanswered question from history and submit an answer. Related work: #14408 added completion delivery. #14607 tests completion reporting turns. Neither records conversational answers on approval cards. ## What Changed - Add a confirmation endpoint backed by a user comment, with schema validation, OpenAPI discovery, and native Plan-mode access. Ask mode remains read-only. - Check company, active run, actor, current session, message provenance, revision, and resolver policy. Save the decision and audit in one transaction. Retries do not repeat effects. Emit resolution telemetry after commit. - Give fresh and resumed chat turns the actual pending confirmation identities. Teach agents to save clear conversational decisions before acting and to clarify ambiguity. - Keep unanswered Agent Chat questions as compact history entries. A newer user message closes the old form. Question cards never contribute to composer pending counts or navigation, including after dismissing a fresh form. The history card is the sole reminder; clicking it restores that exact form and draft. - Preserve Agent Chat questions when later messages or questions arrive. Historical ordinary inputs no longer gate later chat replies. Current-run requests, task execution, and governed approvals keep their gates. Remove the special acknowledgement-publication proof helpers that this rule replaces. - Route answers to finished native runs through durable fresh-wake delivery, with existing idempotency and source-question context. Settle late replies against contiguous completed conversation turns and freeze their history replay; failed, unhandled, and newly arriving messages remain actionable. - Add real-component Storybook scenarios, database and UI regressions, and a three-turn native Claude/Codex E2E case. Capture distinct, UI-ready screenshots and report the individual assertions. ## Verification - Focused decision/publication/UI regressions after merging master: 288 passed; subsequent UI draft, failed-send, and conversation checks: 199 passed. - Native question and durable delivery regressions: 106 passed, including all four terminal run states and exactly-once late delivery. Seven targeted regressions fail against the original implementation and pass with the fix. - Latest conversation/decision/native-delivery regressions after the master merge: 121 passed. Covers completed progress, missing or failed intervening turns, new messages during a late reply, stale sessions, and frozen retry/replay boundaries. Four new assertions fail before the ordering fix. - E2E support suite after the master merge: 792 passed. Negative controls reject expired cards, wrong questions/answers, stale or missing replies, unrelated clarification forms, and unexpected tasks. - The embedded-browser walkthrough caught one additional defect: dismissing a fresh question still showed a composer badge. Both Cancel and close-button regressions failed before the fix. The fix at `65f2ade12` passes 170 chat-thread tests and 792 E2E support tests. After merging master, 232 chat-thread/confirmation tests, server/UI typechecks, and token gates pass. The preview and two-provider live E2E pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at that commit. All 55 checks are now successful at `e5512a206` (four conditional checks skipped), including the aggregate verification gate and clean-install canary test. The first attempt was interrupted by simultaneous CI worker shutdowns; one failed-job rerun passed without code changes. - [Published Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on): nine real-component scenarios. Manually exercised move on, reopen, preserve draft, answer later, answer one of multiple questions, and a custom mobile answer in the embedded browser. Retested fresh Cancel and close-button dismissal in the updated build, then reopened and submitted the preserved Green selection and inspected its answered receipt. Static preview has no live model/backend; its callbacks are fixture responses. - [First live campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/) reproduced the late-answer completion-state defect on both providers despite correct saved answers and acknowledgements. It also exposed a valid imperative clarification rejected by the old oracle. Both issues are fixed with regression controls; this failing run is retained as evidence. - [Four-cell qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/) passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous confirmation, each on native Claude and Codex. Inspected saved state, source-message decisions, visible cards, and agent replies. Both late-answer chats settled to waiting; no unrequested tasks were created. [Final branch rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/) passed 2/2 at `142630720`: the same unanswered-question journey after merging master, plus an additional screenshot and browser assertion for the actual late-answer acknowledgement. - [Composer-reminder E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/) passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each, with explicit no-badge assertions before and after reload. Inspected saved pending/answered state, both screenshots with a clear composer, and actual Blue acknowledgements; all five behavioral matchers passed per provider and neither created tasks. Cost coverage is partial; this is bounded workflow qualification. - [Fresh-dismissal E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/) passed 2/2 at `e5512a206`: native Claude and Codex, including fresh Cancel, clear composer, reopen, unrelated message, reload, late Blue answer, and actual agent acknowledgement. All five behavioral matchers pass per provider. Inspected the fresh-dismissal screenshots and saved pending/answered identity; neither created tasks. Cost coverage is partial (4/6 runs). - Prior evidence remains available in [the earlier campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/). Its early loading screenshot and overwritten final capture prompted the UI-ready, distinct screenshot fixes. ## Risks - The model interprets intent. The server verifies permission and provenance; it does not infer consent from text. Ambiguous and unrelated replies are not approvals. - Historical questions can accumulate. They remain visible, pending, and answerable; no automatic answer or expiry is invented. - The change to completion gates is scoped to Agent Chat and ordinary historical inputs. Current-turn and governed approvals retain their existing controls. - Live qualification is limited to the selected stories. Broader native onboarding finalization remains separate work. - No database migration. Telemetry adds no fields or values; the contract and README document the commit boundary. Privacy review was requested on the PR. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser-test orchestration. The exact model ID and context-window size are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d30b03bd8c |
test: add persistent E2E coverage for human blocker decisions (#14707)
## Thinking Path > - Paperclip manages work for AI agents. > - Agents use the coordination skill when work needs human authority or a scope decision. > - PR #14188 replaced automatic manager escalation with direct blocker handling. > - This behavior needs real browser, server, database, and provider tests. > - The test must verify saved human input, task ownership, and resumed work. > - This pull request adds six reusable Product E2E cases and improves the skill examples that they exercise. ## Linked Issues or Issue Description Refs #14188. The merged change needs repeatable behavior coverage. The new suite tests missing administrator access, missing hiring permission, and requester scope questions. Searches found no duplicate blocker-guidance suite. This extends the existing eval system described in ROADMAP.md. ## What Changed - Add the explicit-only `blocker-guidance` Product E2E suite. It has three local scenarios on legacy Codex and legacy Claude. - Use the production UI and public APIs to create work, save a human-only question or confirmation, answer it after reload, and resume the same task. - Check requester identity, ownership history, manager activity, hiring, saved answers, and completion. Keep direct text input as a separate UX result. - Save pending and final screenshots, API checkpoints, skill hashes, provider evidence, and billing data through the existing report pipeline. - Isolate the Claude fixture home. Verify the served skill bytes before dispatch so an old installed skill cannot silently replace the evaluated skill. - Improve the coordination and hiring skill examples. Include the human-only policy, requester address, wake behavior, and handling of authorized scope changes. - Grader v5 requires the exact approved public welcome note as a new worker comment. Browser input checks reject unwritable scope cards before clicking, and confirmation direction must be saved in the resolution before the worker wakes. - Add grader calibration and browser-input tests. Update the fixture guide and generated capability inventories. ## Verification - `pnpm build`: passed after rebasing onto current master. - `pnpm -r typecheck`: passed. - `pnpm test:e2e:runner:typecheck`: passed. - `pnpm test:e2e:runner:unit`: 742 tests passed. - `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests passed. - `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells found. - Capability inventory and generated-contract checks: passed. - `pnpm exec vitest run server/src/__tests__/hiring-operational-examples.test.ts`: four tests passed after synchronizing the generated API reference and section anchor. - Full general and serialized test suites: passed in CI on `6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm test:run` was interrupted after complete CI coverage passed; it is not claimed as a completed local full-suite run. - Final GitHub checks: 54 passed, two optional Storybook checks skipped. The runtime-exposure startup test hit a 10-second readiness timeout once, passed a targeted local reproduction, and its CI shard passed the single retry without code changes. - Current-head Greptile: 5/5, clean check, zero unresolved threads. - Historical live measurement on September 29 at `4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex `gpt-5.6-sol` passed 8/9. These runs predate this rebase. - Version 5 changes the scope answer to an exact approved publication draft. The historical runs do not qualify that new requirement; the two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec` passed Codex and failed Claude. Claude posted the correct salary-free sentence but omitted its required reference line from that comment, placing the reference in a separate completion message. The `public-welcome-note` check correctly failed. An earlier Claude database-startup failure was retained separately; its fresh-instance retry reached the model. This pilot is not a six-cell qualification. - The failed Codex scope case omitted `addresseeUserId`. The strict routing check remains. All 18 attempts had clean evidence manifests and passed cleanup. - To repeat with provider credentials: `pnpm test:e2e:runner -- --suite blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is excluded from `--all`. ## Risks - The live suite measures variable model behavior. The retained 17/18 historical result and the current 1/2 scope pilot are not all-pass qualifications. These paid cases are opt-in; their observed model failures remain visible independently of deterministic CI checks. - A separate generic task-replacement diagnostic still exposed a Claude refusal. The ordinary cases use specific business decisions. The diagnostic is not a standalone catalog case in this change. - Earlier measurements included an old installed Claude skill and test defects. Their grades remain retained and are not combined with the three final repetitions. - Skill examples can affect when agents ask for human input. Downstream permission checks still apply. - Native runners, Daytona, agent-requester routing, and real external connection authorization are outside this suite. - Raw provider traces and credentials remain private. No screenshots, raw reports, secrets, workflow changes, or lockfile changes are committed. ## Model Used OpenAI GPT-6 through Codex assisted with this change. The exact deployed variant and context window size are not exposed in this session. The assistant used reasoning, repository edits, tool use, and shell execution. The evaluated models were `gpt-5.6-sol` and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4f9a7c613 |
test(runner): guard continuation after journals exceed 2 MiB (#14312)
Add an actual runner resume regression above the former 2 MiB journal boundary and an explicit-only three-turn Daytona workflow that grows real execution history. Verify journal size and distinct completed tool calls without exporting private payloads. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
fc36d1ccd3 |
test(e2e): qualify large Daytona Git workspace continuation (#14316)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can keep one agent process and workspace across human review turns. > - Large workspaces must survive each transfer between Daytona and the host. > - Small fixtures do not cross the previous 32 MiB filename-output limit. > - Correct files alone do not prove that finalization and recovery have settled. > - This pull request adds an explicit three-turn test with 60,000 files and independent host checks. > - The test detects lost files, replaced processes, and stale finalization retries. ## Linked Issues or Issue Description Refs: #14253, #14314, #14315, #14402, #14420. This test covers the workspace Git streaming fix in #14253 and the Daytona archive validation and native finalization fixes consolidated from #14315 and #14314 into #14402. These runtime changes are merged into master. This PR adds regression coverage. ## What Changed - Add one explicit-only native Codex Daytona test. Broad matrix runs do not select it. - Create 60,000 small untracked files through ordinary provider execution. Independently check all host file contents and the 39,828,890-byte filename manifest after each turn. Each continuation changes all generated file contents, so stale host copies fail. - Check spaces, newlines, leading hyphens, Unicode, and glob characters in filenames. - Require committed native finalization, successful workspace receipts, no active transfer or runless cleanup, and no scheduled recovery before each continuation and after the final turn. - Use a fixed external instruction bundle so the same runner PID and process identity can continue across the three browser-driven turns. Managed agent folders intentionally stop the process for file collection after #14420. Set the 20-minute idle window on the environment, then verify the admitted policy on every run. Set a 25-minute Daytona auto-stop window for this large-file test. - Save public workspace-operation evidence when an E2E attempt fails. - Bound this large-file fixture to 15 minutes per turn and 50 minutes total, reserving five minutes outside the turns for setup, host verification, and cleanup. A measured CI continuation succeeded in 11 minutes 5 seconds, exceeding the ordinary warm fixture's 10-minute deadline. The ordinary fixture and all file, process, finalization, and cleanup assertions remain unchanged. ## Verification - Final PR head: `b7eed5517426507f5d7912edb8ae49163d1783b0`. - `pnpm test:e2e:runner:unit` — 56 files and 708 tests passed locally. The catalog retains all 407 existing cells and adds this one explicit-only cell. - `pnpm test:e2e:runner:typecheck` — passed locally. - [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/36640920233) — all gates passed, including repository typechecks, tests, build, browser suites, and canary dry run. The PR has 54 successful checks and two expected skips. Fresh Greptile reviewed all seven files at 5/5; all review threads are resolved. - [Live single-cell Daytona verification](https://github.com/paperclipai/paperclip/actions/runs/36637283902) — passed on the first attempt in 1,822,299 ms (30m 22s) on `227074b7b`, using native Codex `gpt-5.6-sol` and the verified Daytona image. All three turns independently verified every one of the 60,000 file contents, 39,828,890 filename bytes, and five unusual names. All runs committed with one stable runner PID/process fingerprint, native/provider sessions, runner instance, and sandbox. Final browser/download assertions, all nine matchers, explicit cleanup, and report publication passed. Results are published through the [Product E2E history](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/) and [eval hub](https://pages.paperclip.ing/evals/). - The final `b7eed5517` follow-up only extends the total allowance from 45 to 50 minutes, updates its catalog assertion/version, and documents the setup/cleanup margin. Per-turn limits and behavioral assertions are unchanged from the successful live run; that paid run was not repeated for this allowance-only follow-up. - The [first integrated-head live attempt](https://github.com/paperclipai/paperclip/actions/runs/36632127364) is retained: turn 1 passed, then turn 2 hit the old 10-minute deadline while finalizing. Its trace records successful completion after 11m 5s and successful cleanup. Process diagnostics also showed the intentional managed agent-folder stop boundary introduced in #14420. These observations motivated the larger turn budget and fixed external instructions. - The broad local `pnpm test:run` began alongside the build and encountered server setup and port-test failures. The affected setup suites and assertions passed on focused reruns after the build, using canonical macOS temporary paths where needed. The redundant broad local run was stopped; full-suite success is established by final-head CI, not by that local run. ## Risks - The live test makes provider calls and creates a billable Daytona sandbox. It runs only when explicitly selected and deletes its sandbox during cleanup. - The test creates 60,000 files and can take several minutes per transfer. Its longer retention window applies only to this fixture. - The test will fail on a runtime that does not include all three required fixes. ## Model Used OpenAI GPT-6 in Codex assisted with code, terminal tools, and test analysis. The exact serving model suffix and context-window size were not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
83076d7e7c |
feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage. Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3f4f8b37ba |
fix: grade Codex clarification and refusal outcomes from evidence (#14570)
Grade clarification lists, obsolete unstarted wakes, and refusal cancellation from persisted evidence. Preserve execution and ownership assertions, add boundary regressions, and version the affected eval definitions. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3ca196b0a6 |
feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - An agent needs personal files across tasks and sessions. > - AGENTS.md is one file in that directory. Supporting files need the same persistence. > - The Instructions Editor and agent runs must share one current directory. > - Concurrent runs should apply only the files they change. The last sync of the same file wins. > - This pull request uses existing file transport and removes temporary copies after sync. > - Old instruction-only sessions keep their restore contract. New saves do not create revision history. ## Linked Issues or Issue Description Refs #14325. This replaces its revision-oriented design with persistent agent files. Keep #14325 unmerged. Transport prerequisite #14416 merged first at `d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master and remains below 100 changed files. Related work: #4513 and #8798 cover instruction tooling. This change handles run synchronization, cross-task personal files, browser editing, and old-session restoration. ## What Changed - Keep one current directory per company and agent. Point AGENT_HOME at a temporary working copy for each active run. Keep task files and provider HOME separate. - Restore text, binary files, and nested folders through workspace transport. Exclude remote agent files from task Git snapshots with a self-ignoring file inside the reserved runtime directory; never write through repository-controlled Git metadata. - Collect after the provider and child processes have stopped. Keep resumable conversation state. - Apply changed and deleted files under the agent lock. The last sync wins for the same file. Unrelated concurrent changes survive. - Remove temporary copies after successful sync, rejected sync, and staging failure. Register ownership before copying so restart recovery can remove interrupted preparation. Retry transient synchronization up to three times. Preserve the original remote lease reference until deletion succeeds; restart cleanup never acquires a replacement sandbox. Do not create captured directories or a conflict-review queue for new runs. - Keep browser editing, stale-draft protection, and streaming binary downloads. Keep the instruction entry and text editor limited to 1 MiB. - Keep historical agent-folder sync failures on their affected runs instead of repeating them above current saved instructions. Preserve legacy candidate review and current browser-save errors. Avoid duplicate quota warnings while retaining separate sync failures when they describe a different problem. - Require target-scoped caller grants for peer instruction access, while preserving self edits, responsible-user checks, and protected-change consent. - Treat full storage as a nonblocking run warning, never an agent pause or run-admission failure. Restore already-over-quota saved folders so ordinary agent cleanup can recover; warn on each run until cleanup. The run detail view shows the warning. - Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash large files as streams. Check editor-save quotas with metadata instead of hashing unrelated files. - Preserve old native inputs, instruction-only copies, paths, digests, and pending legacy candidates. Adopt old revision heads once. New writes do not append history rows. - Add idempotent migration 0287 and verify upgrades from the preview tables and receipts. - Add nine interactive stories under **Agents / Persistent files**, including automatic incoming edits, stale browser drafts, and storage-limit diagnostics. ## Verification - Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after merging current master and the landed transport prerequisite. Integration required no manual conflict resolution; the feature remains 99 changed files. Full workspace typecheck, production build, token gates, and 715 focused tests passed on this merge candidate. Fresh Greptile review is 5/5 with no unresolved findings. All 55 checks passed, with four conditional skips, including the build, typecheck, browser E2E, and canary dry run. A single retry recovered four jobs interrupted by runner shutdowns; no source changes were required. - Historical-warning UI fix: all 6,834 UI tests across 640 files passed, including regression coverage for three old failures, legacy preserved edits, and warnings scoped to the affected run. Full workspace typecheck, production build, Storybook build, and token gates passed. Browser-verified Storybook playtests passed for Historical Failures After Successful Save, Storage Limit, and Full Storage Run Warning. - Review follow-ups at `4e20c9fb2`: all 18 focused tests passed, including external Git directories, linked worktrees, symlinks, hardlinks, and distinct I/O failures alongside storage warnings. Server and UI typechecks, token gates, and the production build passed. - Storage warning regressions at `0724f3012`: all 33 directory tests and all five heartbeat-list tests passed, with no skips in their successful runs. They cover repeated runs while full, an already-over-quota saved folder, cleanup, warnings retained after unrelated save failures, and bounded warnings in large result JSON. Server typecheck passed after the final warning fixes. - Full workspace typecheck, production build, and token gates passed during this follow-up. Product E2E harness: 631 tests passed across 52 files; harness typecheck passed. Earlier native session/context and directory/legacy collection suites passed 537 tests; Runner unit/transport suites passed 329 tests. - **Real E2E at `0724f3012` (before this follow-up):** legacy local Codex and native Daytona Codex each passed six tasks, one server restart, seven independent assertions, and cleanup verification. Both prove browser-to-agent edits, agent-to-browser edits, nested/binary restoration, per-file last-sync-wins, a successful run after an oversized save rejection, and cleanup clearing the warning. - Native local Codex also passed the six-task quota flow before the final warning-retention fixes. That pass began at `918d1ed02` while the bounded-result warning fix was being edited, so it is not claimed as exact-final-head evidence. Its final-head rerun failed during embedded PostgreSQL bootstrap before any provider run: the macOS host had 87,365 of 87,381 SysV semaphores occupied. No unrelated services or kernel limits were changed. - The final-source report intentionally records **2/3 cells passed**, preserving the blocked native-local attempt: `tests/runner-e2e/results/agent-files-quota-final-20260928-report/`. Earlier failed attempts and provenance notes remain under `tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and the original campaign directories. - Daytona used immutable image `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf` and its exact Linux runner binary. Controller source is `0724f3012`; image source is recorded separately. - Legacy-session compatibility and all three ACP Stop/resume browser regressions passed on the prior validated feature head `169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same provider session is retained and interrupted writes are not replayed. Migration upgrade tests also passed earlier. - Nine interactive stories are under **Agents / Persistent files**, including **Full Storage Run Warning**. Its playtest and visual browser inspection passed; the warning states that runs continue and the editor remains available. - Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs skipped, no failures or pending checks. All eight browser E2E shards and their aggregate passed. Fresh Greptile review is 5/5 with no findings; all review threads are resolved, the security scan passed, and GitHub reports no merge conflicts. - The broad local follow-up test run was interrupted after host semaphore exhaustion affected isolated PostgreSQL instances. It also encountered the existing macOS long-path fixture failure and two timeout failures. This is not a claim that the full local suite passed. Logs are retained; focused storage/warning tests passed. ## Risks - A later sync can overwrite an earlier edit to the same file, including a saved browser edit. There is no text merge or retained version. This is the intended last-sync-wins policy. - A save that exceeds a storage limit is rejected and its temporary copy is discarded. The run itself continues normally, and later runs restore the last saved files with a warning until cleanup. Transient sync failures get bounded retries. An I/O failure partway through a sync can leave some files updated; a failed receipt does not claim whole-folder success. - Larger folders increase copy time, network traffic, and temporary disk usage. Active runs still need working copies. Terminal runs do not accumulate archives. Operators must provision disk for agents and configured concurrency; these limits are not company-wide quotas. - A restored old native session remains instruction-only until a fresh session starts. Its original conflict fence and existing pending candidates remain compatible. - Provider processes close at the collection boundary. Conversation resume remains available, but warm process reuse is lost. - Backups must include the instance filesystem and database. External bundles keep their existing behavior until explicitly moved to managed storage. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and browser tools assisted this change. Real provider E2E uses `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing> |
||
|
|
992f720262 |
fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task descriptions, comments, continuation data, skills, and execution rules enter several agent adapters. > - The same source can be rendered by more than one automatic input carrier. > - Failed resumes can also rebuild input from stale or compact context. > - This pull request gives each Paperclip-owned source one delivery owner and preserves the required transport boundaries. > - It adds deterministic adapter, interaction, runner, and browser tests for these boundaries. > - The benefit is more predictable context delivery with explicit evidence for later live qualification. ## Linked Issues or Issue Description Related: #13144 removes a duplicate environment payload and bounds wake lists. Related: #11360 addresses Hermes resume behavior. This pull request preserves compatible active-session formats while repairing context ownership and stale question creation. **What happened?** Task descriptions and comments could enter more than one automatic context block. Native transports could wrap a complete model input in a second task envelope. Some legacy and gateway adapters could omit the owned assignment on ordinary tasks or rebuild a failed resume with stale compact context. A continuation could also request a question after newer human comments had arrived. **Expected behavior** Each task or comment source has one automatic model-facing owner. Distinct comment IDs and repeated wording remain distinct. Fresh fallback attempts rebuild the required full context. A question request is rejected when newer queued human direction makes it stale. Harness access policy remains owned by execution configuration. **Steps to reproduce** 1. Build a task with a description and current comments. 2. Capture the actual adapter or runner input. 3. Compare source ownership and task-envelope nesting. 4. Queue a human comment before a continuation requests a question. 5. Trigger a failed resume and inspect the fresh retry input. 6. Run the focused adapter, interaction, runner, and browser checks. ## What Changed - Add shared prompt-section selection at the provider-attempt boundary. - Deliver owned assignment context through native, legacy CLI, ACP, gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and Hermes paths. - Rebuild full or compact context after resume recovery changes the attempt. Add native and Claude ACP tests of actual recovery requests. - Preserve custom templates, loaded instruction files, execution policies, and older active-session formats. - Record continuation source metadata and reject stale question creation under the issue-row lock. - Add explicit Product E2E context-integrity profiles, prerequisite gates, credential-isolation checks, and report fixtures. - Bypass service-worker forwarding for same-origin Vite development modules. A real Chromium test fails with resource exhaustion before the repair and passes after it. Production asset caching keeps its existing policy. - Add browser diagnostics and service-worker module-loading regressions. - Add an explicit zero-retry eval option. The default retry behavior remains unchanged. Each campaign records its effective policy. - Remove the model-facing working-directory sentence from four prompt builders. Existing workspace, sandbox, permission, and custom-template configuration remains unchanged. - Align the everyday workflow assertion with the current 47-entry catalog. Compared with current upstream master, the branch carries the context-ownership implementation and its tests, the explicit context-integrity catalog and evidence harness, and the focused browser regression checks. ## Verification **Merge assessment:** focused regression evidence supports merge. This is not full completion of the original broad qualification matrix. The maintainer has authorized merge after fresh verification of the master integration. - Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`. All 14 conflicts are resolved. Cancellation checks, workspace finalization, native Grok support, and both sets of tests are retained. - Current-head Greptile: **5/5**, with no blocking findings. The review names this exact commit. All **59 reported checks are terminal: 55 successful, 4 skipped, zero pending or failing**. This includes the full root general and serialized suites, separate runner checks, typecheck, build, canary, browser E2E, Docker, and security checks. The successful legacy security status is included in that total. - After integration: workspace typecheck and full build passed. Separate runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust tests, and 39 preparation checks**. Other passing checks include 621 Product E2E harness units, 376 focused shared/adapter tests, 160 real-database/API tests, 86 Hermes tests, 18 browser-support checks, and Product E2E typechecking. The complete root suite passed in CI. The duplicate local monolithic root run was stopped after that CI result; it is not counted as a completed local pass. - New native recovery coverage retains full assignment, completion contract, and explicit skill selection after safe replacement, for old and prepared input formats. Full native session test file: **136/136 passed**. - New Claude ACP coverage captures actual fresh, resumed, and missing-session fallback requests. It verifies one assignment copy, comment order, identical text under distinct comment IDs, and full fallback context. Full file: **33/33 passed**. Both affected TypeScript checks passed. - Existing deterministic tests cover source revisions, approval and trust boundaries, completion validation, custom templates, compatible sessions, standalone driver wrapping, and maintained adapter transport requests. - Provider-free browser support: **17/17 passed** after the master merge. Service-worker unit tests: **33/33 passed**. The module-overload regression failed before the repair and passed after it in real Chromium. ### Fresh live comparisons The new batch ran exactly four Product E2E attempts. **All four passed on the first attempt; no retries.** Each has six terminal matchers plus the existing browser lifecycle and invariant checks. | Exact case ID | Control | Candidate | |---|---|---| | `core-compatibility.runner-codex.local.plan-revise-accept` | Passed | Passed | | `local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume` | Passed | Passed | The plan case checks a revised canonical plan and revision-bound approval before completion. The question case restarts the server before submitting the answer, then verifies the continuation completes. Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical frozen definitions and provider versions: Codex `0.156.0` with `gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with `claude-sonnet-5`. The September 24 head added master browser recovery and test-only changes. The September 28 head also integrates newer master changes, including cancellation, workspace finalization, and native Grok. These are frozen-source live results, not exact-head live runs. The candidate received one description copy where the control initially received three. The submitted initial plan envelopes were 7,969 versus 19,097 characters. Question envelopes were 7,592 versus 18,919. These are structural measurements, not whole-provider token or dollar savings. ### Earlier evidence and failed attempts - The preceding fresh batch has four effective passing pairs: OpenCode comment continuation and assigned skill, native Codex comment continuation, and native Claude comment continuation. It retains **11 attempts: eight passed and three failed**. - Original failures remain recorded: missing local PostgreSQL library links before task creation; host-sleep cleanup after task/page checks passed; and a Claude **control** session-open rejection before a model turn. Setup was repaired identically on both worktrees. The permitted unchanged infrastructure retries passed. The underlying Claude provider startup error was not retained and remains unknown. - Older R2 retains **17 passes and one failure** across 18 attempts, including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode blank-page failure led to the service-worker repair. R2 is historical evidence: master changed the native fixed prompt and removed duplicate wake environment data afterward. - The September 24 CI run initially failed one unrelated preview readiness test (`ECONNREFUSED` on its local fixture). Its test and production code match master. Isolated local verification passed **28 tests, 3 skipped**. One unchanged CI retry passed the full shard: **831 passed, 1 skipped**, including all **31 preview-exposure tests**. The aggregate CI gate passed afterward. The precise startup cause remains unknown; a port race is a hypothesis, not a proved cause. ### Limits The original wider profile/workflow matrix, repeated trials, and remote Daytona qualification are incomplete. These results support a focused merge recommendation, not statistical equivalence or universal harness qualification. Some usage receipts are missing in both variants, so no token or dollar savings are claimed. The $500 ceiling was preserved using conservative allowances; failed attempts and unknown charges remain in the ledger. Reproduce the focused additions with `pnpm exec vitest run packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/native-session-runtime.test.ts`. Full checks use `pnpm -r typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner checks. Paid evals require the frozen definitions, profiles, and credentials; do not use `--all` as a substitute for the selected cases. ## Risks - Context placement changes can affect model behavior. Deterministic checks cover the selected paths, but live qualification remains incomplete. - The stale-question guard can reject a request when queued human comments arrived during the run. This is intended. - New stored inputs and model envelopes retain compatibility readers for older active sessions. - Custom templates may intentionally repeat content. - Removing a model-facing working-directory sentence does not change filesystem, command, sandbox, or permission configuration. - The worker bypass applies only to same-origin development module paths. Cache-policy tests preserve private-response handling and production asset caching. Mounted HTTP fixture changes remain test-only. - This PR does not claim measured token savings or statistical equivalence across every harness. ## Model Used OpenAI Codex, exact model gpt-6-astra, with repository tools and code execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The serving context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR using the required issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused local checks and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect these changes - [x] I have considered and documented risks above - [x] All current-head Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups for the current head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1a394bd30 |
feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3447609d22 |
fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.
## Linked Issues or Issue Description
**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.
**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.
**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.
Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.
## What Changed
- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.
## Verification
- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.
## Risks
- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.
## Model Used
OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
cea8dda472 |
test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users can delegate work through onboarding and Agent Chat. > - A completed task does not prove that its result reached the original conversation. > - Existing tests do not isolate completion after the source chat becomes idle. > - This pull request adds four explicit native-runner probes across Claude and Codex. > - The probes preserve the result and reply so we can separate delivery failures from inaccurate answers. ## Linked Issues or Issue Description Refs #13775. Refs #13813. These evals extend native-runner qualification. They measure completion updates before we choose a product change. ## What Changed - Add the opt-in `completion-updates` suite with two stories for each native provider. - Test completion in the existing onboarding task flow and after an Agent Chat handoff becomes idle. - Gate the chat worker on a brief inside its managed project workspace. Prove the source is idle before releasing the worker. - Check durable task completion, saved output, a subsequent source reply, and rendered access to the result. - Preserve replies, task state, screenshots, run events, and a separate semantic review rubric. - Add grader regression tests and update the documented eval contract. - Preserve the suites added on master and include four completion cases in the 306-cell catalog. Production behavior and prompts are unchanged. ## Verification - Passed all 565 eval support tests across 45 files after merging current master: `node node_modules/vitest/vitest.mjs run --config tests/runner-e2e/vitest.config.ts`. - Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p tests/runner-e2e/tsconfig.json`. - Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs tests/runner-e2e/launch.ts --list --suite completion-updates`. - Four-cell behavior campaign on source `ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`: https://github.com/paperclipai/paperclip/actions/runs/36072337485. - A screenshot-only follow-up waits for the restored source reply to render after result-link navigation. Its one-cell Claude onboarding verification passed on final head: https://github.com/paperclipai/paperclip/actions/runs/36075716141. Corrected report: https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/. The original four-cell onboarding screenshots caught navigation loading; its saved reply evidence remains valid. The follow-up again found stale wording: "That work will run next" was posted 38 seconds after the child was Done. The four-cell campaign keeps its original source and measurements. - Published evidence: https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/. - Suite definition: `afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`, version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local execution, one attempt per cell. All four cleanup checks passed. Onboarding billing coverage is partial; reported zero cost must not be read as a free run. | Story | Automated delivery/access | Separate semantic review | | --- | --- | --- | | Codex onboarding | Pass | Pass: accurate completion reply with an accessible result | | Claude onboarding | Pass | Fail: reply says it will save the note once the task runs, after the note is already saved and the task is Done | | Codex idle chat handoff | Fail | Worker completed and saved the note; no completion reply during the full observation window | | Claude idle chat handoff | Fail | Worker completed and saved the note; no completion reply during the full observation window | Both chat cases positively recorded the source waiting and the worker at the brief gate before release. Both saved outputs include the brief-only start time. The opt-in campaign is red because it exposes current behavior. It is not a required merge gate. The PR does not fix that product behavior. Semantic review is a recorded human/agent assessment of retained evidence; it is not an automated prose-quality judge. - Second campaign: https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex chat reached the idle boundary and completed its task, then received no completion reply during the full window. Claude onboarding again returned a stale handoff answer. Claude chat exceeded the prior 110-second handoff setup budget; this revision raises that bounded setup window to 180 seconds. - Retained baseline: https://github.com/paperclipai/paperclip/actions/runs/36069427676. Onboarding passed delivery/access for both providers, but Claude gave a stale handoff answer. Chat cases stopped at fixture problems; they do not establish a completion-delivery failure. This revision fixes the workspace path and competing reference requirements. - On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR checks passed and two were skipped, including typecheck, tests, and build. Broad checks ran in CI, not locally. That head received Greptile 5/5 with no unresolved findings. The unchanged mobile repository-settings browser test passed on one targeted retry after a detached/disabled Save-button timeout. - Merged current master in `9b4491e1f` and resolved the catalog-count conflict. Eval support tests and eval TypeScript pass locally. All individual CI jobs passed on this merge commit, including build, typecheck, server tests, runner checks, and browser shards. The final aggregate check also passed: 54 checks passed and two were skipped. Greptile reviewed this exact commit at 5/5 with no unresolved findings. ## Risks - These explicit probes can expose current product failures. They do not change the default paid test selection. - Mechanical delivery and result access do not establish answer accuracy. The preserved reply still requires semantic review. - A fixture failure before the idle boundary or worker completion cannot establish a completion-update failure. - The handoff setup window lasts three minutes. The worker brief wait is bounded at four minutes. The observation window lasts two minutes after worker completion. It retains later replies without erasing earlier accessible delivery. ## Model Used OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository inspection, code execution, and GitHub tool use. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f2ed0b65c4 |
fix(runner): enable API tools by default (#14186)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner exposes tools for company tasks.
> - API search and call tools cover operations without a dedicated tool.
> - The current default hides these tools unless an operator sets an
environment variable.
> - This pull request enables the tools when that variable is absent.
> - Operators can still disable the tools or restrict them to selected
companies.
## Linked Issues or Issue Description
Refs #13003, which added the guarded API tools.
**What happened?**
The native runner does not advertise `search_api` or `call_api` with the
default server configuration.
**Expected behavior**
The tools are available without a special environment variable. Existing
authorization checks still apply.
**Steps to reproduce**
Remove `PAPERCLIP_RUNNER_API_TOOLS_ENABLED` and
`PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS`. Create a normal runner
authority. Inspect its tool definitions.
**Paperclip version or commit**
|
||
|
|
96bf004a79 |
fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |