Commit Graph
14 Commits
Author SHA1 Message Date
DottaandPaperclip caf120105c test: prepare neutral native connection guidance evals (#15407)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need to discover connections, obtain consent, and continue
from saved decisions.
> - We want to reduce repeated instructions only when measured behavior
supports the change.
> - The existing decline tasks tell the model not to retry. One
provider-decline check can pass without an explanation or an observed
service counter.
> - This PR adds neutral tasks and stricter saved-evidence checks before
any connection instruction reduction.
> - Production instructions remain unchanged. The new cells are
configured, not live-qualified.

## Linked Issues or Issue Description

Refs #15218. Refs #15389.

**What existing behavior does this improve?**

The Product E2E connection workflow evaluation and its instruction
measurement provenance.

**Current behavior**

Some decline prompts supply the policy they intend to test. The
provider-decline workflow does not require a saved post-decision
explanation. Its old no-call check can use a missing fixture counter as
zero. OpenCode has no connection cases in the original Everyday matrix.

**Proposed behavior**

Add an explicit-only suite with five connection stories on native Codex,
ACPX Claude, and OpenCode. Require an explanation attributed by exact
run ID after a saved decline. Observe the provider fixture counter.
Preserve the original cases and grades.

## What Changed

- Add fifteen configured cells with one attempt, twelve-minute
deadlines, and verified 1,000-cent company and agent budget stops.
- Remove procedure hints from the three new decline prompts. Keep a
user-permitted explanation fallback and the existing positive controls.
- Require saved decline state, one decision, unchanged connections,
observed zero service calls where applicable, and a post-decision
explanation from a successful run on the same task.
- Add negative grader calibration and test the actual fixture budget
payloads. Exclude the suite from default and generic selection.
- Extend the existing full-catalog measurement source manifest with
connection descriptions and schemas. Add an audit of fixed text, tool
descriptions, returned instructions, and unqualified behavior.
- Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and
preserve its new Cursor suites. No production, credential, workflow, or
lockfile change.

## Verification

- Before rebase: Product E2E support passed 1,424 TypeScript tests and
128 Node checks. Six catalog measurement tests, repository
typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery
passed.
- The full pre-rebase repository test run was stopped when master
advanced. Its partial result is not a pass.
- After rebase and the review correction: repository build/typecheck,
Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128
Node checks, six measurement tests, and exact fifteen-cell discovery
pass. The duplicate local full-suite run was stopped incomplete after
about 20 minutes once complete CI passed; no local full-suite pass is
claimed.
- Review found that the initial grader read `runId` instead of public
`createdByRunId`. A regression calibration reproduced both rejection of
valid public comments and acceptance of the wrong alias. The fix uses
the actual field and binds the evidence type to the shared
`IssueComment` contract. A subsequent type-only import path correction
passes Product E2E typecheck.
- Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes
[complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545):
51 successful checks and two intentional Storybook skips, plus separate
Snyk success. Fresh Greptile review is 5/5 with the single review thread
resolved and no new findings. The PR is clean and mergeable.
- Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm
test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm
test:e2e:runner -- --list --suite native-connection-guidance`. The
measurement uses
`PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json
pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/native-procedure-measurement.test.ts`. Validation
used pinned pnpm 9.15.4.
- No paid provider campaign was started. There is no baseline/candidate
behavior result for these new cells.
- The audit records 654 UTF-8 bytes of fixed connection guidance. A
clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41
supplied tools, 53,341 normalized bytes at start/resume, 50,949 at
compact continuation, and a 48,195-byte authenticated OpenCode MCP
catalog. These are byte counts, not tokens, bills, vendor-private prompt
sizes, or savings from this PR.

## Risks

- This is eval preparation. Passing support tests do not establish live
model behavior or qualify an instruction reduction.
- The explanation oracle checks attributed saved output. It does not
prove cognition or arbitrary prose truthfulness. One saved interaction
also does not prove the absence of repeated idempotent tool calls.
- Successful new authentication and tool refresh, existing-connection
agent grants, independent work while waiting, explicit retry after
decline, and blocking when mandatory work remains still need separate
coverage.
- Notion setup decline does not execute a real Notion service. Positive
service approval uses an already installed deterministic service; it
does not qualify new connection creation.
- The original historical failures remain unchanged. Future comparisons
must freeze source, fixture, model, input, and grading controls and
retain every actual attempt.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code
execution. The exact serving model ID and context-window size were not
exposed in this session; they are not inferred. No model provider was
invoked by the eval suite in this PR preparation.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 21:28:16 -05:00
DottaandPaperclip a6306ba606 feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native Runner keeps provider sessions under company authority,
approvals, budgets and durable recovery.
> - Cursor work was spread across candidate branches. The published
branch lacked later plan, permission and cleanup fixes.
> - Production also needs public installation and matching runtime
assets for local and Daytona execution.
> - This pull request consolidates Cursor onto current mainline recovery
behavior and completes that installation path.
> - The installed v11 release passed focused local and Daytona
qualification after the generic mode and lifecycle cleanup. The later
model-selection correction and current mainline merge produce v14
artifacts that need matching release qualification.
> - Cursor admission is enabled in source; publish only an artifact
combination with matching qualification. Native AskQuestion and complete
per-run dollar accounting remain excluded.

## Linked Issues or Issue Description

Refs: #14435, #14631, #14669, #14699, #14724.

This completes the Cursor implementation by @cryppadotta from combined
source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer
mainline recovery, completion and warm-directory behavior. Pi and
Copilot remain gated.

## What Changed

- Generate named Rust and TypeScript ACPX release profiles from one
manifest. Share runtime pins with packaging and server verification.
Preserve vendor runtime versions; bind the updated ACPX patch to Cursor
profile v14 and reject stale generated declarations at build/typecheck.
- Remove ACPX model allowlists, including the former Codex and Pi
restrictions and the duplicate developer test-drive gate. Send any
explicit model ID unchanged to its provider and verify the effective
selection before prompting. The bundled ACPX package forwards unlisted
IDs, rejects mismatched acknowledgements, and restores the exact
selection after session load. It does not expand Cursor model aliases.
Provider rejection, mismatch, or missing model controls fails without a
fallback. Model examples live in evaluation fixtures, outside runtime
declarations.

- Add pinned Cursor execution, contained instructions, exact model
verification and Agent/Plan/Ask modes.
- Carry an opaque generic `mode` identifier in shared native execution,
sidecar, Rust and recovery contracts. The provider adapter owns
supported modes, defaults, native translation and acknowledgement.
- Keep native RPC recognition, accepted-plan interpretation and
permission evidence behind provider adapters. Shared settlement and
recovery verify normalized facts and their committed evidence.
- Replace the Cursor-only warm-attachment branch with a runner-owned
capability. Only Cursor opts into it. Move profile compatibility and
optional usage parsing into provider metadata and adapters.
- Write generic plan-wait receipts. Read exact historical Cursor
receipts through a separate compatibility decoder. Reject mixed formats
and preserve existing authority checks.
- Carry native plans, semantic questions, todos, child activity,
permission identities and partial usage diagnostics through the Runner.
- Preserve durable response delivery, cancellation, warm ownership and
process retirement.
- Finish accepted planning runs successfully. Keep their tasks open for
explicit direction. Acceptance does not start implementation.
- Ship `paperclipai runtime setup cursor` and its provisioner through
the public package. npm installation does not download Cursor. Setup
uses the OS account's closure-keyed cache so system-wide npm packages
can remain read-only. Run it as the Paperclip service account.
- Include Cursor in normal provider packs and Daytona images for macOS
ARM64/x64 and Linux x64.
- Reject stale release packs by source revision and current ACPX/Cursor
pins before assembly writes files. Verify current Cursor
version/profile/closure again at runtime.
- Ship all three daemon targets and the expected Linux image-pack
identity. A macOS controller uses its packaged Linux daemon for Daytona.
Image mismatches fail before provider launch.
- Use the vendored Runner boundary for installed readiness probes.
Verify the actual installed Cursor probe.
- Verify compiled public Daytona plugins and their release versions in
installed smokes.
- Record exact artifacts, the acceptance matrix, retained failures,
supported capabilities and rollback behavior in the [readiness
report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md).

## Verification

- Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges
mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor
plan and cancellation guards alongside mainline historical-question
filtering. The evaluation catalog includes both Cursor and expanded
adapter accounting cases (683 total). Recursive typecheck, full build,
696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck
passed. Current-head CI passed: 56 successful checks, one neutral and
four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535).
The fresh Base Greptile review is 5/5 on this exact head, with 304 files
reviewed, zero new comments and zero unresolved threads. The user
authorized overriding the CODEOWNER review gate after checks passed; no
failing checks are overridden. Prior results below retain their own head
identities.
- Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the
post-merge Apex finding. Automatic-review and new-evidence
reconciliation preserve pending child results and recheck delivery under
the status lock before completing. Account repair now excludes unrelated
secret consumers and requires the failed agent's identity. Regression
coverage includes the commit race, delivery statuses,
current-run/current-intent exclusions, repeated reconciliation, both
database reconciliation paths, and credential consumer boundaries. All
184 affected tests, server typecheck and server build passed.
Current-head Base Greptile review is 5/5, with 304 files reviewed, zero
new comments and zero unresolved threads. Current-head CI passed: 56
successful checks, one neutral and four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724).
This Base review is distinct from the earlier Apex review.
- Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves
both accepted-plan waits and pending-child-completion checks, current
provider selectors, task-creation response identities, and mainline ACPX
missing-file handling. The combined patch is bound to Cursor profile
v14; historical records keep their original identities.
- Merge head `5957c257a` passed recursive typecheck, full build, 43
installed ACPX/package contracts, 107 provider UI and plan/recovery
tests, 593 database-backed lifecycle tests, 49 profile/native contract
tests, 45 Product E2E fixture tests, fixture typecheck, token gates,
three provider-free browser task-creation cases, and Runner
conformance/replay checks. Its complete CI passed (55 successful checks,
one neutral and four skipped), while Apex returned 2/5 with a
child-delivery finding addressed below.
- The local full-suite attempt again failed the unchanged Git streaming
test (360-second timeout) and was stopped. The concurrent local Rust
attempt failed four unchanged Codex process/deadline tests; all four
passed serially without code changes in 7.29 seconds after removing the
competing test load. These failed commands are retained and are not
reported as full-suite passes; the fresh Linux CI runs are tracked
separately.
- The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned
Apex 5/5 with zero comments after fixing all three findings: per-user
install cache, stale release-pack rejection, and public Linux smoke
account/home handling. Its real built installer passed from read-only
public packages on macOS ARM64 and Linux x64. All 137 release-registry
checks and 64 ACPX package contracts passed. That review does not cover
this mainline reconciliation.
- Prior `beadd3654` passed the full CI matrix; its one unchanged chat
test failure and successful single retry remain in the [CI
history](https://github.com/paperclipai/paperclip/actions/runs/37521449327).
Historical results below remain attributed to their original builds.

- Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes
provider-specific model choices from generic offline ACPX tests. The
fake sidecar preserves the model and session identity selected at open
through suspension. Affected verification passed: 106 Rust tests and 73
TypeScript tests. This commit changes test code only; the
production-code checks below retain their recorded identities. Its CI
and Greptile review later passed; those results belong to that
historical head.
- Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`:
252 focused Runner tests passed (six platform skips), covering all six
ACPX agents, native model acknowledgement, rejected selections,
installation integrity and recovery identity. The merged branch passed
recursive typecheck, full build, token gates, server admission (19
tests), and the Product E2E catalog (45 tests). The acceptance catalog
passed all four tests. The full Rust suite passed: 643 tests, 2 ignored.
It verifies sidecar acknowledgement of unlisted models and rejection of
model mismatches. The final commits only update Rust tests; production
sources match the verified build at
`65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls
were made.
- The merge preserves both Cursor and the new mainline public-MCP
fixture cases. Auto-merge remains disabled; the latest follow-up status
is recorded above. The local `pnpm test:run` attempt hit the unchanged
Git streaming test's 300-second timeout and was interrupted before
merging mainline. The broad Runner attempt found obsolete single-model
assertions plus three macOS fixture-path failures caused by a
`/private/tmp` override. The assertions are corrected; affected
TypeScript checks passed with the standard macOS temporary directory,
and the complete Rust suite passed. Neither interrupted command is a
full-suite pass.
- Earlier declaration-cleanup head `6f4a5e9e2` passed recursive
typecheck, build, Rust and focused tests. Its CI later exposed a test
expecting duplicated Grok digest literals. The current source fixes that
assertion to compare launcher bytes with the shared manifest. Historical
successes and failed attempts are retained; no new live provider
qualification is claimed.
- Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed
complete CI (56 successful checks, one neutral, four skipped) and
Greptile 5/5. [Historical complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305).
Those results are not claimed for the cleanup head.
- Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`.
Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The
declaration cleanup preserves release pins and does not relabel that
tested artifact as a build of the new source. Mainline through
`e34abee670` was reconciled while preserving accepted-plan waits,
provider-capacity handling, and both Cursor and public-MCP fixtures.
- Clean normal installation, explicit Cursor setup and daemon resolution
passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm
lifecycle hooks ran without silently downloading Cursor.
- Historical v11 live matrix: **18/18 passed with cleanup** (nine local,
nine Daytona) after the generic mode and lifecycle cleanup. The campaign
has 23 attempts; all five failures and their diagnoses remain recorded.
Exact case identities, hashes and limits are in the readiness report.
All provider calls are real, use the explicit Luna model and
company-bound credentials, and run without qualification or
runtime-asset overrides.
- The immutable Daytona image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`.
The public Daytona plugin is installed independently and its version is
checked.
- Recursive typecheck, full build, token gates and Runner
contract/conformance/replay checks passed on the frozen application. Its
complete Linux CI suite passed. The duplicate local full-suite command
was incomplete after timing failures; affected repeats passed, but that
command is not reported as a clean pass.
- Qualification fixtures passed typecheck, 1,675 Vitest tests (one
skip), 128 Node checks, three provider-free browser tests, and 150
focused lifecycle tests after the final diagnostic correction. The
affected legacy Cursor command file also passed all five tests after
removing its shorter 10-second override; it now inherits the suite’s
standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no
unresolved review threads. CI results above are recorded separately from
historical build results.

## Risks

- Cursor v14 includes the updated ACPX dependency patch and release
identity. The v11 live matrix and image below remain historical
evidence. They do not certify new v14 package/image artifacts.

- ACPX accepts models beyond the qualification fixtures. Availability
and entitlement depend on the provider. Successful configuration is not
a claim of live qualification for every model.
- Shared mode is an opaque identifier. Provider adapters own its
meaning. Incompatible historical sessions remain fenced; exact committed
plan waits and task history remain inspectable.
- Native AskQuestion is excluded. Paperclip semantic questions are
supported. Authoritative per-run dollar accounting is unavailable;
partial counters remain diagnostics and unknown cost is not zero.
- Image input, detailed native diffs, deeper child transcripts and
native plan-file export remain follow-ups.
- macOS x64 has clean-install and daemon-startup proof under Rosetta,
not a separate live campaign on Intel hardware.
- Release only the tested package/image combination. Merging this PR
does not publish npm packages or deploy that image. Later builds need
their own release verification. Rollback disables new Cursor admission
while preserving records and recovery inspection.
- A model can fail an exact instruction: one cancelled-plan attempt
returned the wrong summary marker despite correct cancellation. The
unchanged repeat passed; both results remain in the report.

> ROADMAP.md was checked. This completes existing native Runner/Cursor
work; it does not add an independent core feature proposal.

## Model Used

OpenAI Codex, GPT-6. The exact serving variant and context window are
not exposed in this session. The agent used reasoning, repository
inspection, code execution, protocol tests and browser-backed Product
E2E tools. Cursor acceptance uses the explicit
`gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is
the evaluated provider model.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — affected suites passed;
full CI and the retained local failed attempts are recorded separately
above.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green — 56 successful checks, one
neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`;
zero new comments and no unresolved threads. The earlier Apex finding
remains fixed.
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:48:15 -05:00
DottaandPaperclip 72ff3a9f27 Measure native tool context and expand bounded workflow evals (#15218)
## Thinking Path

> - Paperclip manages persistent agents and their assigned work.
> - Agents receive both fixed instructions and tool definitions.
> - Moving a procedure into a tool description still adds model context.
> - We need to measure the complete delivered catalog and test real
outcomes.
> - The tested reductions saved little space and introduced failing
outcomes.
> - This PR keeps measurement, bounded eval coverage and the original
evidence.
> - Production prompts, tools and runtime behavior stay unchanged.

## Linked Issues or Issue Description

Refs: #15151, #14961, #14948, #14985.

This adds the measurement and eval coverage needed to assess further
native
instruction changes. The attempted hiring and dependency reduction
failed
qualification and is excluded from the final diff.

## What Changed

- Measure the actual standard-mode tool authority, including all 39
tools and their input schemas. Capture scripted native start, resume and
continuation payloads and the OpenCode MCP declaration list.
- Add OpenCode to the two explicit-only local hiring/reuse and
delegation/feedback stories. Preserve the original task requests and
independent oracles.
- Apply one attempt per selected story and explicit company and
lead-agent budget stops.
- Select managed hiring credentials from the requested profile,
including OpenRouter.
- Run the existing Node test files under Node instead of collecting them
as Vitest suites.
- Retain sanitized comparison reports, original failed grades, source
hashes and evidence gaps.

## Verification

Final source: `2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8`. Local
verification passes:

- Full repository `pnpm -r typecheck` and `pnpm build`.
- Eval-support unit tests: 1,252 Vitest checks and 128 Node checks.
- Eval TypeScript check and six final-source measurement tests. All 36
normalized
components across nine scripted deliveries and the OpenCode MCP catalog
match
the baseline exactly; the complete standing projection is 49,200 bytes.

Fresh review of this exact head is [5/5 with no remaining
findings](https://github.com/paperclipai/paperclip/pull/15218#issuecomment-5996723170);
all three review threads are resolved.
[Final-head
CI](https://github.com/paperclipai/paperclip/actions/runs/37354539126)
passes all 47 jobs, including repository typecheck, build, tests, runner
checks,
browser shards and the canary check. One initial annotation-test timeout
is
retained in attempt 1; its seven-test suite passed locally, and the
affected
CI lane passed on one targeted retry. No source or paid eval rerun was
needed.

Both ready-transition security scans passed with zero annotations. The
PR is
out of draft and conflict-free. GitHub still requires code-owner
approval for
the `package.json` test-script change; its requested reviewers are
already set.

A byte-for-byte comparison against master context
`a65ca0950834a85bb93bcc4b4042ecacdebfef53` confirms no production
changes under
`packages/` or production `server/` paths. The only server addition is a
measurement test. No new paid rerun is needed to compare unchanged
production
bytes. This does not claim that existing product defects have been
fixed.

The rejected corrected experiment had baseline **5 PASS / 1 FAIL** and
candidate
**3 PASS / 3 FAIL**, including **two newly failing pairs**. The later
readiness
experiment had candidate **3 PASS / 3 FAIL** and baseline **3 PASS / 2
FAIL / one
setup cell without a behavioral grade**. Its five comparable pairs had
two new
failures, two new passes and one unchanged pass. The missing baseline
Codex
cell never reached its provider step because Docker setup timed out.

The observed missing behaviors include parent continuation, waiting for
the
latest child revision, revised ZIP delivery and OpenCode credential
persistence.
Those failures remain failures. Source review and passing CI do not
regrade them.
The original reduction saved 460 bytes; its first repair saved only 125
bytes,
and the larger unqualified runtime repair increased the full projection.
None
of those production changes is shipped here.

Read the
[report](https://github.com/paperclipai/paperclip/blob/2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8/doc/plans/2026-10-05-native-procedure-guidance.md)
and its linked sanitized receipts for exact
sources, original campaign links, pair-level results and evidence
limitations.

## Risks

- Full-catalog bytes are not model tokens, invoices, private vendor
prompts, lazy-loading behavior or truncation proof.
- The two stories are explicit-only and do not prove general coding
quality or arbitrary resume behavior. Single trials do not establish
causation or performance trends.
- The retained OpenCode candidate credential guard failed. Cleanup
removed the original provider database, so the precise persistence
mechanism remains unknown. The guard is unchanged.
- ACPX provider-execution IDs lack proven host-call mapping. No
extra-work or feedback-consumption claim is inferred by matching names,
order or counts.
- PostgreSQL cannot start locally while the host's shared-memory slots
are exhausted. Hosted CI must supply the full database checks; the full
local database suite is not claimed green.

## Model Used

OpenAI Codex, based on GPT-6, with repository tools and code execution.
The exact deployment model ID and context-window limit are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 20:03:35 -05:00
DottaandPaperclip a386a59998 Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native agents receive task constraints and completion tools from
Paperclip.
> - Completion tools already define the procedure for reporting a
result.
> - Repeated procedure text adds instructions to each full task turn.
> - The final reply must still explain a blocker and link a saved
document.
> - This pull request removes repeated procedure text and keeps these
visible outcome requirements explicit.
> - A document receipt supplies the exact link, and stricter evals check
the persisted reply and browser navigation.

## Linked Issues or Issue Description

Refs: #14961. Related: #14948 and #15007.

**What happened?**

Native task envelopes repeat completion procedure text. A reduced
envelope needs explicit final-reply requirements. The `write_document`
receipt also lacks a canonical document link.

**Expected behavior**

Keep the completion tools as the source of procedure details. Require
one accepted completion result before the final reply. A blocked reply
must explain the reason, owner and unblock action. A document reply must
contain a working link to the saved document.

**Steps to reproduce**

1. Run the native assigned-skill document case and native blocker case.
2. Inspect the run-attributed provider final and its persisted comment.
3. Check the blocker explanation or open the final reply's document
link.

## What Changed

- Remove repeated completion procedure text from the native task
constraints and backend instructions.
- Keep explicit blocker and document-link requirements in full task
turns.
- Return a company/task-scoped `documentHref` from `write_document`.
Preserve the link in the idempotent mutation receipt.
- Repeat canonical links for this run's current saved revisions in
accepted completion feedback. Give blocked providers final-response
guidance for the cause, owner and unblock action.
- Keep internal document/comment anchors when Markdown issue links load
cached issue details.
- Add a manual six-cell comparison suite with strict source, build,
default-instruction and budget admission.
- Capture eighteen shared runnerd RPC projections and six direct
OpenCode HTTP projections across start, resume and continuation phases,
using scripted local transports and no provider execution.
- Apply v3 checks only to the manual instruction comparison; preserve v2
checks for the existing native completion suite. Check the actual
persisted blocker reason and exact saved-document link. Click the
rendered document link and check the original content marker in the
classic document card or the new document tab.
- Forward exact OpenCode finishing calls through the controller. Wait
for acceptance, keep accepted feedback and concrete rejection text, and
reject malformed responses. Preserve ordinary dynamic-tool response
handling.
- Settle the completion decision and tool response before mapping a
racing idle/error/abort event or handling explicit close/interruption.
Reject a concurrent finishing call before controller admission.
- Add a provider-free regression through real runnerd, the OpenCode
proxy and a fake provider. Reject the first completion, accept the
corrected report in the same turn, and propose one result.
- Keep all original verdicts unchanged. Treat replay under new checks as
separate diagnostics.

## Verification

- `pnpm -r typecheck` and `pnpm build` pass locally.
- Native document-authority tests pass, including company/run
authorization and idempotent replay.
- Native runtime-context, backend and measurement tests pass.
- Final-answer calibration, protocol scoring, source-admission and
catalog tests pass. Wrong reasons, absent links and wrong link targets
fail.
- `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six
single-attempt local cells with the declared models.
- Exported `prepareNativeInstructionPreflight` then
`verifyNativeInstructionPreflight` pass on this clean committed source.
They build locally and make zero provider calls.
- Corrective live confirmation is incomplete. Source 3a7349d passed both
Claude and both Codex cases. OpenCode saved the correct document but
omitted its final link; its blocker case was canceled before paid
execution. Preserve this failure. The e171282 confirmation was stopped
during build after fresh review found a completion-settlement race; it
executed zero providers. Source c3e0cb303 fixes that race. Two affected
OpenCode cases await fresh review and one bounded confirmation; earlier
results remain attributed to their original source.
- OpenCode proxy parsing, driver, factory and input tests: 81 pass
across retained focused runs, including six settlement races.
Evaluator/scoring/admission checks: 122 pass. The real proxy regression
passes. Fresh local prepare then verify passes with 18 shared and 6
direct scripted captures, fresh SDK/Rust builds and zero providers.
- The full local suite recorded two failures: a webhook timeout and a
Git-scan load count mismatch. Both files pass in isolation with
unchanged assertions/time budgets; preserve the original failure log.
Fresh c3e0cb303 CI and review are pending. This PR remains draft.

## Risks

- Final-answer wording can vary by provider. The checks cover the
declared release-access blocker and saved document fixture, not general
answer quality.
- A single trial does not establish general equivalence, cause, speed,
cost or live resume behavior.
- `documentHref` is an additive receipt field. It points to the current
saved document, not an immutable historical revision. Replaying an older
receipt does not fabricate a new link.
- The correction adds four production paths for document receipts,
accepted completion feedback and UI navigation, plus four OpenCode
controller/proxy paths, beyond the original three instruction paths.
Completion rejection must remain repairable; the production-boundary
regression covers it.
- Preserve the frozen comparison context for live measurement. A
merge-tree check against current master is clean. Do not relabel earlier
live results as results from a later source tree.

## Model Used

- OpenAI Codex, GPT-6 family. The exact serving model ID and context
window are unavailable in this session. Capabilities used: reasoning,
code editing, shell execution, test authoring and evidence review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 08:10:14 -05:00
DottaandPaperclip b17019e14d fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its adapters supply task context and access to Paperclip skills and
tools.
> - The default hire manual and shared prompts also repeat general work
procedures.
> - Those procedures overlap with stock provider instructions and the
Paperclip skill.
> - Existing E2E fixtures supply a QA manual, so they do not qualify the
production default.
> - This pull request reduces the generic instructions and adds real
default-hire coverage.
> - The benefit is less competing guidance, with inspectable evidence
for preserved skills and task context.

## Linked Issues or Issue Description

Refs: #14920. That merged change preserves native Codex base
instructions. This PR covers the default manual, shared legacy prompts,
operational skill guidance, and the narrowly approved ACP
skill-discovery/session-environment repair for measured delivery and
credential-persistence failures.

**What existing behavior does this improve?**

New non-CEO hires without a custom bundle and legacy task/chat startup
and continuation prompts.

**Current behavior**

The shipped default manual contains 602 words. Generic task/chat prompts
and ordinary resume deltas repeat work procedures already available
through the harness and Paperclip skill.

**Proposed behavior**

The default manual contains only the eight-word company identity. Shared
startup prompts retain identity and connection guidance. Ordinary resume
deltas retain current work context without the generic execution
contract.

**Reason and benefit**

Let the stock harness guide general work. Keep Paperclip-specific
capabilities and independently test default hires, skills, ordered
comments, and chat restart.

**Breaking changes**

New default hires receive less guidance. Existing saved manuals,
explicit custom bundles, CEO templates, and specialized wake contracts
retain their behavior. The obsolete includeExecutionContract option
remains accepted for source compatibility.

## What Changed

- Reduce the default hire manual to one sentence.
- Reduce shared task/chat defaults and remove the generic
ordinary-resume contract.
- Keep connection guidance, auth, skills, custom prompts, and
specialized wake context.
- Add credential-free instruction-boundary gates and 26 explicit Product
E2E cells across eight legacy/native profiles, including two focused
Paperclip-storage cases.
- Capture public hire receipts before providers run, then grade
delivered prompts and independent task/chat outcomes.
- Add an early legacy skill API recipe for saving a task document,
checking the saved revision receipt and linking the document. Improve
stock task/heartbeat skill-selection metadata and show a clickable
Markdown UI-link example. Keep native tool completion separate.
- Advertise bounded routing descriptions and exact successfully staged
SKILL.md paths in legacy ACP Claude; keep full bodies on demand and
preserve remote path rebasing.
- Remove only the provider environment from copied persisted ACP session
records, while loading current run credentials and preserving all other
options/conversation state.
- Regenerate both capability metadata inventories and reject stale
manifests/inventories before provider admission.
- Publish the original reduction and focused skill-repair comparisons,
preserving all failures, automatic recovery, cost coverage and
limitations.

## Verification

**Behavioral qualification remains pending.** Original legacy ACP Claude
loses the issue document only in the reduced cohort beneath an unchanged
credential failure. A source-backed diagnosis finds that neither
ordinary assignment reads the staged operational skill, while the
runtime persists provider environment in session state. The new common
repairs expose skill metadata/path and omit persisted env; strict
document and credential guards stay intact. [Inspectable diagnosis and
retained
hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md).

Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb`
incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged
#14961/#15007). All 222 affected adapter tests, adapter-utils/E2E
typechecks, and final 96 variant/grader/retry calibrations pass. The
frozen historical comparator is
`c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths
identical, exactly two production instruction paths plus nine declared
unit expectations differ. The operational skill/discovery/environment
repairs, selected model/profile/task/core grader/auth/permissions/retry
policy are identical. Both actual launcher prepare→verify admissions
pass with zero providers. [Immutable manifest and exact
receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json).

One original legacy ACP Claude cell per variant is authorized, with
enforced single campaign attempts, 12-minute deadlines and company/agent
1,000-cent hard stops; every product recovery run/cost is counted.
Actual live outcomes are pending. Current normal CI has one failed
server shard and failed aggregate verify under diagnosis; other normal
gates including typecheck/build/Rust/all eight browser shards pass.
Fresh review completed successfully; the valid historical startup/resume
masking finding was fixed with per-invocation task/chat checks and
strict complete-snapshot capture, calibrated and resolved. Prior heads,
failures and campaigns below remain historical evidence, not checks on
this repair head.

- Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on
merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52
current-head checks pass with two intentional Storybook skips, including
repository typecheck/test/build and the browser shard. Fresh Greptile is
5/5 with zero unresolved review threads. Exact-head stock prerequisites
pass 599 assertions (598 TypeScript + 1 Rust), all six gates and
retained receipt verification, zero providers/source errors. Fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`.
Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck
and 26-cell stock discovery pass. Canonical contract/inventory checks
and the later issue-derived reference calibration are retained; that
reference-only follow-up is not live-qualified by earlier frozen runs.
- Prior full repository typecheck/build passed. The complete local
Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three
timing failures. All three affected files passed unchanged narrow
reruns; original failures remain retained. Current-head CI now passes
the full general checks; the original local failures remain retained.
- The original 24-pair default-manual/shared-prompt comparison has two
new overall classic Claude/OpenCode document-delivery failures plus an
additional legacy ACP Claude document loss beneath an unchanged
credential-guard failure (not closed by later runs), two newly passing
OpenCode ordered cases, seven unchanged failures and 13 unchanged
passes. Equal 15/24 totals do not establish behavioral equivalence.
[Complete original
report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md).
- The skill-only repair holds the eight-word manual/shared prompts and
merged #14920 fixed. All four matched profile configurations and 203
fixture/behavior files match. Candidate
`abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill
sources against baseline `bc83fe030234439ac51279502a28803958963e2e`.
[Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547)
and [baseline
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047)
each pass 571 exact-source prerequisites before providers; all eight
cells clean up successfully. Failed campaigns publish successfully and
remain failed.
- Repair pairs: Claude original Fail → Pass; Claude explicit Pass →
Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff
worsens beneath the unchanged failing UI-link grade: baseline gives a
clickable API URL, candidate gives a code-formatted path without an
anchor. The request's usable-link wording is narrower in the UI-only
oracle. [Complete repair report and safe
projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md).
- The subsequent narrow stock metadata/link correction has two matched
Pass → Pass cases, zero new machine failures/passes and no pending
pairs. Both original-case handoff links remain deficient: candidate uses
a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the
preserved original oracle only requires a durable document. Both
explicit clickable UI-link cases pass revision/content/link grading. All
four exact-source 587-check gates, single assignment runs and cleanup
pass. This does not establish fix causality because baseline also
succeeds. [Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401)
freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched
baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374)
freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only
comparison with reduced manuals/shared prompts held constant, not a
repeat of the historical-manual comparison. Only original and clarified
explicit classic OpenCode cases are selected, two per variant/four
expected turns. 8,242 other tracked files and both profile hashes match;
protected workflows admit each exact source before credentials.
[Complete qualification
report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md).
Candidate original loads Paperclip/reference before saving publicly;
baseline original loads it after writing locally, then saves publicly
within the same assignment. Reported cost totals are $0.0107824490
candidate / $0.0107909015 baseline, with unmetered runtime. The later
reference-only issue-derived link correction is provider-free calibrated
and **not live-qualified** by these frozen runs; no further paid runs.
- Retained tool calls show the repaired original OpenCode assignment
loads only its assigned output skill before writing locally. Operational
Paperclip is first loaded during automatic disposition recovery; its
early recipe is visible then, but it never saves the missing document.
Explicit candidate loads Paperclip and reads the new reference before
saving successfully. All nine actual runs are counted. Reported LLM
totals are $0.3802537209 baseline and $0.4918990161 candidate; local
runtime is unmetered.
- Initial setup, packaging, cancelled/missing-cell recovery, callback
test and relative-output attempts remain retained. No completed provider
failure was rerun. Frozen measurement branches are unchanged by later
canonical metadata maintenance.
- Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`,
and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly
for paid execution; it is excluded from `--all`.

Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f`
replays this PR on merged hiring #14985
(`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four
explicit custom-CEO-bundle checks, minimal generic manual boundary, and
both suites. The combined fixture catalog and hiring calibrations pass
67 assertions; exact-head stock prerequisites pass 599 assertions (598
TypeScript + 1 Rust), all six gates and retained-receipt verification,
zero providers/source errors, fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E
typecheck and 26-cell stock discovery pass. Fresh current-head CI passes
all 52 checks with two intentional skips, and fresh Greptile is 5/5 with
zero unresolved review threads.

The prior source-plan browser failure is retained: a deterministic
process fixture replayed its last `fixture:plan` command on
`chat_task_completed`, writing revision 2 with identical body after the
approval handoff. This was not paid provider execution. Rebased
current-head CI passes the same assertion without an old-head retry or a
change to that browser fixture.

The merged hiring change was measured separately on immutable matched
unions, with this reduced/shared/operational context and native
completion guidance held constant. [Complete original two-profile
report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md):
[candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208)
/ [historical
baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463),
705 provider-free prerequisites each. Both pairs are unchanged Fail →
Fail on the exact-five count, with six core delivery checks passing all
four cells; 28 actual successful runs include eight automatic completion
wakes, zero retries, four successful cleanups. Source-read coverage is
uncomparable, actual model charges unknown. Separately versioned
provider-free accounting remains analytical work; original verdicts are
preserved. This does not rerun or qualify the completed default-manual
or native campaigns.

## Risks

- Legacy ACP Claude's additional delivery loss is not closed by any
later matched run and blocks the no-extra-failing-behavior merge
criterion. Legacy document delivery may have relied on the prior
manual/shared prompts. The early skill repair improves Claude in one
trial; the later OpenCode pairs pass in both variants and cannot
establish causality or robust recovery. Both original-case links remain
deficient beneath the storage-only grade. The later issue-derived
reference correction has only provider-free validation. Native
finish/block descriptions must not be supplied to legacy agents.
- The comparison holds merged native Codex fix #14920 constant; it
cannot measure that fix's before/after task performance.
- These bounded skill/context/chat workflows do not measure general
coding quality. Unrepresented providers remain unqualified.
- Saved manuals and old Codex sessions are not automatically migrated.
Codex through ACP still has a separate base-instruction follow-up.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution, and delegated PR/eval
work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites and all
three unchanged narrow reruns pass; complete-run timing failures
retained in Verification)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green on the new repair head
(prior-head checks retained above)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
on the new repair head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:32:42 -05:00
DottaandPaperclip dd868ed125 fix(runner): share native completion tool guidance (#14961)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Native Runner agents report completion through finish and block
tools.
> - The providers receive different descriptions for those tools.
> - Completion guidance belongs with the tools that enforce the result.
> - This pull request shares the descriptions and refreshes retained
catalogs.
> - A separate native suite checks completion and blocking on production
defaults.
> - Legacy agents retain their separate skill and API paths.

## Linked Issues or Issue Description

Refs: #14920, #14948, #14985.

**Current behavior**

Native Codex and MCP bridges describe finish and block differently.
Retained provider sessions can keep old descriptions.

**Proposed behavior**

Native providers receive the same finish and block descriptions. The
descriptions cover report selection, validation feedback, returned
outcomes, approval gates and final-answer timing. Retained native
sessions refresh from v13 to v14.

**Reason and benefit**

Put the completion procedure next to its native tool. Preserve stock
base instructions, schemas, permissions and terminal semantics. This PR
now stands alone on master. It contains no reduced manual, shared prompt
or operational-skill changes from #14948.

## What Changed

- Add canonical native finish and block descriptions. Use them in direct
Codex and both native MCP bridges.
- Advance the native tool contract to v14. Cover old-v13 refresh without
replacing task identity or prior history.
- Check authenticated tool catalogs, provider start/resume frames and
serialized daemon catalogs.
- Add an independent, explicit-only native completion suite. Preserve
the original assigned-skill durable-document journey. Pair it with a
concrete whole-task blocker across Codex, ACPX Claude and OpenCode.
- Verify the actual public production default bundle and budgets before
execution. Require independent durable disposition, native
result/terminal receipts and observable provider-final ordering.
- Correct the blocker browser oracle to accept the requested
explanation. Keep exact owner/action/scope checks. Calibrate positive,
missing and contradictory replies.
- Preserve only actual `tool_call` terminal names (`paperclip_finish` /
`paperclip_block`) in the native compatibility run-log projection.
Require the same named call ID through its finishing result; retain all
other redaction boundaries.
- Admit verified hosted shallow checkout/build hydration and bind the
selected runnerd to exact source/archive/binary provenance. Hosted cells
truthfully reuse the existing trusted build; local admission executes
Rust calibration. Forward only public source/run identifiers through
both launcher preflight subprocess paths.
- Enforce single attempts in the launcher for opted-in fixtures. Keep
ordinary retry policy unchanged. Run exact-source, credential-free
admission before credential loading.

## Verification

- Frozen candidate: `d6e59e4712a3158ab4cd7d58deff1389b4578c21`, based on
master `59c07ede72dc08b8aba149a01cc11e0b7a204621`; historical
descriptions: `e74ed61a69fbdd8b3a8f15dd6456bc3140246e33`. Exactly the
five original native production files and six unit tests differ. Both
carry identical corrected fixtures, strict named finishing-call grader,
closed compatibility carrier and admission. Defaults,
profiles/models/auth/permissions and manifest bytes match.
- Actual launcher `prepareNativeCompletionPreflight` →
`verifyNativeCompletionPreflight` admission passes on both exact refs
with zero providers: candidate 132 / historical 127 selected TypeScript
assertions, 128 Node calibrations and one Rust normalization calibration
each; E2E typecheck, manifest checks, selected binary provenance and
six-cell discovery pass. Each has 257 explicitly skipped unrelated
assertions, not coverage. The credential-free environment calibration
exercises both real prepare/verify subprocess options with public hosted
identifiers and rejects credential/ambient overrides. Complete actual
launcher prepare→verify also passes on both frozen refs with explicitly
synthetic hosted metadata/verified archives, separately labeled as
calibration rather than a trusted GitHub run. Exact framed provenance
parsing and mock source identity are calibrated without relaxing the
real verifier.
- [Complete matched qualification
report](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-master-qualification.md),
[immutable
manifest](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-manifest.json)
and [closed retained
audit/hashes](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-results/comparison.json)
are inspectable. All six candidate cells pass; historical descriptions
pass five. Paired outcomes: **zero new failures, one new pass (Codex
blocker), five unchanged passes, zero pending pairs**. [Candidate
campaign](https://github.com/paperclipai/paperclip/actions/runs/37098728980)
and [historical
campaign](https://github.com/paperclipai/paperclip/actions/runs/37098815696)
each execute six original attempt-1 native runs, with no campaign retry
and successful cleanup. Their trusted workflow revision is
`215586d127e97c9301d86e769a39a15c13298ca2`, separate from measured
source. [Candidate public
HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098728980-1/index.html)
and [historical public
HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098815696-1/index.html)
retain declared screenshots.
- Independent candidate evidence agrees with all original grades: 51
strict native checks, 12 served-default/budget checks and 21 original
skill/document checks pass. The historical Codex blocker saves the
correct whole-task blocker but omits the required marker from its actual
provider final and identical saved reply. This is not semantic-summary
fallback. Its original browser/matcher failure stays retained; the
additional native snapshot/grade and workspace before/after digest were
never written and are not fabricated by the separate API/PRP audit.
Historical Codex completion has one failed finish followed by success
within the same native run; the public receipt records no failure
reason. All twelve runs and their usage remain counted. Reported
model-cost subtotals are $0.00421482 historical/$0.00437391 candidate;
Codex/Claude zero entries have unknown billing type, actual invoices are
unverified and hosted execution cost is unmetered. One matched trial
supports no extra failure within these six cases, not broad statistical
or coding-quality equivalence.
- Initial hosted `e18c2cf9` / `459455ac` and subsequent `0a9c5a7` /
`00a761b` cohorts each stopped before providers in all twelve cells. The
latter failed a mocked-receipt unit test under ambient hosted metadata;
all source/build proofs passed. [All twelve later setup
receipts](https://github.com/paperclipai/paperclip/blob/402ee94c52273ad58de355ae9a7d562dd22f8101/doc/plans/2026-10-02-native-completion-qualified-hosted-setup.json)
are retained. [Exact failed setup
receipts](https://github.com/paperclipai/paperclip/blob/27653eb1a8f8ce839776d760f4563f672e5a706c/doc/plans/2026-10-02-native-completion-master-hosted-setup.json)
and the original manifest remain intact. Local sandbox-denied loopback
and stale anchor-expectation attempts are retained separately; unchanged
appropriate assertions were corrected/admitted before paid dispatch. Old
anonymous OpenCode streams are not assigned inferred tool names or
retroactively passed.
- Full provider-free E2E support previously passed 927 tests in 67
files. Exact-head d6 normal CI run `37098409915`, attempt 1 passes full
repository typecheck/build/tests, Runner Rust/static checks, all browser
shards/aggregate and canary: 52 check-runs pass, four intentional skips,
Snyk passes. Fresh Greptile check `111132956342` is 5/5 with zero
unresolved threads. Source-specific deterministic tests do not
substitute for the bounded live comparison.
- Earlier native source `9138f570c341c251a5727c32d6615ce238bc8e03` is
archived. Its [complete reduced-manual-context
report](https://github.com/paperclipai/paperclip/blob/9138f570c341c251a5727c32d6615ce238bc8e03/doc/plans/2026-10-02-native-completion-live-comparison.md)
remains intact, including original failures, grader limits and
provider-free replay. It is not current-master-context qualification.

## Risks

Changed tool text can change model behavior. The completed six-pair
qualification shows no extra failing outcomes in this bounded trial;
other tasks and repeated-run variance remain unmeasured. Observable
final ordering does not prove provider feedback consumption. Public
evidence can fail closed if a provider does not expose the required
result sequence. This slice does not remove native fixed prompts or
measure general coding quality. No database, schema, permission or
legacy completion changes occur.

## Model Used

OpenAI Codex, GPT-6 family, with code inspection, execution and tool
use. The exact deployment ID and context-window size are not exposed in
this session. They are unavailable rather than inferred from the model
menu.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 06:20:38 -05:00
DottaandPaperclip 018993140f feat: let agents name prompt-only tasks (#14761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users create tasks with a title and a description.
> - A required title adds work when the prompt already explains the
request.
> - An agent can name the task once it reads that request.
> - This pull request accepts prompt-only tasks and starts them with a
short prompt slice.
> - A scoped title tool lets the assigned agent replace that slice early
without changing execution state.
> - A live browser eval checks the real agent call, saved title, audit
entry, and preservation of user titles.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: task creation, shared contracts, database, server, runner
tools, and board UI.

**Problem or motivation**

Users must currently write a title before they can submit a detailed
task prompt. The agent has enough context to write a useful title
itself.

**Proposed solution**

Make the title optional when a description is present. Save the first
120 characters of the normalized prompt as a provisional title. Ask the
assigned agent to call `set_task_title` early. Use an atomic
provisional-title guard to preserve titles supplied or edited by users.
Keep explicit titles supported.

Related: #14543 and #14556 concern empty-title submission. This change
intentionally enables that submission when a prompt is present, instead
of requiring a title.

## What Changed

- Add the `titleNeedsGeneration` field with an idempotent migration.
Keep existing titles unchanged.
- Add `PUT /api/issues/:id/title` and the native and legacy
`set_task_title` tool. Enforce company access, active-run ownership,
shared, bounded retry receipts across native/HTTP calls, and
transactional audit logging. Refresh external-object links after commit,
with the same feature gate and plugin detectors as ordinary title edits.
- Add early naming guidance in Standard, Ask, and Plan task context.
Preserve the description, status, and assignment.
- Allow prompt-only root and child task creation, plus draft restoration
in the New Task dialog. Keep user titles supported.
- Add an opt-in Product E2E suite for prompt-only Standard and Ask
tasks, plus an explicit-title control. It checks actual provider calls
within the first five tools, persisted state, audit attribution, and the
reloaded UI.
- Preserve a closed vocabulary of API key maintenance phrases in
declared prose while rejecting opaque credential suffixes. Add one
bounded naming retry after wording is rejected, without treating the
rejected call as a saved title.
- Repair the native cleanup receipt check exposed during full
verification: accept matching input digests, retain legacy input checks,
and reject conflicting receipts.

## Verification

- Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3
passed** with native Codex `gpt-5.4-mini`, first attempts only,
automatic retries disabled. Standard and Ask each saved “Rotate expired
API key” on their first tool call, with matching persisted state and a
single same-run audit entry. The explicit-title control retained its
user title with zero title writes. All three verified the reloaded
browser UI.
- Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns
are retained separately; they exposed credential-prose handling and
prompted the naming recovery fix. No failed result was regraded or
deleted.
- Reproduce with `pnpm test:e2e:runner -- --id
task-titles.runner-codex-mini.local.prompt-title-standard --id
task-titles.runner-codex-mini.local.prompt-title-ask --id
task-titles.runner-codex-mini.local.preserve-explicit-title
--max-automatic-retries 0` and an authorized provider key.
- Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit.
The runner build used the configured external eval source tree.
- Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck
and UI token gates passed.
- Title API/native regressions cover prompt-only and explicit child
creation, user edits, ownership/company isolation, external reference
refresh, cross-surface retry replay, and the 64-key limit without
receipt eviction. All passed. Prompt-context coverage: **44 tests
passed**.
- Rust credential regressions: **35 tests passed**, including benign
maintenance qualifiers and opaque credential rejection in every declared
prose field. Catalog/report reconciliation: **28 tests passed**. Native
recovery: **560 tests passed**.
- Broad local `pnpm test:run`: **14,555 tests passed** in the general
server group; two suites failed to initialize embedded PostgreSQL and
the existing 40,000-file Git streaming stress test exceeded its
300-second macOS timeout. All three suites then passed in isolation (**5
tests passed**) without code or timeout changes. The original full local
command exited nonzero and is not being represented as a clean full run.
- Latest-head GitHub checks are green: **53 passed, 4 skipped, zero
failed or pending**, including all test shards and the canary packaging
dry run. Greptile reviewed the same commit at **5/5**, with zero
unresolved review threads.

## Risks

- The additive database field must reach the server and UI together. The
migration uses `IF NOT EXISTS` and defaults existing tasks to a final
title.
- Title generation depends on the assigned agent running. Tasks without
a run keep their provisional title.
- Live qualification covers the native Codex path in Standard and Ask
modes. API/legacy and Plan behavior have deterministic coverage.
- The credential-prose exception validates the entire suffix against a
closed maintenance vocabulary. Unknown suffixes, assignments, quoted
values, credential prefixes, and diagnostics retain strict checks.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed in
this session. The live eval uses the native Codex `gpt-5.4-mini`
profile.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 16:48:24 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00
DottaandPaperclip 9d19f98b50 fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agent chat uses native runner sessions to plan, delegate, and track
that work.
> - A user can press Stop while the native session is still starting.
> - The server can acknowledge that Stop without dispatching it, then
let the session submit a turn.
> - This leaves chat recovery waiting for an execution that the user
expected to stop.
> - This PR waits for the startup handle, dispatches cancellation, and
prevents a late startup from submitting a turn.
> - New full-stack evals check the resulting records and outputs across
Claude and Codex.
> - Those evals also exposed missing ACPX readiness fields, unbounded
polling, and an old-run identity check that rejected valid warm
handoffs.

## Linked Issues or Issue Description

**What happened?**

Stop during native startup could record an acknowledged cancellation
with `dispatched: false`. The provider could then begin work. A
subsequent `/new` stayed queued. A remote Claude follow-up also
exhausted the command journal while probing warm-session readiness: ACPX
never returned the readiness fields required by the shared transport.
Once readiness worked, attachment incorrectly compared the next run
descriptor against the old run ID. The 25 ms polling loop could issue
4,800 commands during its two-minute wait, beyond the 500-command bound.
The existing chat eval treated lifecycle logs as proof of an active
provider turn, so it did not distinguish startup cancellation from
active-turn cancellation.

**Expected behavior**

A Stop during startup must reach the pending session. A late session
must not submit a prompt after Stop. Recovery must retain control when
startup exceeds the bounded wait. Chat evals must check saved task
state, document contents, worker identity, account binding, and
duplicate effects.

**Steps to reproduce**

1. Start a native Claude or Codex chat turn.
2. Press Stop after process startup is requested but before the provider
turn starts.
3. Send `/new`, then send a fresh message.
4. On the affected base, cancellation can be acknowledged without
dispatch and the reset stays queued.

**Paperclip version or commit**

The live Claude baseline reproduced this on `29d6b3509`. The branch also
includes master commit `0f5fafe16`.

Related work: #13678, #13686, #13693, #13291, #13738. A separate runner
reliability branch also contains a startup-wait fix. Its overlap must be
reconciled before merging; this branch additionally prevents prompt
submission after a late startup.

## What Changed

- Wait for a pending native startup before acknowledging a run-scoped
Stop. Preserve the existing recovery error when that wait expires.
- Keep a Stop guard on startup. Cancel a late handle before it can
submit a provider turn.
- Add regression tests for normal handle publication and publication
after the Stop deadline.
- Back off blocked warm-attachment probes. Keep the fast two-snapshot
barrier, fail closed, and record changed blockers.
- Add red/green tests for delayed readiness, persistent blockers,
alternating readiness, and readiness near the deadline.
- Publish ACPX readiness and blockers. Preserve the old authority’s
event acknowledgement barrier; only settled sessions can proceed to
attachment.
- Bind warm ACPX descriptors to the validated next authority while
retaining old-run event correlation until activation. Preserve session
identity and provider profile checks.
- Exercise two consecutive run rotations through a qualified fake
sidecar, verifying checkpointing, provider identity, pre-activation
rejection, and new-run work admission.
- Separate startup and active-turn cancellation checkpoints in the
browser eval.
- Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells
across Claude and Codex.
- Cover hiring and reuse through managed AI accounts, source-based
review, current blocked-task status, request replay after a lost HTTP
acknowledgement, server restart continuity, and Stop/reset continuity.
- Use ordinary production agent instructions. Enable API tools only for
the two coordination cases that need them.
- Calibrate the matchers with invalid records and outputs. Require
remembered context after restart and a structured status snapshot that
distinguishes the current blocker from history and task status from
active execution. Compare the public issue mutation contract and
relationships during read-only reporting. Preserve before/after source
records in failed eval evidence.
- Fix the lost-ack browser harness and verify it against a real HTTP
server. Check the chat composer after restart instead of waiting for an
unrelated document lifecycle event.
- Document the scope and limits of each case.

## Verification

- The startup regression failed on the unfixed executor and passed after
the fix.
- `pnpm test:e2e:runner:typecheck` passed.
- `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files.
- `pnpm exec vitest run
server/src/services/native-runtime/native-session-executor.test.ts`
passed: 385 tests.
- [Baseline live
campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868):
Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed.
Codex delegation was blocked by provider capacity.
- [Eval-only startup
campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786):
both providers failed as expected. Both persisted `dispatched: false`
and left `/new` queued.
- [First fixed
campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706)
on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both
providers. Failed cases exposed eval harness defects and remote
continuity failures. All attempts remain available.
- [Original workflows and stronger memory
checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649)
on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and
startup Stop passed for both providers; Codex remote restart passed.
Claude remote restart exposed the missing readiness contract. Two Codex
planning cells hit provider capacity.
- [Unchanged-model
retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548):
Codex planning and backlog creation both passed.
- [18-cell campaign with ACPX
readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963)
on `6a98ef743`: 16/18 passed, including all local/remote Stop and
committed-send cases. Claude remote continuity exposed the
next-authority check, now fixed. Codex hiring produced its checklist,
but the runner redacted the requested marker after it appeared as
“Tracking token: …”. That content-redaction policy is unchanged and
remains an explicit limitation.
- [Structured status
grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725)
on `50448c228`: both providers passed on their first attempt, including
cleanup.
- [Complete read-only state
grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011)
on `551e13892`: both providers passed.
- [Final ACPX handoff and hiring
retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456)
on `cbd637587`: all three Claude Daytona cases passed (restart
continuity, active Stop/reset, and lost-ack replay). Codex hiring
reproduced the content-redaction failure: the saved checklist contained
`Tracking token: [REDACTED]` instead of the required business marker.
All four cases completed cleanup successfully. [Published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/).
The only subsequent commit adds the qualified-sidecar integration test;
production code is identical to this live proof.
- `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without
paid models.
- Runner TypeScript typecheck passed. All 5 warm-readiness tests pass;
two failed with the prior fixed-rate loop, and the late-readiness test
failed before the pacing correction.
- ACPX readiness and warm-identity regressions each failed before their
fixes. All 292 runner-core Rust library tests passed. The
qualified-sidecar integration test passes. Rust formatting is checked.
- Status-grader regressions for misleading historical mentions and
previously unchecked mutations each failed before tightening the oracle
and pass now.
- [Latest-head
CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307)
passed on `a4093c8f1`: full build, type checks, test partitions, browser
E2E, and native runner checks. Two unrelated tests initially failed
(Sentry fixture release attribution and local-service fixture
readiness); both passed locally together (35 passed, 5 optional SDK
tests skipped) and on the failed-job retry. No changes were made to
those tests.
- Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed
and all review threads are resolved.
- The paid live suite is not fully green: the reproducible
content-redaction case remains red. This is separate from the passing PR
merge checks. No production content-redaction, prompt, model, or
completion-policy change is included.
- Managed-account hiring and review cases explicitly enable API tools;
these do not qualify default new-user onboarding.

## Risks

- Stop can wait up to 30 seconds for startup, then use the existing
pending-recovery path. This does not prove that remote cleanup has
finished.
- Blocked warm readiness adds up to 750 ms between later probes with the
two-minute remote budget, or about 32 ms with the default five-second
budget. Ready sessions retain the short second barrier.
- Paid evals can fail because of provider capacity or agent decisions.
Each failure needs evidence-based classification.
- The HTTP request replay case checks comment idempotency and duplicate
effects. It does not prove replay safety for an ambiguous provider tool
call.
- The new suite is opt-in. It does not increase the default paid
campaign.
- No production prompts or model selection change. Review-handoff
behavior and content-redaction policy remain separate product decisions.
The latter can remove harmless business content that looks like
credential syntax; the failing attempt is retained.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 10:35:38 -05:00
DottaandPaperclip 11921075a4 Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The first task helps a new user define and approve useful work.
> - That workflow needs reusable instructions and tests against the
production experience.
> - Native Codex and Claude must load the assigned skill, including
after resume.
> - Maintainers need recorded conversations and precise failed checks to
judge regressions.
> - This pull request adds the first-task skill and a suite in the
shared Runner E2E harness.
> - It keeps behavior results separate from informational quality scores
and incomplete recordings.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The first onboarding task and the Runner E2E report used to review it.

**Current behavior**

Onboarding embeds its policy in a hidden brief. Native Codex drops the
skill-instructions setting at the Rust boundary. The shared E2E harness
has no onboarding suite or full conversation view.

**Proposed behavior**

Assign and invoke `/first-task` for the onboarding task. Send selected
Codex skills as structured protocol inputs. Run twelve scenarios across
legacy Codex, legacy Claude, native Codex, and native ACPX Claude.
Include all 48 cells in full campaigns. Show recorded chat, question and
approval cards, exact checks, instructions, and billing in the shared
dashboard.

**Reason and benefit**

Measure the real onboarding experience before changing prompts.
Distinguish infrastructure failures, behavior failures, and unexercised
journey steps.

**Breaking changes**

No database migration or production API change. First-task instructions
now live in an assigned skill. The user-edited persona is preserved; the
skill includes the maintainer-approved proposal-mode mapping and
saved-plan requirement.

Related: #11043 is earlier onboarding work. #13422 already fixes native
Claude model pinning, context delivery, and read permissions on master;
this branch includes those fixes through its base. The new Claude
recovery test supplements them.

## What Changed

- Extract and assign the first-task skill while retaining the production
greeting and opening question.
- Carry the Codex skill-instructions flag through thread start and
resume. Resolve explicit task skill references only against assigned
skills and send native skill inputs.
- Invoke an unambiguously selected assigned skill through Claude ACPX’s
native slash-command parser on initial and resumed turns, retaining the
entire task/wake envelope as its argument. Do not carry that invocation
into ordinary tasks.
- Restore the saved single-task proposal modes: confirmation card, or
saved plan with revision-targeted checkbox approval. Explicit plan
requests also require a saved plan.
- Add first-response and complete-journey cases with fixed user facts,
acceptance checkpoints, durable outcome checks, and accounting for child
runs.
- Fail the eval when choice questions have fewer than two real options.
Recognize planning documents without treating them as completed work.
- Add optional, bounded quality judging as explicit post-processing.
- Render full conversations and static interaction cards in the shared
report. Conversations start folded. Show original and regraded results
and incomplete journeys distinctly.
- Keep credential-persistence scanning outside the first-task behavioral
suite; retain public evidence redaction.
- Refresh generated capability references after the API-reference edits.
- Correct shared native question guidance and tool schemas: choices need
at least two meaningful options; open-ended questions use canonical text
fields with the required compatibility payload. Verify both formats
through real tool-authority persistence.
- Disable announcements automatically for every isolated Runner E2E
process and label the gallery environment/provider/target explicitly.
- Remove CI races in the GitHub connection browser test and native
session recovery test by waiting for the actual async work before
asserting its results.

## Verification

- `pnpm exec vitest run
server/src/services/onboarding-first-task-assets.test.ts
server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19
passed.
- `pnpm --dir packages/paperclip-runner exec vitest run
src/drivers/acpx/runtime-host.test.ts
src/drivers/acpx/native-skill-prompt.test.ts
src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command
forwarding and the 1 MiB input boundary both failed before their fixes
and passed afterward. Coverage includes changed skills on reopen,
approval context, and an ordinary subsequent task.
- Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64
first-task fixture and grader tests also pass.
- Full repository typecheck and build passed locally. Server typecheck
and Runner build passed again after the native-command change.
- Full GitHub Actions CI passed on `23e56447b`: all
server/workspace/browser shards, Runner verification, typecheck/release
registry, build, canary, policy, and Docker checks. Greptile reviewed
this exact head at 5/5 with no unresolved threads. The earlier broad
local run had database startup/timing failures that passed isolated
retries; the complete remote suite is green.
- Merge verification against current master: 312 harness tests and 13
native recovery tests passed. Regenerated semantic contracts and fixture
hashes pass their consistency check. Full local typecheck and build also
passed on the stacked queue branch. After merging the latest master and
preserving the GitHub setup timing regression in the split browser
suite, both focused GitHub browser tests passed. Three CI timing/startup
flakes passed local verification and one remote retry; all latest-head
checks are green.
- Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local
mock API confirmed that `/skill-name` expands the assigned skill body
before the model request and retains the task arguments. A prose mention
does not. The probes made no paid model calls. The ACP probe used the
current first-task skill body and retained the wake arguments.
- [Full 48-case campaign and
report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/):
44 passed after three interrupted Codex cases completed in targeted
reruns. Original results, regrades, and all 51 executions remain in the
report provenance.
- [Claude campaign after the shared-question
fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/):
10/12 passed with zero single-option failures. All 12 recorded the
current assigned skill and corrected guidance. The failures exposed
skipped skill invocation and a missing saved plan. This PR adds native
command invocation and explicit saved-plan instructions; the subsequent
report below still shows behavior failures.
- [Fresh 12-case Claude
report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/)
at `78452129e`: 10/12 pass after correcting two false proposal-matcher
failures. The recordings said “Here is the task I will create and
run/complete” in approval cards; the old matcher missed that word order.
Regression tests failed before the fix and pass after it. Original
results and offline regrade provenance remain linked. No agent rerun was
needed. Zero single-option-question failures; two behavior failures
remain: direct work before acceptance on a plain first message, and an
explicit plan request without a saved plan. Neither check was relaxed.
The follow-up `82087ac7e` fixes command-prefix size accounting;
`94aefb1f3` fixes only that proposal matcher.
- Report browser checks confirm folded conversations, rendered cards,
explicit Local/Daytona labels, and no page errors. The published-object
audit scanned 1,306 text files across 2,154 objects with no
credential-format findings or prohibited files. Image pixels and unknown
token formats are outside that scan.

## Risks

- Model behavior is nondeterministic. One campaign is evidence, not a
guarantee. The two remaining Claude behavior failures are visible in the
report and require further product work; this PR does not claim all
onboarding scenarios pass.
- The suite checks persisted Paperclip effects. It cannot prove the
absence of arbitrary external effects.
- Historical recordings can miss later journey steps. These remain
incomplete, never passes.
- Native profiles switch runtime after the production onboarding wizard
because it does not yet expose a native option.
- Quality scores are informational and cannot override behavioral
failures.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, and code
execution. The exact deployed model identifier and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 14:24:32 -05:00
cceeb0aa66 test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner must support project work, delegation, hiring, and
service access.
> - Browser tests exposed lost connection access, rejected helper
events, and stalled recovery.
> - Some eval failures also came from incorrect fixtures and decision
controls.
> - This pull request fixes those paths and adds eight everyday workflow
stories.
> - The tests retain observed failures and verify delivered files
independently.
> - The benefit is repeatable evidence for common user tasks and their
remaining gaps.

## Linked Issues or Issue Description

Related work: #13404 contains earlier workflow fixes. #13300 and #13470
changed the CI contracts used by the harness security tests. Merged
companion:
[paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22).

**What happened?**

Native ACPX sessions did not receive the assigned connection gateway.
Codex helper events could arrive before their spawn receipt and fail
thread validation. A parent continuation could take a shared workspace
before its child retried. A failed native continuation could leave the
task status without a clear recovery blocker. The eval harness also
confused tool approvals with new connection requests and could reject a
valid delegated download.

**Expected behavior**

Keep assigned gateway access and its approval checks. Verify helper
lineage before accepting helper progress. Let a waiting child proceed
before automatic parent recovery. Preserve a failed task's recovery
ownership. Grade the actual requested workflow and its delivered files.

**Steps to reproduce**

Run the everyday workflow suite with the native Codex and Claude
profiles. Exercise service approval, connection refusal, delegated
project work, and teammate reuse. The commands and case requirements are
in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm
test:runner-recovery` for controlled crash and replacement cases.

## What Changed

- Pass the scoped connection gateway binding through the native ACPX
host and sidecar.
- Recognize Codex helper lineage from parent metadata and spawn
receipts. Verify early helper events with `thread/read`. Keep helper
events separate from root completion authority.
- Guide agents to use persistent hiring, child tasks, dependency
records, and a blocked handoff while waiting for a child.
- Defer automatic parent recovery while a child has an active execution
path in the same shared workspace. Allow parent recovery when the child
needs review.
- Record Blocked status and recovery evidence when a failed native
continuation needs reconciliation, including existing active or
escalated incidents. Preserve their owner and retry budget.
- Add eight browser-driven workflow cases. Use real decision controls,
explicit child feedback delivery, managed hiring credentials, and
independent ZIP checks inside a bounded Docker sandbox. Verify sandbox
availability before task creation. Record screenshot SHA-256 at capture.
- Keep runner crash probes in controlled recovery tests. Preserve the
original failure when cleanup also fails.
- Display missing accounting and replay revisions as unavailable. Align
harness security assertions with the approved CI changes.
- Make the channel-rejection browser fixture bind its file after the
send captures its payload. This prevents live refresh from removing the
file before the simulated race.

## Verification

- Full workspace `pnpm -r typecheck` passed after merging current
master.
- Runner E2E typecheck passed. Harness unit tests passed: 216/216.
- Wake-queue database tests passed: 55/55. The two added
existing-incident tests failed before the fix and pass after it.
- Docker artifact calibration passed: 12/12. Host-file and host-loopback
isolation tests failed before the fix and pass after it. Read-only
delivery and output limits are also verified.
- Full `pnpm build` passed. Targeted recovery tests passed: 83/83.
- The channel-rejection browser test passed five consecutive runs after
fixing the fixture race found in CI.
- Local general-server (12,351 tests), UI (6,250), CLI (485), and
workspace package groups passed. The monolithic run stopped at an
unchanged lock-heartbeat fixture race; the isolated workspace group
passed on rerun (shared: 747/747). A separate local serialized run
passed 97 files before two socket errors in the unchanged issue-list
route suite; that suite passed 15/15 on isolated rerun. These local full
commands did not finish uninterrupted; the complete CI matrix below
covers the remaining suites.
- Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful
checks, 2 expected skips**, including every server/workspace shard,
browser shard, native runner verification, build, and typecheck. [Final
CI
run](https://github.com/paperclipai/paperclip/actions/runs/34989136700).
- Greptile reviewed this exact head at **5/5**; all review threads are
resolved. Both Superagent security checks are successful.
- ACPX credential-boundary tests passed: 118/118. Superagent accepted
the runner/sidecar versus provider-environment trace and cleared its
finding.
- The latest paid local campaign on source
`f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8,
Claude 7/8, Mini 7/8. These results predate the merge with current
master.
- The two remaining failures are in `hire-reuse`: Claude exceeded the
attempt deadline during final review; Mini made invalid deliverable tool
calls and remained Blocked.
- Six Daytona cases were not run because the matching immutable runner
image was unavailable. This PR does not claim new remote model results.

## Risks

The changes affect connection admission, helper identity, and recovery
scheduling. Assigned gateway grants and user approval still govern
service calls. The workspace admission gate still exists; the broader
folder-sync design is separate work. Provider behavior can still cause
the two recorded hiring failures. No database migration is required.
Paid cases are opt-in and have bounded attempt deadlines. Project
stories now require Docker and the documented pinned Python image on the
harness host.

## Model Used

OpenAI `gpt-6-astra` performed implementation, diagnosis, and
substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR
preparation, and review tracking. Both used repository tools and code
execution. Context-window sizes were not recorded.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks and
isolated reruns; full-run limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 11:04:16 -05:00
Dotta 1dceee9a4e fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path

> - Paperclip manages AI agent work and the execution state for each
task.
> - Remote agents run in sandbox environments such as Daytona.
> - Daytona keeps files while a sandbox is stopped, but deletion removes
those files.
> - Runner Codex did not copy successful remote workspace changes back
to the host workspace.
> - A warm sandbox could therefore hide data loss until Daytona replaced
or deleted the sandbox.
> - This pull request makes the host workspace durable after every
successful turn and keeps verified reusable sandboxes warm.
> - The benefit is reliable multi-turn work across warm reuse, restart,
stop, and sandbox replacement.

## Linked Issues or Issue Description

**What happened?**

A successful native Codex turn in Daytona could leave workspace changes
only in the remote sandbox. A later warm turn appeared to work because
it reused that filesystem. A replacement sandbox could start from stale
host data and lose the successful changes.

**Expected behavior**

Paperclip must merge each successful remote turn into the authoritative
host workspace before it completes the run. A verified warm lease may
reuse its remote files. A replacement lease must reconstruct the exact
durable workspace seed.

**Steps to reproduce**

1. Run Codex in a reusable Daytona environment.
2. Write a file during one successful turn.
3. Replace the Daytona sandbox before the next turn.
4. Observe that the next turn can start without the prior file on the
unpatched code.

Related remote workspace foundation: #10070.

## What Changed

- Added explicit `host_current`, `durable_seed`, and `adopt_remote`
workspace preparation modes.
- Added atomic, versioned native workspace descriptors and seed archives
under `PAPERCLIP_HOME`.
- Added real native sandbox export and three-way host merge before
terminal result completion.
- Added workspace-only recovery after a proposed result. Recovery does
not submit another provider turn or consume the provider retry budget.
- Added fail-closed handling when a sandbox with unexported changes is
gone.
- Kept healthy reusable Daytona sandboxes started for legacy Codex and
Runner Codex.
- Kept the Runner Codex process and provider session across verified
warm turns.
- Added the paid `daytona-warm-continuity` browser suite. It contains
exactly the legacy Codex and Runner Codex cells. Each cell performs
three measured turns.
- Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`.
No package script was added.
- Added no database migration. The metadata format is backward
compatible and idempotent.

## Verification

- `pnpm typecheck`
- `pnpm test:e2e:runner:unit` — 114 passed
- Native workspace, finalizer, session, and environment tests — 232
passed
- Daytona provider tests — 150 passed
- Workspace staging and merge tests — 98 passed
- Runner transport tests — 63 passed
- Legacy Codex restore tests — 5 passed
- Rust format and compile checks pass through root typecheck
- The paid Daytona suite was not run locally because the required
Daytona, OpenAI, and immutable image credentials are not present.

## Risks

- The main risk is an incorrect workspace identity or merge after a
crash. Durable descriptors bind the run, workspace, lease, provider
lease, local root, remote root, and baseline digest. Ambiguous evidence
fails closed.
- The host merge may conflict with concurrent host edits. The existing
three-way merge and exclusion rules handle this case and surface
failures.
- A deleted sandbox cannot recover unexported bytes. Paperclip now
blocks with `workspace_sync_out_unrecoverable` instead of reporting
success or rerunning the provider.
- There is no database migration. Descriptor writes and recovery are
atomic and idempotent.

## Model Used

OpenAI Codex with GPT-5. The run used agentic reasoning, repository
inspection, code execution, test execution, Git, and GitHub CLI tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-05 13:00:57 -05:00
Dotta 5716fe907e test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner subsystem executes agent work across local and managed
provider backends.
> - The lower pull requests restore the task runtime, provider backends,
and managed-provider control plane.
> - The restored system needs repeatable full-stack checks before it can
ship safely.
> - Paid live checks also need clear access, cost, and secret controls.
> - This pull request adds acceptance, live evaluation, chaos, and
release gates for the restored runner stack.
> - The benefit is measurable runner parity with safer release
decisions.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting. This change covers runner tests, release workflows,
server contracts, and evaluation tools.

**Problem or motivation**

The runner stack did not have one complete acceptance surface for native
Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could
miss provider drift, task-view regressions, cost-policy errors, and
destructive cleanup errors.

**Proposed solution**

Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid
workflows. Add live evaluation, chaos, cost-limit, redaction, and
release contract checks. Add AWS AgentCore infrastructure and guarded
provisioning tools. Keep the native runner experimental flag off by
default.

**Alternatives considered**

We considered manual smoke tests only. They do not give repeatable
evidence and they do not protect release branches. We also considered
one large pull request. The stacked pull requests keep each review below
the Greptile file limit.

**Roadmap alignment**

This work supports the shipped Cloud / Sandbox agents milestone and the
shipped Agent evals & feedback milestone in `ROADMAP.md`.

Related stack:

- #12699 adds managed provider backends and lifecycle support.
- #12691 adds qualified OpenCode and ACPX provider backends.
- #12685 restores task runtime rendering and steering.

## What Changed

- Add the runner full-stack harness with 57 catalog cells and 60 unit
tests.
- Add a Daytona runner image with digest-pinned base images and
base-aware image-content checks.
- Add guarded live evaluation and chaos workflows with a fixed
40-execution matrix; live and full-stack paid schedules now run only on
Sundays or by manual dispatch.
- Add in-flight reported-usage cost stops, post-turn cost caps,
exact-threshold failure classification, secret redaction, retry
classification, and actor authorization.
- Reattach stream and hard-budget listeners before restart-recovery
continuations so restored paid sessions cannot bypass in-flight
interruption.
- Preserve OpenCode usage and cost across tool-loop messages and turns
while exposing an explicit current-run delta to durable accounting.
- Keep PNG/WebM evidence in access-controlled artifacts only, reject
SVG, and publish only pruned inert structured per-attempt evidence.
- Add AWS AgentCore infrastructure, provisioning checks, and smoke
tools; reject unsafe model identifiers, require exact stack ownership
markers, and make failed-stack replacement explicit.
- Add evaluation-session contracts and capability reports.
- Add release workflow checks for immutable action pins, frozen
dependency installs, exact weekly cron shape, paid-run guards,
provider-secret isolation, and chaos test paths.
- Reauthorize the original and triggering numeric actor IDs as the first
step of every provider-secret job, including partial reruns, before
checkout or provider access.
- Give each full-stack matrix cell only its matching provider
credential, expose Daytona only to Daytona cells, and disable shared
dependency caches anywhere paid credentials or OIDC write access are
present.
- Protect the legacy manual E2E workflow with the same default-branch,
allowlist, environment, and per-job authorization boundary.
- Rotate live-eval candidates by week and retain 120 days of compatible
history so the seven-week trend window remains viable.
- Restore the root runner-acceptance commands and reconcile reported
snapshots,
raw receipts, and terminal usage without double counting or losing late
usage.
- Mark ACPX token deltas exact only when every budget field is present,
keep
cumulative cost/request authority separate, reject non-USD cost
labeling,
  and include thought tokens in output-token budgets.
- Keep `enableNativeRunner` off by default. The acceptance harness
enables it only in its isolated test instance.

## Verification

Passed locally:

- `pnpm --filter @paperclipai/paperclip-runner typecheck`
- `pnpm test:runner-acceptance:typecheck`
- `pnpm test:runner-acceptance` (19 tests)
- focused OpenCode proxy, driver, runnerd transport, live-session, and
turn-stream tests (106 tests)
- `pnpm --filter @paperclipai/paperclip-runner exec vitest run
src/live/clean-room-server.test.ts` (22 tests)
- `pnpm test:e2e:runner:typecheck`
- `pnpm test:e2e:runner:unit` (62 tests)
- `node --test scripts/__tests__/release-verify-workflow.test.mjs`
- `pnpm --filter @paperclipai/paperclip-runner
test:runner-workflow-evals` (22 tests)
- `pnpm -r typecheck`
- `pnpm build`
- `node --test
packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs`
(6 tests)
- `git diff --check`
- `cargo test --manifest-path
packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core
--lib --locked` (161 tests)
- focused ACPX provider-event tests (10 tests)
- The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged.

I did not run paid live provider jobs or provision AWS resources. Those
checks need credentials and can create cost.

## Risks

The paid workflows can create provider cost. They require an allowlisted
original and triggering actor, the protected `runner-e2e-paid`
environment, explicit opt-in variables, and cost limits. The four
provider credentials exist only in that master-only environment, which
requires allowlisted reviewer approval and disables administrator
bypass; repository and organization Actions scopes contain no copies.

Provider usage arrives after a billable request, so the live guard
cannot prevent one request from crossing a threshold. It interrupts
immediately on the first reported threshold hit and permits no
continuation.

Visual evidence can contain secrets rendered as pixels. PNG/WebM remain
only in access-controlled workflow artifacts; SVG and per-attempt XML
are excluded, and S3/Pages receive a pruned structured dashboard.

The AWS scripts can create cloud resources. They use explicit commands,
least-privilege roles, KMS encryption, saved nonsecret metadata, and
explicit teardown.

This pull request does not enable the experimental native runner for
existing instances.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex with GPT-5. The model used extended reasoning, tool use,
code execution, and parallel subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-02 08:55:08 -05:00