9 Commits
Author SHA1 Message Date
DottaandPaperclip dd868ed125 fix(runner): share native completion tool guidance (#14961)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Native Runner agents report completion through finish and block
tools.
> - The providers receive different descriptions for those tools.
> - Completion guidance belongs with the tools that enforce the result.
> - This pull request shares the descriptions and refreshes retained
catalogs.
> - A separate native suite checks completion and blocking on production
defaults.
> - Legacy agents retain their separate skill and API paths.

## Linked Issues or Issue Description

Refs: #14920, #14948, #14985.

**Current behavior**

Native Codex and MCP bridges describe finish and block differently.
Retained provider sessions can keep old descriptions.

**Proposed behavior**

Native providers receive the same finish and block descriptions. The
descriptions cover report selection, validation feedback, returned
outcomes, approval gates and final-answer timing. Retained native
sessions refresh from v13 to v14.

**Reason and benefit**

Put the completion procedure next to its native tool. Preserve stock
base instructions, schemas, permissions and terminal semantics. This PR
now stands alone on master. It contains no reduced manual, shared prompt
or operational-skill changes from #14948.

## What Changed

- Add canonical native finish and block descriptions. Use them in direct
Codex and both native MCP bridges.
- Advance the native tool contract to v14. Cover old-v13 refresh without
replacing task identity or prior history.
- Check authenticated tool catalogs, provider start/resume frames and
serialized daemon catalogs.
- Add an independent, explicit-only native completion suite. Preserve
the original assigned-skill durable-document journey. Pair it with a
concrete whole-task blocker across Codex, ACPX Claude and OpenCode.
- Verify the actual public production default bundle and budgets before
execution. Require independent durable disposition, native
result/terminal receipts and observable provider-final ordering.
- Correct the blocker browser oracle to accept the requested
explanation. Keep exact owner/action/scope checks. Calibrate positive,
missing and contradictory replies.
- Preserve only actual `tool_call` terminal names (`paperclip_finish` /
`paperclip_block`) in the native compatibility run-log projection.
Require the same named call ID through its finishing result; retain all
other redaction boundaries.
- Admit verified hosted shallow checkout/build hydration and bind the
selected runnerd to exact source/archive/binary provenance. Hosted cells
truthfully reuse the existing trusted build; local admission executes
Rust calibration. Forward only public source/run identifiers through
both launcher preflight subprocess paths.
- Enforce single attempts in the launcher for opted-in fixtures. Keep
ordinary retry policy unchanged. Run exact-source, credential-free
admission before credential loading.

## Verification

- Frozen candidate: `d6e59e4712a3158ab4cd7d58deff1389b4578c21`, based on
master `59c07ede72dc08b8aba149a01cc11e0b7a204621`; historical
descriptions: `e74ed61a69fbdd8b3a8f15dd6456bc3140246e33`. Exactly the
five original native production files and six unit tests differ. Both
carry identical corrected fixtures, strict named finishing-call grader,
closed compatibility carrier and admission. Defaults,
profiles/models/auth/permissions and manifest bytes match.
- Actual launcher `prepareNativeCompletionPreflight` →
`verifyNativeCompletionPreflight` admission passes on both exact refs
with zero providers: candidate 132 / historical 127 selected TypeScript
assertions, 128 Node calibrations and one Rust normalization calibration
each; E2E typecheck, manifest checks, selected binary provenance and
six-cell discovery pass. Each has 257 explicitly skipped unrelated
assertions, not coverage. The credential-free environment calibration
exercises both real prepare/verify subprocess options with public hosted
identifiers and rejects credential/ambient overrides. Complete actual
launcher prepare→verify also passes on both frozen refs with explicitly
synthetic hosted metadata/verified archives, separately labeled as
calibration rather than a trusted GitHub run. Exact framed provenance
parsing and mock source identity are calibrated without relaxing the
real verifier.
- [Complete matched qualification
report](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-master-qualification.md),
[immutable
manifest](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-manifest.json)
and [closed retained
audit/hashes](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-results/comparison.json)
are inspectable. All six candidate cells pass; historical descriptions
pass five. Paired outcomes: **zero new failures, one new pass (Codex
blocker), five unchanged passes, zero pending pairs**. [Candidate
campaign](https://github.com/paperclipai/paperclip/actions/runs/37098728980)
and [historical
campaign](https://github.com/paperclipai/paperclip/actions/runs/37098815696)
each execute six original attempt-1 native runs, with no campaign retry
and successful cleanup. Their trusted workflow revision is
`215586d127e97c9301d86e769a39a15c13298ca2`, separate from measured
source. [Candidate public
HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098728980-1/index.html)
and [historical public
HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098815696-1/index.html)
retain declared screenshots.
- Independent candidate evidence agrees with all original grades: 51
strict native checks, 12 served-default/budget checks and 21 original
skill/document checks pass. The historical Codex blocker saves the
correct whole-task blocker but omits the required marker from its actual
provider final and identical saved reply. This is not semantic-summary
fallback. Its original browser/matcher failure stays retained; the
additional native snapshot/grade and workspace before/after digest were
never written and are not fabricated by the separate API/PRP audit.
Historical Codex completion has one failed finish followed by success
within the same native run; the public receipt records no failure
reason. All twelve runs and their usage remain counted. Reported
model-cost subtotals are $0.00421482 historical/$0.00437391 candidate;
Codex/Claude zero entries have unknown billing type, actual invoices are
unverified and hosted execution cost is unmetered. One matched trial
supports no extra failure within these six cases, not broad statistical
or coding-quality equivalence.
- Initial hosted `e18c2cf9` / `459455ac` and subsequent `0a9c5a7` /
`00a761b` cohorts each stopped before providers in all twelve cells. The
latter failed a mocked-receipt unit test under ambient hosted metadata;
all source/build proofs passed. [All twelve later setup
receipts](https://github.com/paperclipai/paperclip/blob/402ee94c52273ad58de355ae9a7d562dd22f8101/doc/plans/2026-10-02-native-completion-qualified-hosted-setup.json)
are retained. [Exact failed setup
receipts](https://github.com/paperclipai/paperclip/blob/27653eb1a8f8ce839776d760f4563f672e5a706c/doc/plans/2026-10-02-native-completion-master-hosted-setup.json)
and the original manifest remain intact. Local sandbox-denied loopback
and stale anchor-expectation attempts are retained separately; unchanged
appropriate assertions were corrected/admitted before paid dispatch. Old
anonymous OpenCode streams are not assigned inferred tool names or
retroactively passed.
- Full provider-free E2E support previously passed 927 tests in 67
files. Exact-head d6 normal CI run `37098409915`, attempt 1 passes full
repository typecheck/build/tests, Runner Rust/static checks, all browser
shards/aggregate and canary: 52 check-runs pass, four intentional skips,
Snyk passes. Fresh Greptile check `111132956342` is 5/5 with zero
unresolved threads. Source-specific deterministic tests do not
substitute for the bounded live comparison.
- Earlier native source `9138f570c341c251a5727c32d6615ce238bc8e03` is
archived. Its [complete reduced-manual-context
report](https://github.com/paperclipai/paperclip/blob/9138f570c341c251a5727c32d6615ce238bc8e03/doc/plans/2026-10-02-native-completion-live-comparison.md)
remains intact, including original failures, grader limits and
provider-free replay. It is not current-master-context qualification.

## Risks

Changed tool text can change model behavior. The completed six-pair
qualification shows no extra failing outcomes in this bounded trial;
other tasks and repeated-run variance remain unmeasured. Observable
final ordering does not prove provider feedback consumption. Public
evidence can fail closed if a provider does not expose the required
result sequence. This slice does not remove native fixed prompts or
measure general coding quality. No database, schema, permission or
legacy completion changes occur.

## Model Used

OpenAI Codex, GPT-6 family, with code inspection, execution and tool
use. The exact deployment ID and context-window size are not exposed in
this session. They are unavailable rather than inferred from the model
menu.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 06:20:38 -05:00
DottaandPaperclip d30b03bd8c test: add persistent E2E coverage for human blocker decisions (#14707)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Agents use the coordination skill when work needs human authority or
a scope decision.
> - PR #14188 replaced automatic manager escalation with direct blocker
handling.
> - This behavior needs real browser, server, database, and provider
tests.
> - The test must verify saved human input, task ownership, and resumed
work.
> - This pull request adds six reusable Product E2E cases and improves
the skill examples that they exercise.

## Linked Issues or Issue Description

Refs #14188.

The merged change needs repeatable behavior coverage. The new suite
tests missing administrator access, missing hiring permission, and
requester scope questions. Searches found no duplicate blocker-guidance
suite. This extends the existing eval system described in ROADMAP.md.

## What Changed

- Add the explicit-only `blocker-guidance` Product E2E suite. It has
three local scenarios on legacy Codex and legacy Claude.
- Use the production UI and public APIs to create work, save a
human-only question or confirmation, answer it after reload, and resume
the same task.
- Check requester identity, ownership history, manager activity, hiring,
saved answers, and completion. Keep direct text input as a separate UX
result.
- Save pending and final screenshots, API checkpoints, skill hashes,
provider evidence, and billing data through the existing report
pipeline.
- Isolate the Claude fixture home. Verify the served skill bytes before
dispatch so an old installed skill cannot silently replace the evaluated
skill.
- Improve the coordination and hiring skill examples. Include the
human-only policy, requester address, wake behavior, and handling of
authorized scope changes.
- Grader v5 requires the exact approved public welcome note as a new
worker comment. Browser input checks reject unwritable scope cards
before clicking, and confirmation direction must be saved in the
resolution before the worker wakes.
- Add grader calibration and browser-input tests. Update the fixture
guide and generated capability inventories.

## Verification

- `pnpm build`: passed after rebasing onto current master.
- `pnpm -r typecheck`: passed.
- `pnpm test:e2e:runner:typecheck`: passed.
- `pnpm test:e2e:runner:unit`: 742 tests passed.
- `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests
passed.
- `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells
found.
- Capability inventory and generated-contract checks: passed.
- `pnpm exec vitest run
server/src/__tests__/hiring-operational-examples.test.ts`: four tests
passed after synchronizing the generated API reference and section
anchor.
- Full general and serialized test suites: passed in CI on
`6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm
test:run` was interrupted after complete CI coverage passed; it is not
claimed as a completed local full-suite run.
- Final GitHub checks: 54 passed, two optional Storybook checks skipped.
The runtime-exposure startup test hit a 10-second readiness timeout
once, passed a targeted local reproduction, and its CI shard passed the
single retry without code changes.
- Current-head Greptile: 5/5, clean check, zero unresolved threads.
- Historical live measurement on September 29 at
`4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell
runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex
`gpt-5.6-sol` passed 8/9. These runs predate this rebase.
- Version 5 changes the scope answer to an exact approved publication
draft. The historical runs do not qualify that new requirement; the
two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec`
passed Codex and failed Claude. Claude posted the correct salary-free
sentence but omitted its required reference line from that comment,
placing the reference in a separate completion message. The
`public-welcome-note` check correctly failed. An earlier Claude
database-startup failure was retained separately; its fresh-instance
retry reached the model. This pilot is not a six-cell qualification.
- The failed Codex scope case omitted `addresseeUserId`. The strict
routing check remains. All 18 attempts had clean evidence manifests and
passed cleanup.
- To repeat with provider credentials: `pnpm test:e2e:runner -- --suite
blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is
excluded from `--all`.

## Risks

- The live suite measures variable model behavior. The retained 17/18
historical result and the current 1/2 scope pilot are not all-pass
qualifications. These paid cases are opt-in; their observed model
failures remain visible independently of deterministic CI checks.
- A separate generic task-replacement diagnostic still exposed a Claude
refusal. The ordinary cases use specific business decisions. The
diagnostic is not a standalone catalog case in this change.
- Earlier measurements included an old installed Claude skill and test
defects. Their grades remain retained and are not combined with the
three final repetitions.
- Skill examples can affect when agents ask for human input. Downstream
permission checks still apply.
- Native runners, Daytona, agent-requester routing, and real external
connection authorization are outside this suite.
- Raw provider traces and credentials remain private. No screenshots,
raw reports, secrets, workflow changes, or lockfile changes are
committed.

## Model Used

OpenAI GPT-6 through Codex assisted with this change. The exact deployed
variant and context window size are not exposed in this session. The
assistant used reasoning, repository edits, tool use, and shell
execution. The evaluated models were `gpt-5.6-sol` and
`claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:56:54 -05:00
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00
DottaandPaperclip 96bf004a79 fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its control plane decides when a task can continue, wait, stop, or
complete.
> - Legacy continuation could change when an agent changed its wording
without changing task state.
> - Shared attempt counts also let repair and infrastructure retries
affect each other's limits.
> - This pull request uses persisted state and separate, bounded
allowances for these decisions.
> - If automatic repair stops, the task explains what happened and
offers a guarded retry.
> - Paired tests and real-provider evaluations verify that Stop,
approvals, ownership, and spending limits remain authoritative.

## Linked Issues or Issue Description

Related work: Refs #13761, Refs #11126, Refs #13610. These cover
obsolete continuation dispatch and retry storms. Open and closed issues
and PRs were searched for related lifecycle, continuation, and retry
work.

**What happened?**
Legacy continuation depended on English wording and progress heuristics.
Repair, failure retry, and productive continuation could consume shared
counts. When bounded repair stopped, the task showed a technical
recovery message without a clear next action.

**Expected behavior**
Persisted disposition and owned execution paths determine the next
action. Missing disposition prompts bounded agent repair. Explicit work
mode determines planning mode. Narrative changes and raw activity counts
cannot replenish allowances. An exhausted repair shows a readable
notice. An explicit retry checks current controls and preserves the
assigned agent.

**Steps to reproduce**
Run `pnpm test:lifecycle-baseline`. The paired probes keep structured
state constant while varying completion, planning, blocker, and progress
prose. Run the explicit `lifecycle-baseline` and
`continuation-accounting` Product E2E suites for real-provider coverage.
In Storybook, open **Design previews / Recovery notice** to inspect the
production component's normal, pending, acknowledged, unavailable,
failure, and mobile states.

## What Changed

- Hide the image attachment button, icon, and drop/paste hint in answer
composers. Image paste and drop support remains available.
- Merge current master and retain both browser regression sets. Use a
production-stamped service worker in the offline recovery browser
fixture.
- Share one state-based legacy continuation decision across immediate,
delayed, and recovered dispatch. Bind bounded repairs to their source
run and episode.
- Remove title and description wording from work-mode authority. Agents
can still write requested plans in execution mode.
- Persist separate failure-retry and productive-continuation counters.
Disposition repair and resource waits cannot consume or reset those
allowances.
- Validate delayed repair identity, then recheck current gates before
provider dispatch. Fence native startup cancellation.
- Show **Agent needs attention**, a plain-language explanation, **Retry
agent**, and expandable details in both task interfaces. Report request
progress, acknowledgement, and errors inline.
- Store typed recovery notice metadata. Recognize older active notices
only through exact stored action and run IDs. Notice text never grants
retry authority.
- Use the existing recovery-action endpoint for retry. Recheck current
action, status, owner, agent availability, dependencies, active runs,
pending questions and confirmations, approvals, pause controls, and
budget. Duplicate requests do not wake twice.
- Add component, page, route, database, contract, and Storybook
coverage. Keep the scenario inventory and executable evals here.
Historical reports and snapshots live in the [commit-pinned
paperclip-evals
archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md).
- Preserve unsaved project fields while the same project URL changes to
its canonical alias. Do not reuse data across projects or companies.
This separate fix addresses the repeated repository-editor browser
failure without changing the browser test.
- Keep the development service worker from intercepting Vite module
reloads. Update the connection-intent browser fixture to record progress
and completion through the agent API.

## Verification

Merge preparation on September 25, commit
`c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`:

- Merged master `bd2030932` and resolved the browser test-list conflict
by keeping both sets of regressions.
- Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no
failures, skips, or missing selected evidence. Unit 423, runner 184,
database integration 397, grading 86.
- Browser support: 17/17 passed. The offline recovery test first failed
with an unstamped development worker, then passed with the production
stamp. Its assertions are unchanged.
- Focused interaction UI and offline fallback tests: 19/19 passed.
Verified the custom-answer composer in Storybook: no attachment controls
or hint; entering an answer enables Next.
- Recursive typecheck, production build, token gates, and diff checks
passed. The worktree is clean. No new real-provider campaign was run.
- Current CI and review: [Current PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243):
55 successful checks and two optional Storybook skips. Greptile scored
this exact commit 5/5. Hiding the question attachment controls is an
intentional UI change; paste/drop remains available.

Earlier recovery UI verification, commit
`21be0fec0e90e86b6d662b8ee4831847cd041cdb`:

- Recursive typecheck, production build, token gates, and diff checks
passed.
- Focused UI coverage: 338 tests passed across six suites (336 before
the interaction guard, with the two affected suites rerun at 149 passed
after it). Covers both task interfaces, the real page mutation,
pending/error acknowledgement, stale state, and unavailable controls.
- Recovery database integration: 352 tests passed before the interaction
guard. The complete recovery-action and mutation-route suites passed 181
tests after it. The two new pending question/confirmation regressions
failed before the fix and passed afterward, including
resolved-interaction controls. Shared validator suite: 31 passed. E2E
catalog suites: 34 passed.
- Browser inspection passed for light/dark themes, mobile layout,
expandable details, pending retry, acknowledgement, failure, and
disabled retry. Storybook renders the production component; its request
is simulated.
- The broad local run hit two chat callback-order wait failures and was
stopped after all CI unit/database/runner shards passed. Both local
failures passed when rerun without the competing full-suite process.
- CI exposed a repeated project-repository draft-loss race during
canonical redirects. A new unit regression failed before the fix; all
nine project-page tests now pass, including controls for other projects
and companies. Both unchanged repository browser tests passed against a
fresh local server. UI typecheck, production UI build, and token gates
passed after this fix.
- [Earlier PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798)
on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two
optional Storybook skips, and no failed or pending checks. The
repository browser shard passed with the production fix. Greptile is 5/5
on this exact commit with no unresolved review threads. The PR is
mergeable.

Historical, source-qualified lifecycle evidence:

- Lifecycle baseline: 1,074 assertions. Native session coverage: 447
tests. Product E2E support: 515 tests. Browser support: 11 tests. Full
earlier verification is retained in the archive.
- [Real-provider campaign: 8/8 passed, zero
retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html),
source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup
checks passed. This includes deliberately exhausted repair cases that
correctly remain blocked; it does not mean every task finished Done.
This campaign predates the recovery UI change.
- Archive migration verified all 16 original JSON files byte-for-byte
and all 24 checksum entries. App tests do not need private archive
access. [Archive PR
#27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged.

## Risks

- Agents that omit durable disposition receive at most two repair
attempts by default. Prose-only completion exposes missing state rather
than silently changing scheduling.
- A retry is an explicit board action. The server rechecks current
controls. A successful response confirms the task returned to To do; it
does not claim that the provider has already started.
- Existing notice metadata remains valid. Only older active notices with
matching structured evidence receive the new UI. Historical notices
without that evidence keep their existing rendering. No schema migration
is required.
- Old run records require conservative retry accounting. Tests cover old
counters, alternating retry lanes, restarts, and exhausted repairs.
- Historical snapshots require private `paperclip-evals` access. The app
index retains public campaign links. Live campaigns qualify specific
sources and scenarios; no new real-provider campaign has run for the
recovery UI commit.

> This fixes existing lifecycle and recovery behavior and does not
duplicate planned core work.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, reasoning, code
execution, and review. The exact serving model ID and context window are
not exposed in this task. Historical real-provider evaluations used
Codex model `gpt-5.6-sol`, separately from the implementation assistant.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 15:28:11 -07:00
DottaandPaperclip efce9356b5 fix(ui): offer recovery when the app fails before React starts (#13970)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The browser must load its JavaScript before React can render a task.
> - A failed import can stop that process before the React error
boundary exists.
> - The HTML entry then leaves an empty page with no recovery action.
> - This PR adds a small recovery screen that works without React.
> - The user can retry the same page and return to saved task content.

## Linked Issues or Issue Description

Refs #13824 and #13895. This is a follow-up to their browser startup
investigation.

**What happened?**

Interrupting the app bundle or a required import leaves an empty React
root. A startup exception has the same effect. A React error boundary
cannot handle these failures because React has not started.

**Expected behavior**

The page must explain the startup failure and offer a manual retry. A
late successful load must dismiss the recovery message without a reload.

**Steps to reproduce**

1. Open a saved task in the browser.
2. Abort the application bundle request, or make a required module
return HTTP 503.
3. Observe the empty page before this change. With this change, use
Reload page after the fault clears and verify the saved task and
comment.

**Paperclip version or commit**

The failing regression baseline used master at `8781f06a8`.

**Deployment mode**

Local source build and compiled UI. Tests cover both initial navigation
and a page controlled by the production service worker.

The exact cause of the older intermittent Vite stall remains
unconfirmed. Forty app loads and thirty replays of retained responses
did not reproduce it. This PR fixes the missing recovery path; it does
not claim to remove that historical cause. A normal HTTP 304 response is
not a failure.

## What Changed

- Add an inline startup guard and recovery screen in the HTML entry. It
does not depend on the app module graph.
- Show a manual reload action after a startup error or after 30 seconds
without rendered root content.
- Remove the notice, timer, observer, and error listeners when the app
starts. Never reload automatically.
- Keep the recovery screen outside the React root so it cannot satisfy
app-readiness checks.
- Add browser tests for interrupted imports, a stalled import, an
evaluation error, service-worker-controlled retry, repeated offline
retry, and cleanup after successful startup.
- Return a static, uncached HTML retry screen when a
service-worker-controlled navigation fails offline. It contains no task
content.
- Add a full-app test that retries an interrupted compiled bundle and
checks the saved task, comment, composer, route, and absence of agent
runs.
- Document the coverage and the limits of the historical diagnosis.

## Verification

- Red baseline: four recovery cases failed; the normal-startup case
passed. After the change, all five recovery cases passed. The review
found an offline retry gap; that additional case failed before the
worker fix and passed afterward.
- Full provider-free browser-support suite: 16 passed.
- Compiled-app browser tests: four passed, including saved-task reload,
interrupted-bundle recovery, slow-CPU service-worker reload, and sidebar
navigation.
- Expanded service-worker, offline response, PWA, and worker build-ID
unit tests: 37 passed. The two old plain-text offline expectations were
reproduced as failures and updated for the HTML retry contract.
- UI production build, full local repository typecheck (`pnpm -r
typecheck`), runner-E2E typecheck, and design token checks passed.
- Manual browser check: a temporary server failed the compiled bundle
once. The recovery screen appeared. Clicking Reload page restored the
same saved task, comment, and composer.
- Full local `pnpm build` passed.
- Full local `pnpm test:run` was attempted with a bounded deadline and
stopped after it timed out. Workspace runtime/cleanup tests reported
timeouts on this host. The monolithic local run is not a pass. The
focused tests above and the complete Linux CI run provide the successful
verification.
- Final-head [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36072201966)
passed. All 53 check runs succeeded; the two Storybook jobs were
intentionally skipped. The legacy security status also passed.
- Greptile reviewed `f83e0f51fb760541d83353f2c1df4e182f3948f9`: 5/5.
Both review findings are fixed and resolved.

## Risks

- The guard only handles startup before React renders root content.
Existing React boundaries handle later rendering errors.
- A slow startup can show the message after 30 seconds. A later
successful render removes it; the page does not reload by itself.
- The fallback uses native HTML when the app stylesheet is unavailable.
- The worker changes only its offline navigation response. It returns
static HTML with a reload button and `Cache-Control: no-store`. Its
cache allowlist, private-response protections, task state, provider
prompts, and grading rules stay unchanged.
- This does not establish or fix the unknown cause of the historical
intermittent Vite stall.

## Model Used

OpenAI GPT-6 through Codex. The session exposes the GPT-6 family but not
an exact served model ID or context window size. Used reasoning, code
editing, shell tools, and browser testing. No subagents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 19:36:45 -05:00
DottaandPaperclip 9a76390cfa test(runner): improve blank-page diagnostics and infrastructure coverage (#13824)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner E2E tests verify real tasks and retain evidence for failures.
> - Some exposure tests assumed that port 42000 was free.
> - A blank task page could also fail without enough browser startup
evidence.
> - This pull request tests occupied ports and task reloads, and records
private startup diagnostics.
> - These changes make test failures easier to reproduce and explain.

## Linked Issues or Issue Description

Refs #13815. Report publication already received its production fix in
#13750; this PR adds a regression check for that workflow.

**What happened?**

Three exposure tests assumed that the allocator would select port 42000.
The synthetic failure disappeared when another process occupied that
port. A prior Daytona run also retained an empty task page after
navigation, but its evidence did not record pending modules or
service-worker control.

**Expected behavior**

Exposure tests must exercise the intended failure on the actual assigned
port. Browser failure evidence must distinguish an empty root from
loaded content. A saved task must remain usable after navigation and
reload.

**Steps to reproduce**

Run the exposure regression with the base port pair marked unavailable.
Run the browser-support tests with an unresolved entry module. Open a
saved task under the service worker, navigate to the same URL, and
reload it.

**Paperclip version or commit**

Based on master at a68f3d8e3.

## What Changed

- Make synthetic exposure failures use the assigned port. Test both free
and occupied base pairs.
- Record document readiness, root children, service-worker control,
pending module paths, and recent module error or 304 statuses in private
runner evidence.
- Add browser tests for pending startup, failed module responses, and
diagnostic reset after navigation.
- Add a provider-free browser test for persisted task content and the
composer across three navigation/reload cycles.
- Guard trusted report-job lockfile resolution before frozen install and
AWS credential setup.
- Document the new evidence and test commands.

## Verification

- `pnpm test:e2e:runner:unit`: 443 tests passed.
- `pnpm test:e2e:runner:browser-support`: 10 tests passed.
- Exposure tests: 28 passed; 3 platform-specific skips.
- Harness typecheck, workspace typecheck, and full build passed.
- Task-reload browser test passed against a throwaway local instance. It
checks three navigation/reload cycles with a controlling service worker.
- All PR verification suites passed, including server, chat, workspace,
serialized server, Rust, runner, build, typecheck, and browser shards.
The existing sidebar navigation case showed an empty page on the first
CI attempt and passed on one retry.
- The monolithic local `pnpm test:run` was started, then stopped after
CI completed the equivalent suites. Additional local browser
reproduction attempts also hit embedded-Postgres startup failures; the
focused checks listed above completed successfully.
- Greptile: 5/5 on the current head, with no findings.

## Risks

This change affects tests and private test evidence. It does not change
production prompts or runtime behavior. Module paths omit queries; the
collector reads no response bodies or headers. The existing sanitizer
and publication allowlist still apply. A 304 response does not cause a
test failure.

The blank-page cause remains unconfirmed. The first CI attempt
reproduced it in an existing sidebar navigation case: the trace shows an
empty page after a service-worker-mediated 304 response for the large
editor module. The single retry passed. This PR adds coverage and
diagnostics; it does not claim to fix that intermittent symptom.

## Model Used

OpenAI Codex, GPT-6, with repository tools, code execution, and browser
tests. The exact deployed model variant and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:04:13 -05:00
DottaandPaperclip 846336e5a0 test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path

> - Paperclip lets people manage agents through ongoing conversations.
> - Chat users can change instructions while a provider is already
working.
> - Existing chat evals wait for each turn to settle before the next
message.
> - They cannot prove delivery during active work or the saved effect of
a correction.
> - Existing fixtures also enable Agent Chat through the API rather than
the settings UI.
> - This PR adds bounded browser workflows and checks their persisted
outcomes.

## Linked Issues or Issue Description

Refs #13741, #13752, #13750.

**What happened?**

The chat suites cover planning, delegation, status, and recovery. They
lack active-turn follow-ups and the experimental settings lifecycle. A
sequential conversation can pass even if messages sent during work are
lost.

**Expected behavior**

A follow-up submitted during a provider turn survives and affects the
final reply. A changed launch day appears in the saved plan. Disabling
Agent Chat rejects new messages while preserving history; re-enabling
resumes the same conversation.

**Steps to reproduce**

Run the explicit `agent-chat-stories` suite. It selects three local
cases for each native Claude and Codex profile. An ordinary provider
command waits for a fixture brief file so the browser can send the
follow-up at an observed active-run boundary.

## What Changed

- Add six opt-in Product E2E cells for settings, active follow-ups, and
plan corrections.
- Drive experimental settings through the UI and verify disabled sends
are rejected by the public API.
- Use a bounded file wait in the actual isolated agent workspace, with
provider-written readiness and an undisclosed brief reference.
- Grade persisted user messages, final replies, native run outcomes, and
exact saved plan fields.
- Accept active-turn steering or one queued successor; reject lost
input, duplicate input, and stale outputs.
- Allow one steered run or two sequential runs throughout the shared
harness, while preserving exact counts for other cases.
- Require a single marker-bearing response attributed to the final
provider run.
- Unload the development browser client before restarting the server,
avoiding reconnect/navigation races without weakening the post-restart
memory check.
- Add browser regressions for restart isolation and asynchronously saved
settings switches.
- Document prepared-agent setup, native onboarding limits, and the
separate API-tool rollout gate.

## Verification

- Eval TypeScript check passed.
- Eval support suite: 436 tests passed in 39 files.
- New oracle calibration: six tests passed, including plausible invalid
outcomes.
- Browser support regressions: seven tests passed; the restart
regression was observed failing before the fix.
- Catalog discovery selects exactly six local native cases and leaves
default paid selection unchanged.
- [Consolidated existing native chat
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/):
master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a
browser navigation timeout across restart; the page request returned 200
and the chat rendered.
- [Nine targeted restart/replay
cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/)
passed on `1fe2fe275`, including the original failure, across native
Claude/Codex and local/Daytona; all cleanup passed.
- [Initial six-story
campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817)
retained all six failures: asynchronous switch assertions, unavailable
fixture paths, and rich-text escaping in raw command comparisons. The
corrected fixtures preserve the same behavioral assertions.
- [Six-story campaign
v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/)
on `8232773a0`: 4/6 passed (both settings cases and both Claude
interruptions). Codex could not see the host-temp fixture outside its
workspace; this failed before follow-up delivery was exercised.
- [Four affected interruption
cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/)
all passed, including cleanup, on definition v3 / `ad6ac0545`. Files
live inside the actual agent workspace and the observed run workspace is
verified. Both providers saved Friday in the real plan with the
undisclosed brief reference; follow-ups persisted while the original run
was active. Together with both unchanged settings cases from v2, all six
new scenario variants have passing live evidence.
- Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful
checks, two intentional skips, zero pending/failing checks; mergeable
and clean. Fresh Greptile 5/5, zero unresolved findings.
- Full typecheck, tests, build, and browser CI passed remotely. One
earlier head encountered a signoff-policy browser timing failure; the
final head passed that shard.
- Local pnpm wrapper could not fetch its version/signature metadata in
the restricted environment; local eval checks used the installed Node
executables. Repo-wide validation was completed by GitHub Actions.

## Risks

These are eval-only changes. The file wait is a timing fixture in the
isolated agent workspace, not a production runner hook. Native Codex
host-filesystem isolation stays unchanged. It has a two-minute limit and
is released in `finally`. The prepared-agent settings case is not full
native onboarding: the wizard currently offers legacy adapters. The
disabled-entry assertion uses full document navigation, which clears the
prior React Query cache; preserved history is checked through the public
API and re-enabled chat. No production prompt, rollout default, adapter
behavior, or credential policy changes. Active-task reassignment and
worker-crash recovery remain outside these new cases.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 16:18:06 -05:00
DottaandPaperclip 9d19f98b50 fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agent chat uses native runner sessions to plan, delegate, and track
that work.
> - A user can press Stop while the native session is still starting.
> - The server can acknowledge that Stop without dispatching it, then
let the session submit a turn.
> - This leaves chat recovery waiting for an execution that the user
expected to stop.
> - This PR waits for the startup handle, dispatches cancellation, and
prevents a late startup from submitting a turn.
> - New full-stack evals check the resulting records and outputs across
Claude and Codex.
> - Those evals also exposed missing ACPX readiness fields, unbounded
polling, and an old-run identity check that rejected valid warm
handoffs.

## Linked Issues or Issue Description

**What happened?**

Stop during native startup could record an acknowledged cancellation
with `dispatched: false`. The provider could then begin work. A
subsequent `/new` stayed queued. A remote Claude follow-up also
exhausted the command journal while probing warm-session readiness: ACPX
never returned the readiness fields required by the shared transport.
Once readiness worked, attachment incorrectly compared the next run
descriptor against the old run ID. The 25 ms polling loop could issue
4,800 commands during its two-minute wait, beyond the 500-command bound.
The existing chat eval treated lifecycle logs as proof of an active
provider turn, so it did not distinguish startup cancellation from
active-turn cancellation.

**Expected behavior**

A Stop during startup must reach the pending session. A late session
must not submit a prompt after Stop. Recovery must retain control when
startup exceeds the bounded wait. Chat evals must check saved task
state, document contents, worker identity, account binding, and
duplicate effects.

**Steps to reproduce**

1. Start a native Claude or Codex chat turn.
2. Press Stop after process startup is requested but before the provider
turn starts.
3. Send `/new`, then send a fresh message.
4. On the affected base, cancellation can be acknowledged without
dispatch and the reset stays queued.

**Paperclip version or commit**

The live Claude baseline reproduced this on `29d6b3509`. The branch also
includes master commit `0f5fafe16`.

Related work: #13678, #13686, #13693, #13291, #13738. A separate runner
reliability branch also contains a startup-wait fix. Its overlap must be
reconciled before merging; this branch additionally prevents prompt
submission after a late startup.

## What Changed

- Wait for a pending native startup before acknowledging a run-scoped
Stop. Preserve the existing recovery error when that wait expires.
- Keep a Stop guard on startup. Cancel a late handle before it can
submit a provider turn.
- Add regression tests for normal handle publication and publication
after the Stop deadline.
- Back off blocked warm-attachment probes. Keep the fast two-snapshot
barrier, fail closed, and record changed blockers.
- Add red/green tests for delayed readiness, persistent blockers,
alternating readiness, and readiness near the deadline.
- Publish ACPX readiness and blockers. Preserve the old authority’s
event acknowledgement barrier; only settled sessions can proceed to
attachment.
- Bind warm ACPX descriptors to the validated next authority while
retaining old-run event correlation until activation. Preserve session
identity and provider profile checks.
- Exercise two consecutive run rotations through a qualified fake
sidecar, verifying checkpointing, provider identity, pre-activation
rejection, and new-run work admission.
- Separate startup and active-turn cancellation checkpoints in the
browser eval.
- Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells
across Claude and Codex.
- Cover hiring and reuse through managed AI accounts, source-based
review, current blocked-task status, request replay after a lost HTTP
acknowledgement, server restart continuity, and Stop/reset continuity.
- Use ordinary production agent instructions. Enable API tools only for
the two coordination cases that need them.
- Calibrate the matchers with invalid records and outputs. Require
remembered context after restart and a structured status snapshot that
distinguishes the current blocker from history and task status from
active execution. Compare the public issue mutation contract and
relationships during read-only reporting. Preserve before/after source
records in failed eval evidence.
- Fix the lost-ack browser harness and verify it against a real HTTP
server. Check the chat composer after restart instead of waiting for an
unrelated document lifecycle event.
- Document the scope and limits of each case.

## Verification

- The startup regression failed on the unfixed executor and passed after
the fix.
- `pnpm test:e2e:runner:typecheck` passed.
- `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files.
- `pnpm exec vitest run
server/src/services/native-runtime/native-session-executor.test.ts`
passed: 385 tests.
- [Baseline live
campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868):
Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed.
Codex delegation was blocked by provider capacity.
- [Eval-only startup
campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786):
both providers failed as expected. Both persisted `dispatched: false`
and left `/new` queued.
- [First fixed
campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706)
on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both
providers. Failed cases exposed eval harness defects and remote
continuity failures. All attempts remain available.
- [Original workflows and stronger memory
checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649)
on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and
startup Stop passed for both providers; Codex remote restart passed.
Claude remote restart exposed the missing readiness contract. Two Codex
planning cells hit provider capacity.
- [Unchanged-model
retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548):
Codex planning and backlog creation both passed.
- [18-cell campaign with ACPX
readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963)
on `6a98ef743`: 16/18 passed, including all local/remote Stop and
committed-send cases. Claude remote continuity exposed the
next-authority check, now fixed. Codex hiring produced its checklist,
but the runner redacted the requested marker after it appeared as
“Tracking token: …”. That content-redaction policy is unchanged and
remains an explicit limitation.
- [Structured status
grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725)
on `50448c228`: both providers passed on their first attempt, including
cleanup.
- [Complete read-only state
grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011)
on `551e13892`: both providers passed.
- [Final ACPX handoff and hiring
retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456)
on `cbd637587`: all three Claude Daytona cases passed (restart
continuity, active Stop/reset, and lost-ack replay). Codex hiring
reproduced the content-redaction failure: the saved checklist contained
`Tracking token: [REDACTED]` instead of the required business marker.
All four cases completed cleanup successfully. [Published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/).
The only subsequent commit adds the qualified-sidecar integration test;
production code is identical to this live proof.
- `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without
paid models.
- Runner TypeScript typecheck passed. All 5 warm-readiness tests pass;
two failed with the prior fixed-rate loop, and the late-readiness test
failed before the pacing correction.
- ACPX readiness and warm-identity regressions each failed before their
fixes. All 292 runner-core Rust library tests passed. The
qualified-sidecar integration test passes. Rust formatting is checked.
- Status-grader regressions for misleading historical mentions and
previously unchecked mutations each failed before tightening the oracle
and pass now.
- [Latest-head
CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307)
passed on `a4093c8f1`: full build, type checks, test partitions, browser
E2E, and native runner checks. Two unrelated tests initially failed
(Sentry fixture release attribution and local-service fixture
readiness); both passed locally together (35 passed, 5 optional SDK
tests skipped) and on the failed-job retry. No changes were made to
those tests.
- Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed
and all review threads are resolved.
- The paid live suite is not fully green: the reproducible
content-redaction case remains red. This is separate from the passing PR
merge checks. No production content-redaction, prompt, model, or
completion-policy change is included.
- Managed-account hiring and review cases explicitly enable API tools;
these do not qualify default new-user onboarding.

## Risks

- Stop can wait up to 30 seconds for startup, then use the existing
pending-recovery path. This does not prove that remote cleanup has
finished.
- Blocked warm readiness adds up to 750 ms between later probes with the
two-minute remote budget, or about 32 ms with the default five-second
budget. Ready sessions retain the short second barrier.
- Paid evals can fail because of provider capacity or agent decisions.
Each failure needs evidence-based classification.
- The HTTP request replay case checks comment idempotency and duplicate
effects. It does not prove replay safety for an ambiguous provider tool
call.
- The new suite is opt-in. It does not increase the default paid
campaign.
- No production prompts or model selection change. Review-handoff
behavior and content-redaction policy remain separate product decisions.
The latter can remove harmless business content that looks like
credential syntax; the failing attempt is retained.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 10:35:38 -05:00
DottaandPaperclip e26d787928 Shorten continuation prompts and verify question tool guidance (#13574)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents must continue tasks using user answers without losing earlier
requirements or approval gates.
> - The wake prompt mixed human decisions with prior tool evidence and
repeated detailed question instructions.
> - Those instructions belong with the question tool, with a short
routing hint in the wake.
> - The Runner evals need to prove that answers, approvals, and
completed work survive later turns.
> - This PR shortens the prompts, separates authenticated answers, and
adds continuation tests with useful screenshots.

## Linked Issues or Issue Description

Refs #13517. This is a follow-up to the merged onboarding skill and
Runner E2E work. Related #13539 covers responses received while a run is
active; this PR preserves its cases and adds continuation coverage.
Existing continuation/recovery and question PRs were searched; none
covers this prompt/documentation and eval change.

**What existing behavior does this improve?**

The instructions sent when an agent continues a task, the native
human-input tool documentation, and the evidence captured by Runner
full-stack E2E.

**Current behavior**

The wake repeats a long question-tool guide. Human answers appear
alongside untrusted prior results. Screenshot capture can finish at DOM
load while the task still shows a spinner, even when backend behavior
checks pass.

**Proposed behavior**

Keep earlier requirements unless the user changes them. Treat
clarification as distinct from approval. Give authenticated human
responses a scoped field. Keep tool and agent results as evidence. Put
detailed question behavior in the tool descriptor and retain one routing
sentence in the native wake. Wait for the correct task and loaded
conversation before taking screenshots.

**Reason and benefit**

Reduce repeated prompt text and make authority boundaries clear. Test
that real question cards, later answers, approval gates, and completed
child tasks still work. Make screenshots useful for human review.

## What Changed

- Shorten shared continuation instructions for legacy and native
runners. Separate authenticated user responses from tool results and
agent summaries.
- Remove the detailed question guide from native wake prompts. Keep its
behavior in the canonical `request_human_input` descriptor and existing
payload schema. Regenerate semantic contracts and fixture hashes.
- Add five continuation cases across four local profiles. Add a
dedicated choice-then-text case for native Codex and native Claude. All
22 cells join the shared full E2E campaign.
- Cover revised scope, clarification without approval, hostile
instructions in a handoff file, and reuse of a completed child after
restart. Keep production instructions and fixed user facts.
- Capture continuation screenshots only when the intended task and
conversation have rendered. Add provider-free browser regressions for
loaders and wrong-task capture.
- Preserve current master’s extra tool and onboarding cases. The default
campaign now contains 166 cells; 35 manual everyday cells remain
separate.

## Verification

- `pnpm -r typecheck`: passed after replay on current master.
- `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed.
- `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm
test:e2e:runner:browser-support`: 4 passed. These tests failed against
immediate screenshot capture and passed after the fix.
- Focused continuation and native-input tests: 36 passed locally. The
tool-authority suite could not initialize embedded PostgreSQL locally,
including one isolated retry; its 17 assertions did not run locally. The
full remote server shards passed on this PR commit.
- `pnpm build`: passed after replay on current master. `pnpm test:run`
was attempted locally but hit the same embedded PostgreSQL
initialization failure; the remaining local run was stopped after
complete remote CI passed. This is not claimed as a full local test
pass.
- [Full PR
CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685):
passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All
server/chat/workspace/serialized shards, browser shards, Runner checks,
typecheck, build, canary and policy checks passed. The isolated native
Runner build and security checks also passed: 57 successful checks, with
two expected Storybook skips.
- Greptile reviewed the exact PR head at 5/5, with no findings or
unresolved review threads. The PR has no merge conflicts.
- [Live question-docs
report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/):
3/3 passed at source `83dd132f2` before replay on master. Native Codex
and Claude each asked a choice, waited, asked a text question, and saved
both answers. Claude also passed a completed-child restart case. All
three native turns are checked for absence of the old question block.
- [Earlier continuation
report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/):
all five continuation cases passed on native Claude. The report retains
campaign and revision provenance and separately shows two unresolved
onboarding behavior failures.
- [Before/after prompt
report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/):
full text, current recorded Claude inputs, and reproducible
reference-token counts. The controlled wake comparison removes 401
reference tokens; the net counted input reduction is 339 after charging
the larger tool description. These are text-size estimates, not measured
billing savings.

## Risks

- Prompt wording affects model behavior. Live results cover the stated
cases, not every provider or conversation. Legacy profiles are
registered but were not rerun for this change.
- The optional continuation field changes prompt data only; there is no
database migration or new production API.
- Authenticated answer projection excludes generated summaries and
agent-resolved interactions. It preserves the answer’s question or
approval scope.
- The screenshot guard can expose UI loading failures that earlier runs
hid. Backend grading alone no longer makes those captures valid.
- The two prior onboarding failures remain separate product issues: work
before acceptance and a missing saved plan. This PR does not claim the
entire onboarding suite passes.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, code execution,
and browser verification. The exact deployed model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted tests above; the
full local database-startup limit is documented
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 09:32:54 -05:00