## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
12 KiB
Runner E2E fixture authoring
The fixture catalog is executable production-contract data. Keep it small, typed, deterministic, and free of raw credentials.
Suites and matrices
A RunnerSuiteFixture declares one durable testing purpose: stable ID, label,
description, profiles, environments, cases, expected size, and definition or
ranking metadata. Its execution IDs are globally prefixed as
<suite>.<profile>.<environment>.<case>. Add a new suite when the testing
purpose or desired cross-product differs; do not inflate an existing suite with
unrelated dimensions.
The suite definition fingerprint is historical comparison metadata. Any profile, model qualification, environment, task, or ranking-snapshot change must change that fingerprint automatically so the dashboard can annotate the boundary instead of silently joining unlike totals.
Agent profiles
Add RunnerProfileFixture entries in catalog.ts. A profile declares:
- a stable ID and searchable groups;
- legacy or native generation;
- adapter/provider and required credential;
- a model imported from its adapter constant or qualified runner profile;
- supported environment IDs;
- expected runtime metadata; and
- an agent payload factory.
Do not duplicate model IDs, qualification decisions, CLI versions, or runner
artifact rules. Codex profiles import DEFAULT_CODEX_LOCAL_MODEL, OpenCode
profiles import QUALIFIED_OPENCODE_MODEL, and ACPX profiles import
QUALIFIED_ACPX_PROFILES. Add or qualify models at their owning production
source first.
OpenRouter breadth profiles are generated from openrouter-models.json, not
written by hand. That reviewed snapshot must contain exactly five unique,
available, tool-capable models with rank, canonical ID, display name, supported
parameters, source URL, capture time, and verified content hash. Refresh it
manually with pnpm test:e2e:runner:models:update; nightly campaigns never
change fixture definitions.
Agent adapterConfig.env values must be {type:"secret_ref", secretId, version:"latest"} objects supplied to the factory. A fixture source containing
a raw secret-looking value is rejected by catalog validation.
Environments
An EnvironmentFixture declares driver/provider, credential requirements,
attempt deadline, lifecycle behavior, expected execution target, and a payload
factory validated by the shared environment schema.
The local environment is instance-managed: company creation ensures it exists, and the public API intentionally rejects a second local environment. The setup registry therefore discovers that row through the public environments API. This still provides full isolation because every cell starts a new Paperclip instance and database.
Daytona creates sandbox environments through the public API. The core fixture
keeps reuseLease:false and runnerLifecycleMode:"per_turn". The dedicated
warm-continuity fixture uses reuseLease:true and
runnerLifecycleMode:"warm"; its distinct configurationKey is part of the
suite fingerprint even though both fixtures report environmentId:"daytona".
Keep short provider cleanup backstops, a Daytona secret reference, and an
immutable image digest. Teardown
must delete the environment with reusable-lease destruction and must fail the
cell if cleanup cannot be confirmed. Keep CPU, memory, and disk explicit: lease
metadata and the per-test public-list-price runtime estimate depend on that
pinned billable resource shape. Changing it requires updating billing tests and
reviewing the versioned Daytona rates in billing.ts.
Usage and billing data
Do not add fixture-authored token or dollar expectations. The live harness
reads usage from selected public heartbeat-run records and records coverage per
run. Provider-reported dollars remain distinct from runtime estimates. A zero
or missing native usage payload is unavailable unless a real token-bearing
receipt or provider cost proves otherwise. New execution environments must
provide lease/resource metadata for a runtime estimate or explicitly remain
unavailable; never infer that missing billing data means free execution.
Future providers (SSH, E2B, Modal, Cloudflare, Kubernetes, Novita, exe.dev)
should implement the same setup/probe/cleanup contract before being added to a
matrix. Unsupported profile/environment combinations belong in
supportedEnvironments, not in ad hoc test conditionals.
Task cases and matchers
A RunnerTaskFixture owns a work mode, a typed flow, expected run count,
nonce-based title/prompt/marker factories, per-environment attempt deadlines,
deterministic matchers, and expected terminal state. Single-turn prompts should
make one bounded request with observable output and no nondeterministic judging.
The plan_revision_acceptance flow must also provide revision-request and Plan
marker factories. question_resume_completion must define the deterministic
browser answer and prove exactly two successful runs with no pending
interaction. plan_approval_completion must target the exact two-step
canonical Plan revision, capture its pending UI, approve in the browser, and
prove exactly two successful runs. warm_three_turn provides exactly two
browser follow-up messages, preserves one project/execution-workspace scope,
verifies host file contents after every turn, and finishes within three
ten-minute turn deadlines. Native turns 1 and 2 include an actionable human review in the completion report's attentionRequests. Paperclip creates the review gate from that report. An explicit question-tool wait yields the turn and suppresses its final prose, so it is not interchangeable with this completion-review fixture. Turn 3 reports Done without another review.
Every selected case runs in its own isolated Paperclip process, and independent cases may run concurrently. Follow-up turns inside one case retain their shared task state. Each case creates and tears down its own company, secrets, environment selection, agent, and browser-created task. The current plan case proves three runs on the same issue: publish a two-step Plan, request a three-step revision through the UI, and accept the exact new revision through the UI before verifying implementation and Done.
The matcher union supports message exact/contains/regex/ordered checks, issue
and run state, runtime/environment metadata, files, artifacts, JSON paths, and
JSON Schema. The initial cases use normalized message_contains plus state,
runtime, and environment assertions; the plan flow additionally verifies
canonical document revision IDs, bodies, step counts, interaction targets, and
visible previews. Add matcher behavior and credential-free tests together.
Adding a task expands its suite's matrix. Update the suite's intentional size, the complete-catalog size, and credential-free unit tests in the same change. Paid tests never silently skip a missing credential or unsupported artifact.
New Paperclip object fixtures
The explicit-only lifecycle-baseline suite reuses this registry and existing
continuation, chat and governed-action flows. Its narrative pairs require actual
agent/run-attributed comments or exact visible responses. See
the live baseline contract for selectors and proof boundaries.
Register new objects in live-fixtures.ts with explicit dependencies in
FixtureRegistry. Setup must use a public API. Teardown runs in reverse order
and is invoked after partial setup failures. Direct database writes and private
test-only runner endpoints are prohibited.
The expected dependency shape is:
company
└── encrypted secrets
└── environment
└── agent
└── browser-created task
Projects, goals, apps, and configuration fixtures can be inserted into that graph without changing the launcher. Keep returned fixture state to IDs and sanitized metadata; never retain raw secret values.
Required checks
Run before a fixture change is reviewed:
pnpm test:e2e:runner:unit
pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner -- --list
Then run the narrowest paid cell that exercises the fixture. A full matrix is a manual or scheduled campaign, not a PR requirement.
Persistent chat fixtures
chat-cases.ts defines the eight-case agent-chat suite; chat-flow.ts drives the
production composer, plan revision/approval controls, questions, reset command,
and project cards. Keep its 28 local cells intentional. expectedRunCount
counts provider turns, including cancelled and handed-off task runs, but excludes
synthetic /new runs. Assertions must inspect all company runs because ordinary
issue lists exclude the source conversation. assertChatHandoff rejects missing
projects/plans, chat children, wrong assignees, and execution before plan commit.
Retained api-state.json, chat-handoff.json, and plan-revision evidence
include persisted comments, session generations, run context and logs, project
workspaces, task documents, and ordering. They pass through the normal sanitizer.
Screenshots are allowlisted to the exact disposable agent chat. Cleanup cancels
all active runs in the isolated company, including handed-off work; usage from
failed and cancelled runs must not disappear from campaign totals.
Warm three-turn continuity grades the exact workspace file after each turn, task completion, and sandbox/session identity. It also requires a visible persisted final reply with each turn marker once and in order. It does not grade exact final-reply wording; the hello and continuation fixtures retain those exact-response checks. This separates workspace persistence failures from model response-format variance.
chat-hardening.ts adds the explicit-only agent-chat-hardening journeys. Use
the ordinary public APIs to seed source documents and blockers. Keep the answer
out of the user's status/review request. Grade the exact source values, latest
blocker, preserved task identities, worker-authored output, and real executions.
The status request asks for JSON so the grader can distinguish the current
blocker from a historical mention and compare active-run count separately from
task status. The request must not reveal those expected values.
Capture the source after seeding and compare every field in the public issue
update contract, plus labels, dependencies, and dedicated-endpoint settings.
Derived inbound references may change when the chat legitimately cites a task.
The lost-acknowledgement probe may interrupt only the fixture browser's own
comment request after the real server has committed it. Retain its request ID
and replay that same request through the public API after restarting the server.
Never fabricate tool results or repair task state after a failed assertion.
chat-stories.ts seeds an ordinary file wait in the isolated agent's actual
home workspace; native Codex intentionally cannot see arbitrary host temp files.
The observed run workspace must match the fixture location. This is a deterministic interruption
boundary. The real provider command writes the readiness file and waits at most
two minutes. The harness must persist the next browser message while the same
run is active before supplying the brief. Always release the wait in finally.
Save boundary observations independently of the final outcome. The final answer
must recover a brief reference absent from both prompts; the revision oracle
also reads the actual conversation plan. Fixture setup never enables native API
tools for this suite. Do not describe its prepared-agent settings case as a
production onboarding qualification.
The agent-chat-qualification local fixtures use public APIs to seed two workers
and a task with a saved plan, or read-only tasks with contradictory historical
comments. Ordinary Node file waits in the isolated agent workspace establish
observable active execution; no provider output or database outcome is fabricated.
A worker-crash case sends SIGKILL only to a positively identified running native
worker PID, then uses the production Retry button. Each gate is released in a
finally block. Source facts and boundary state are retained with the attempt.
The lifecycle suite also includes two legacy disposition-repair probes. Their
first provider turn intentionally omits task disposition, and their second turn
must be an automatic, causally bound repair that records completion. They use
public task comments/status APIs and run-detail evidence; no private runtime
hooks or database mutations are used by the fixture.